Create Mcp EvalSAFE
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
Overview
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
a08ca1bc70aaOBSERVED · 2026-10-08What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: create-mcp-eval
description: Generate comprehensive eval tests for any MCP server using @mcpjam/sdk. Supports Jest and Vitest with deterministic and LLM-driven test patterns.
---
# create-mcp-eval
Generate eval tests for MCP servers using **@mcpjam/sdk**.
Read this file first. It carries the two things you need before writing any code — what to ask the user, and the rules the generated tests must follow — and routes you to the rest only when you actually need it.
## Reference map
Load a reference when you reach the step that needs it, not before.
| You are about to... | Read |
|---|---|
| Scaffold `package.json`, `tsconfig.json`, `.env.example`, `.gitignore` | `references/project-setup.md` |
| Call `MCPClientManager`, `HostRunner`, `PromptResult`, `EvalTest`, `EvalSuite`, validators, or the MCPJam reporter | `references/sdk-api.md` |
| Choose a shape — config block, toggled suites, shared reporter, parameterized agents, save modes, multi-turn, validator coverage | `references/patterns.md` |
| Write the file out | `references/template.md` |
| Debug a test that runs but behaves oddly | `references/common-mistakes.md` |
| Turn an MCPJam **Agent Brief** into tests | `references/agent-brief.md` |
## 1. Context Gathering
Before generating any code, collect the following from the user:
| Question | Options | Default |
|----------|---------|---------|
| **Connection type** | `stdio` (local binary) or `http` (SSE/Streamable HTTP URL) | `http` |
| **Test framework** | `jest`, `vitest`, or `none` (SDK-only) | _(detect from repo; fall back to `vitest`)_ |
| **LLM provider** | See Supported Providers table below. Format: `provider/model` | _(must ask user)_ |
| **Save results to MCPJam** | `none`, `auto` (saves when MCPJAM_API_KEY is set), or `reporter` (shared EvalRunReporter). Use an MCPJam API key (`sk_...`) from **Settings → API keys**; optionally set `MCPJAM_PROJECT_ID` to file results under a specific project (defaults to the org’s Default project). | _(must ask user)_ |
| **Tool list** | Ask user to paste their tool names or an **Agent Brief** (`references/agent-brief.md`) | — |
If the user provides an **Agent Brief** (markdown with `## Tools` table), parse it to auto-populate tool names, descriptions, parameters, and suggested eval scenarios. See `references/agent-brief.md`.
### Provider Selection (REQUIRED)
You MUST ask the developer which LLM provider they want before generating any code. Do not default to any provider.
**Supported Providers:**
| Provider | Model format | Env var | Example model |
|----------|-------------|---------|---------------|
| `openai` | `openai/<model>` | `OPENAI_API_KEY` | `openai/gpt-4o-mini` |
| `anthropic` | `anthropic/<model>` | `ANTHROPIC_API_KEY` | `anthropic/claude-sonnet-4-20250514` |
| `google` | `google/<model>` | `GOOGLE_API_KEY` | `google/gemini-2.0-flash` |
| `mistral` | `mistral/<model>` | `MISTRAL_API_KEY` | `mistral/mistral-small-latest` |
| `deepseek` | `deepseek/<model>` | `DEEPSEEK_API_KEY` | `deepseek/deepseek-chat` |
| `xai` | `xai/<model>` | `XAI_API_KEY` | `xai/grok-2` |
| `openrouter` | `openrouter/<model>` | `OPENROUTER_API_KEY` | `openrouter/openai/gpt-4o-mini` |
| `azure` | `azure/<deployment>` | `AZURE_API_KEY` | `azure/gpt-4o` |
| `ollama` | `ollama/<model>` | _(none, local)_ | `ollama/llama3` |
| Custom | `<name>/<model>` | _(configurable)_ | `litellm/gpt-4` |
Once the user selects a provider, use the corresponding env var name and model format in all generated code:
- `{LLM_ENV_VAR}` — e.g., `OPENAI_API_KEY`
- `{LLM_MODEL}` — e.g., `openai/gpt-4o-mini`
- `{LLM_KEY_EXAMPLE}` — e.g., `sk-...`
### Test Runner Selection
Before generating tests, check what the codebase already uses:
- `package.json` scripts and devDependencies for `jest` or `vitest`
- Config files: `jest.config.*`, `vitest.config.*`, `vite.config.*`
Then:
- If Jest is present, use Jest (and `ts-jest` if TypeScript).
- If Vitest is present, use Vitest.
- If neither is present, default to Vitest.
- If the developer prefers **no test framework**, the `@mcpjam/sdk` classes (`EvalTest`, `EvalSuite`) can run standalone — call `.run()` directly and check results in a plain script without Jest/Vitest.
In all cases, use `@mcpjam/sdk` for the eval harness (`HostRunner`, `EvalTest`, `EvalSuite`, validators).
---
## 5. Generation Guidelines
Follow these rules when generating eval test files:
1. **Deterministic suite first** — always include a deterministic test section using `HostRunner.mock()` that validates the test structure itself without requiring LLM calls or server connections.
2. **One EvalTest per tool** — create a separate `EvalTest` for each tool you want to evaluate. Each test should prompt the runner with a natural-language request and assert the correct tool was selected.
3. **Single-shot LLM tests are non-deterministic** — a single `runner.run()` may not select the expected tool every time. For single-shot tests, prefer saving results to MCPJam without hard-asserting (`expect(...).toBe(true)`). Use `EvalTest` with `iterations >= 3` and assert on `accuracy()` for reliable pass/fail gates. Reserve hard asserts for high-confidence cases (negative tests, multi-turn with clear context).
4. **Write unambiguous prompts for similar tools** — when a server has tools with overlapping descriptions (e.g., `create_view` vs `export_to_excalidraw`), prompts must reference the tool's *unique* action. Mention specific verbs, targets, or outcomes. Bad: "Share my diagram". Good: "Export and upload my diagram to excalidraw.com so I can open it in a browser".
5. **Multi-turn for related tools** — when tools logically chain together (e.g., `get_user` then `list_workspaces`), create a multi-turn test using `{ context: previousResult }`.
6. **Negative test** — always include at least one test that verifies the runner does NOT call tools when given an irrelevant prompt (e.g., "What is the capital of France?"). Use `matchNoToolCalls()`.
7. **Reasonable deTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
a08ca1bc70aafull audit observations/trust-audit/skill/mcpjam__create-mcp-eval.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | a08ca1bc70aa | SAFE | B | 89 | first audit |
Questions
What does the Create Mcp Eval skill do?
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
Is Create Mcp Eval safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Create Mcp Eval access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (a08ca1bc70aa), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.