Explore To Sdk EvalsSAFE
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
Overview
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
a08ca1bc70aaOBSERVED · 2026-10-08What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: explore-to-sdk-evals
description: Convert MCPJam Explore-generated test cases into @mcpjam/sdk eval tests. Produces one test per case, exactly matching the user's prompts and expectations.
---
# explore-to-sdk-evals
## 1. Purpose
This skill converts **pre-existing MCPJam** Explore **test cases** (in the same document above, under `## Explore-generated test cases`) into runnable `@mcpjam/sdk` eval tests.
**Rules:**
- **Generate exactly the cases provided** — do not invent new cases, do not skip any, do not reword user prompts (use the exact query text from each case’s fenced block).
- The brief above already includes `## Tools`, resources/prompts if any, suggested scenarios, and `## Explore-generated test cases` — parse titles, prompts, negative flags, expected tool-call shapes, and expected output from there.
- For full SDK API (advanced validators, suite patterns, error handling), see the `create-mcp-eval` skill in the MCPJam SDK or package documentation — this file stays minimal.
---
## 2. Framework detection
Before emitting imports, inspect the target repo:
1. Read **`package.json`**: `scripts` and `devDependencies` for **`jest`** or **`vitest`**.
2. Look for **`jest.config.*`**, **`vitest.config.*`**, or **`vite.config.*`** that references Vitest.
**Apply:**
| Condition | Imports / runner |
| --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| Jest present | `describe`, `it`, `expect`, `beforeAll`, `afterAll` from `@jest/globals` (or rely on Jest globals); use `ts-jest` if TypeScript. |
| Vitest present (no Jest) | `import { describe, it, expect, beforeAll, afterAll } from "vitest"` |
| Neither present | **Default to Vitest** in generated scaffold (mention adding dep). |
| User explicitly wants no test framework | Use `EvalTest` / `EvalSuite` `.run()` in a plain `async function main()` script with manual assertions — no `describe`/`it`. |
---
## 3. Quick API reference (Explore flows only)
```typescript
import { MCPClientManager, HostRunner, createEvalRunReporter } from "@mcpjam/sdk";
import {
matchNoToolCalls,
matchToolCallWithPartialArgs,
matchToolArgument,
matchToolArgumentWith,
} from "@mcpjam/sdk";
```
**Connection**
- `const manager = new MCPClientManager();`
- HTTP: `await manager.connectToServer(SERVER_ID, { url: MCP_SERVER_URL });`
- Stdio: `await manager.connectToServer(SERVER_ID, { command: "node", args: ["./server.js"] });`
- `const tools = await manager.getToolsForAiSdk([SERVER_ID]);`
- `await manager.disconnectAllServers();` in `afterAll`
**Agent**
- `new HostRunner({ tools, model: MODEL, apiKey: LLM_API_KEY, maxSteps: 8, mcpClientManager: manager })` — pass the **same** `MCPClientManager` instance you used for `connectToServer` / `getToolsForAiSdk`. **`mcpClientManager` is required** for `widgetSnapshots` on reported results (MCP App HTML capture via `readResource`), which powers **MCP App replay in MCPJam eval traces**. Omitting it still runs tools correctly but traces in the app will have **no embedded widget replay**.
- **`apiKey` and `model` must match the LLM provider the user chose** (see §5); there is no default assumption that the key is an OpenAI key.
- `await runner.run(caseQuery)` — use **verbatim** `caseQuery` from Explore; for slow MCP + multi-step chains use `{ timeoutMs }` (or `timeout`) so the prompt does not abort before tools finish (see §6).
- Multi-turn only if a **single** Explore case clearly requires follow-up in one narrative; otherwise one prompt per `it()` (see §4c)
**Result inspection**
- `result.hasToolCall("tool_name")`
- `result.toolsCalled()` — `string[]`
- `result.getToolCalls()` — for validators
**Validators**
- Negative: `matchNoToolCalls(result.toolsCalled())`
- Partial args: `matchToolCallWithPartialArgs("tool", { key: value }, result.getToolCalls())` — use only when those values are **real literals** from Explore, not placeholders (§4).
- Single-arg exact / predicate: `matchToolArgument`, `matchToolArgumentWith` (see `create-mcp-eval` / SDK validators).
**Optional MCPJam upload**
- `createEvalRunReporter({ suiteName, apiKey, strict: true, mcpClientManager: manager, serverNames: [SERVER_ID], ... })` — pass the **same** connected **`MCPClientManager`** as for `HostRunner` so uploads include **`serverReplayConfigs`** (HTTP MCP only; stdio connections do not produce replayable server config in the SDK). That is what sets `hasServerReplayConfig` in MCPJam and enables **Replay this run** / **server-side MCP replay** in the UI.
- `await reporter.recordFromPrompt(result, { caseTitle, passed, expectedToolCalls?, isNegativeTest? })`
- `await reporter.finalize()` in `afterAll` **before** `disconnectAllServers()` (with timeout). `finalize` calls `getServerReplayConfigs()` on the manager; disconnecting first clears client state and uploads **without** replay metadata. **`finalize()` can throw** (network, DNS, 401) even when all eval assertions passed. For best-effort upload: wrap in **try/catch**, log a warning, and continue; or gate uploads on something like `MCPJAM_REPORTING=1` so local/CI runs do not go red purely on reporting.
- **UI replay still needs LLM keys in MCPJam Settings** for whatever `provider` the suite uses (e.g. OpenRouter); the app does not store provider API keys on the run.
---
## 4. Per-case translation
Use the **Explore case title** as the human-readable basis for `it("...")` description (sanitize quotes if needed). Use the **exact** user prompt from the ``` fenced block as `runner.run(\`...)` argument.
**PlaceTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
a08ca1bc70aafull audit observations/trust-audit/skill/mcpjam__explore-to-sdk-evals.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | a08ca1bc70aa | SAFE | B | 89 | first audit |
Questions
What does the Explore To Sdk Evals skill do?
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
Is Explore To Sdk Evals safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Explore To Sdk Evals access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (a08ca1bc70aa), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.