Atlas / Skills / mcpjam / Explore To Sdk Evals

Explore To Sdk EvalsSAFE

skills/mcpjam/explore-to-sdk-evals

Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
NOASSERTION
Stars
2,245
01

Overview

Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.

Read from source at commit a08ca1bc70aaOBSERVED · 2026-10-08
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: explore-to-sdk-evals
description: Convert MCPJam Explore-generated test cases into @mcpjam/sdk eval tests. Produces one test per case, exactly matching the user's prompts and expectations.
---

# explore-to-sdk-evals

## 1. Purpose

This skill converts **pre-existing MCPJam** Explore **test cases** (in the same document above, under `## Explore-generated test cases`)  into runnable `@mcpjam/sdk` eval tests.

**Rules:**

- **Generate exactly the cases provided** — do not invent new cases, do not skip any, do not reword user prompts (use the exact query text from each case’s fenced block).
- The brief above already includes `## Tools`, resources/prompts if any, suggested scenarios, and `## Explore-generated test cases` — parse titles, prompts, negative flags, expected tool-call shapes, and expected output from there.
- For full SDK API (advanced validators, suite patterns, error handling), see the `create-mcp-eval` skill in the MCPJam SDK or package documentation — this file stays minimal.

---

## 2. Framework detection

Before emitting imports, inspect the target repo:

1. Read **`package.json`**: `scripts` and `devDependencies` for **`jest`** or **`vitest`**.
2. Look for **`jest.config.*`**, **`vitest.config.*`**, or **`vite.config.*`** that references Vitest.

**Apply:**


| Condition                               | Imports / runner                                                                                                                     |
| --------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| Jest present                            | `describe`, `it`, `expect`, `beforeAll`, `afterAll` from `@jest/globals` (or rely on Jest globals); use `ts-jest` if TypeScript. |
| Vitest present (no Jest)                | `import { describe, it, expect, beforeAll, afterAll } from "vitest"`                                                                 |
| Neither present                         | **Default to Vitest** in generated scaffold (mention adding dep).                                                                    |
| User explicitly wants no test framework | Use `EvalTest` / `EvalSuite` `.run()` in a plain `async function main()` script with manual assertions — no `describe`/`it`.     |


---

## 3. Quick API reference (Explore flows only)

```typescript
import { MCPClientManager, HostRunner, createEvalRunReporter } from "@mcpjam/sdk";
import {
  matchNoToolCalls,
  matchToolCallWithPartialArgs,
  matchToolArgument,
  matchToolArgumentWith,
} from "@mcpjam/sdk";
```

**Connection**

- `const manager = new MCPClientManager();`
- HTTP: `await manager.connectToServer(SERVER_ID, { url: MCP_SERVER_URL });`
- Stdio: `await manager.connectToServer(SERVER_ID, { command: "node", args: ["./server.js"] });`
- `const tools = await manager.getToolsForAiSdk([SERVER_ID]);`
- `await manager.disconnectAllServers();` in `afterAll`

**Agent**

- `new HostRunner({ tools, model: MODEL, apiKey: LLM_API_KEY, maxSteps: 8, mcpClientManager: manager })` — pass the **same** `MCPClientManager` instance you used for `connectToServer` / `getToolsForAiSdk`. **`mcpClientManager` is required** for `widgetSnapshots` on reported results (MCP App HTML capture via `readResource`), which powers **MCP App replay in MCPJam eval traces**. Omitting it still runs tools correctly but traces in the app will have **no embedded widget replay**.
- **`apiKey` and `model` must match the LLM provider the user chose** (see §5); there is no default assumption that the key is an OpenAI key.
- `await runner.run(caseQuery)` — use **verbatim** `caseQuery` from Explore; for slow MCP + multi-step chains use `{ timeoutMs }` (or `timeout`) so the prompt does not abort before tools finish (see §6).
- Multi-turn only if a **single** Explore case clearly requires follow-up in one narrative; otherwise one prompt per `it()` (see §4c)

**Result inspection**

- `result.hasToolCall("tool_name")`
- `result.toolsCalled()` — `string[]`
- `result.getToolCalls()` — for validators

**Validators**

- Negative: `matchNoToolCalls(result.toolsCalled())`
- Partial args: `matchToolCallWithPartialArgs("tool", { key: value }, result.getToolCalls())` — use only when those values are **real literals** from Explore, not placeholders (§4).
- Single-arg exact / predicate: `matchToolArgument`, `matchToolArgumentWith` (see `create-mcp-eval` / SDK validators).

**Optional MCPJam upload**

- `createEvalRunReporter({ suiteName, apiKey, strict: true, mcpClientManager: manager, serverNames: [SERVER_ID], ... })` — pass the **same** connected **`MCPClientManager`** as for `HostRunner` so uploads include **`serverReplayConfigs`** (HTTP MCP only; stdio connections do not produce replayable server config in the SDK). That is what sets `hasServerReplayConfig` in MCPJam and enables **Replay this run** / **server-side MCP replay** in the UI.
- `await reporter.recordFromPrompt(result, { caseTitle, passed, expectedToolCalls?, isNegativeTest? })`
- `await reporter.finalize()` in `afterAll` **before** `disconnectAllServers()` (with timeout). `finalize` calls `getServerReplayConfigs()` on the manager; disconnecting first clears client state and uploads **without** replay metadata. **`finalize()` can throw** (network, DNS, 401) even when all eval assertions passed. For best-effort upload: wrap in **try/catch**, log a warning, and continue; or gate uploads on something like `MCPJAM_REPORTING=1` so local/CI runs do not go red purely on reporting.
- **UI replay still needs LLM keys in MCPJam Settings** for whatever `provider` the suite uses (e.g. OpenRouter); the app does not store provider API keys on the run.

---

## 4. Per-case translation

Use the **Explore case title** as the human-readable basis for `it("...")` description (sanitize quotes if needed). Use the **exact** user prompt from the ``` fenced block as `runner.run(\`...)` argument.

**Place
03

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha a08ca1bc70aafull audit observations/trust-audit/skill/mcpjam__explore-to-sdk-evals.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-08a08ca1bc70aaSAFEB89first audit
05

Questions

What does the Explore To Sdk Evals skill do?

Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.

Is Explore To Sdk Evals safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Explore To Sdk Evals access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (a08ca1bc70aa), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement