Create AdapterSAFE
Framework for evaluating and improving agents
Overview
Framework for evaluating and improving agents
59ec1b8a02cfOBSERVED · 2026-10-08Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| codex | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: create-adapter
description: Scaffold a new Harbor benchmark adapter by running `harbor adapter init` and then guide implementation using the Adapters Agent Guide as the authoritative spec.
---
# Create Adapter
Bootstrap a new benchmark adapter in the Harbor repository. This skill scaffolds the adapter directory with `harbor adapter init` and then defers to the adapter tutorial for every implementation decision.
## Authoritative reference
The adapter tutorial is the authoritative specification for this skill. Read it in full before taking any action beyond scaffolding:
```
adapters/ADAPTER_CONTRIBUTING.md
```
That path is relative to the Harbor repo root (a skill prerequisite — see below). The tutorial contains:
- Required directory structures for the generated tasks and the adapter code package.
- Step-by-step process (steps 1-9) covering benchmark analysis, oracle verification, parity experiments, dataset registration, and submission.
- Schemas for `task.toml`, `parity_experiment.json`, `adapter_metadata.json`, and `dataset.toml`.
- Parity matching criterion, pre-flight checklist, and debug playbook.
- README format rules (machine-parsed; deviations break automation).
Do not substitute prior knowledge for the contents of that file. Treat it as the contract.
## Prerequisites
- Harbor CLI is installed and available on `PATH` (`harbor --version` succeeds).
- Working directory is the Harbor repository root.
- Harbor checkout is current: run `git fetch origin && git status` and pull `main` if behind. Stale checkouts miss recent adapter and agent fixes and are a common source of spurious parity failures later on.
## Workflow
### 1. Read the tutorial
Before any filesystem or CLI action, Read `adapters/ADAPTER_CONTRIBUTING.md` in full. Pay particular attention to:
- **Required Directory Structures** — the contract for generated task and adapter layouts.
- **Step 1. Understand the Original Benchmark** — what to identify upstream before coding.
- **Step 8. Register the Dataset → Naming rules** — the `name` field and `<org>/<task>` format requirements.
### 2. Gather benchmark context from the user
Collect the following before scaffolding. If the user has not provided an item, ask before proceeding — these inputs shape the scaffold and the tutorial steps that follow.
| Field | Why it matters |
|-------|---------------|
| Adapter name | Lowercase, hyphen-separated. Must match the benchmark's common identifier (e.g., `swe-bench`, `aider-polyglot`). Becomes the directory under `adapters/` and, after underscore conversion, the Python package name. |
| Human-readable name | Passed via `--name`; appears in the generated README. |
| Upstream repo URL | Needed for tutorial Step 1 (benchmark analysis) and for the `original_parity_repo` field later. |
| Oracle solutions available? | If the benchmark ships reference solutions, use them. If not, oracle solutions must be LLM-generated (tutorial Step 3 → "Benchmarks without oracle solutions"). |
| Agent scenario | Which of Step 4's scenarios applies: (1) existing compatible agent, (2) fork + add LLM agent, or (3) custom agent. This shapes the parity plan. |
### 3. Create the adapter branch
Per tutorial Step 2, work on a dedicated branch:
```bash
git checkout -b <adapter-name>-adapter
```
### 4. Run the scaffold
Prefer the non-interactive form when both names are known:
```bash
harbor adapter init <adapter-name> --name "<Human-Readable Name>"
```
Otherwise run `harbor adapter init` interactively and let the CLI prompt.
Expected output: a new directory at `adapters/<adapter-name>/` containing `pyproject.toml`, `README.md`, `src/<adapter_name>/`, and the template task files under `src/<adapter_name>/task-template/`. Verify the directory exists before continuing.
### 5. Hand off to the tutorial
Continue from "Step 1. Understand the Original Benchmark" in the tutorial. Do not invent structure, field names, or workflow beyond what the guide specifies.
**High-priority gotchas** (each is documented in the tutorial, but these are the most common adapter-build failures — surface them proactively as you work through the steps):
- **Every generated `task.toml` must contain a `name` field under `[task]`.** `main.py` is responsible for deriving a sanitized, unique, registry-safe name for every task. Tasks without a `name` cannot be registered. See the tutorial's "Naming rules" table.
- **Task names must be stable across adapter runs.** Unstable names churn registry digests on republish. If upstream lacks stable identifiers, mint a deterministic scheme (e.g., `{dataset}-1`, `{dataset}-2`) from a reproducible sort.
- **Use `schema_version = "1.4"` at the top of `task.toml`.** Set package versions separately with `[task].version` in `task.toml` and `[dataset].version` in `dataset.toml`. Request any desired registry tags in the PR description.
- **`main.py` must support `--output-dir`, `--limit`, `--overwrite`, and `--task-ids`.** These flags are required for reproducible runs and task-level debugging.
- **The generated `README.md` is parsed by downstream automation.** Fill in every section exactly as the template defines; put extra context in the **Notes** section or in the `notes` fields of `parity_experiment.json` / `adapter_metadata.json`. Do not add, rename, reorder, or remove sections.
- **Do not run parity experiments unilaterally.** Tutorial Step 4 requires team coordination on agents, models, and number of runs before incurring API costs. Complete sanity checks first, and execute full runs symmetrically on both sides.
## Reference adapters by scenario
When implementation questions come up, point at an existing adapter that matches the benchmark's shape:
| Scenario | Example adapter | When to use |
|----------|----------------|-------------|
| Compatible agent already exists | `adapters/adebench/` | Upstream already supports Claude-Code / Codex / OpenHands / Gemini-CLI |
| Fork upstream + add LLM agent | `adapters/evoeval/` | LLM-Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (2)
CLAUDE.md
docs/CLAUDE.md
Gates applied: no_behavioural_pass.
59ec1b8a02cffull audit observations/trust-audit/skill/harbor-framework__create-adapter.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 59ec1b8a02cf | SAFE | B | 89 | first audit |
Questions
What does the Create Adapter skill do?
Framework for evaluating and improving agents
Is Create Adapter safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Create Adapter access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Create Adapter work with?
Its documentation mentions codex. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (59ec1b8a02cf), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.