Create TaskCAUTION
Framework for evaluating and improving agents
Overview
Framework for evaluating and improving agents
59ec1b8a02cfOBSERVED · 2026-10-08Install
Commands as the repository documents them. They are shown, not run.
uvx --from 'harbor-rewardkit==0.2.*' rewardkit /tests
uvx harbor-rewardkit
uvx --with pytest==8.4.1 pytest /tests/test_outputs.py
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: create-task description: Create a new Harbor task for evaluating agents. Use when the user wants to scaffold, build, or design a new task, benchmark problem, or eval. Guides through instruction writing, environment setup, verifier design (pytest vs Reward Kit vs custom), and solution scripting. --- Guide the user through creating a new Harbor task end-to-end. Don't just dump commands — walk them through each decision, especially around the verifier (which is usually the hardest part). ## Step 1: Scaffold the task ```bash harbor task init "<org>/<task-name>" ``` Useful flags: - `--description "..."` - `--author "Jane Doe <[email protected]>"` (repeat for multiple authors) - `--no-pytest` — skip the pytest test template (use if planning Reward Kit or custom verifier) - `--no-solution` — skip solution/ directory - `--metadata-template path.toml` — pre-populate task.toml Produces: ``` <task-name>/ ├── instruction.md # Task prompt for the agent ├── task.toml # Config and metadata ├── environment/Dockerfile # Container definition ├── solution/solve.sh # Reference solution (optional) └── tests/test.sh # Verifier script ``` If the user wants a **multi-step task** (ordered steps with per-step instructions, tests, and early stopping against a shared container), scaffold the single-step layout first, then convert to the `steps/` layout described in the *Multi-step tasks* section below. ## Step 2: Write instruction.md This is the prompt the agent receives. Help the user write it clearly: - **State the goal concretely** — what file to create, what behavior to produce - **Specify expected outputs** — paths, formats, content - **Include constraints** — language, tools, approach - **Don't leak the tests** — describe what "done" looks like, not how you'll check it Example (from the ssh-key-pair tutorial): ```markdown # SSH Key Pair Generation Generate an SSH key pair in the files `~/.ssh/id_rsa` and `~/.ssh/id_rsa.pub`. Don't make them password protected. ``` ## Step 3: Build the environment Edit `environment/Dockerfile` to install dependencies the task needs. The agent works inside this container. ```dockerfile FROM ubuntu:24.04 WORKDIR /app # Install what the task requires — NOT the solution RUN apt-get update && apt-get install -y openssh-client && rm -rf /var/lib/apt/lists/* ``` For multi-container setups, use `environment/docker-compose.yaml` instead (note: most cloud sandbox providers only support Dockerfile). **Test the environment interactively** before writing the solution or tests: ```bash harbor task start-env -p "<task-path>" -e docker -a -i ``` This is usually where task authors realize something is missing from the Dockerfile. ## Step 4: Decide how to verify **This is the most important decision.** Ask the user: *"How do you want to grade this task?"* Then help them pick: Also ask: *"Should the verifier run in the same environment as the agent, or in a separate verifier environment?"* - Use the default shared environment when tests need to inspect the agent's full workspace, installed tools, or services. - Use a separate verifier environment when grading code, dependencies, API keys, or OS requirements should stay hidden from the agent, or when verification should run from a clean image. For a separate verifier container, `tests/` is the verifier image build context and the image must provide `/tests/test.sh` (Linux) or `/tests/test.bat` (Windows). Harbor copies `/logs/artifacts` and configured artifacts into the verifier environment, not the agent's whole workspace. ```toml [verifier] environment_mode = "separate" [verifier.environment] docker_image = "ubuntu:24.04" ``` ### Option A: Reward Kit (recommended for most cases) Use when the verifier has multiple criteria, needs partial credit, uses an LLM/agent judge, or would benefit from composable reusable checks. See the `rewardkit` skill. Good fit signals: - Multiple things to check (file exists + content correct + command works) - Subjective quality dimensions (readability, correctness of prose) - Want partial credit rather than pass/fail - Want to compose built-ins like `file_contains`, `command_succeeds`, `json_key_equals` `tests/test.sh`: ```bash #!/bin/bash uvx --from 'harbor-rewardkit==0.2.*' rewardkit /tests ``` Note: the package is named `harbor-rewardkit` but the executable is `rewardkit`, hence `--from 'harbor-rewardkit==0.2.*' rewardkit`. Running `uvx harbor-rewardkit` directly will fail. Then add `tests/checks.py` and/or `tests/judge.toml`. Invoke the `rewardkit` skill to design the criteria. ### Option B: pytest (good for deterministic unit-style checks) Use when the verification is straightforward assertion-style Python. Default template if `--no-pytest` wasn't passed. `tests/test.sh`: ```bash #!/bin/bash apt-get update && apt-get install -y curl curl -LsSf https://astral.sh/uv/0.9.7/install.sh | sh source $HOME/.local/bin/env uvx --with pytest==8.4.1 pytest /tests/test_outputs.py if [ $? -eq 0 ]; then echo 1 > /logs/verifier/reward.txt else echo 0 > /logs/verifier/reward.txt fi ``` Example `tests/test_outputs.py`: ```python from pathlib import Path def test_file_exists(): assert (Path.home() / ".ssh" / "id_rsa").exists() ``` ### Option C: Custom shell For simple single-command checks (e.g. a binary pass/fail from one command): ```bash #!/bin/bash if diff -q /app/output.txt /tests/expected.txt; then echo 1 > /logs/verifier/reward.txt else echo 0 > /logs/verifier/reward.txt fi ``` ### Reward file format (all options) - `/logs/verifier/reward.txt` — single number (usually `0` or `1`) - `/logs/verifier/reward.json` — `{"accuracy": 0.95, "runtime_sec": 1.2}` for multiple metrics **Always use absolute paths in `test.sh`.** ## Step 5: Write the solution Write `solution/solve.sh` — a script that actually solves the task. The Oracle agent runs this to sanity-check that the task is solvable and the tests pass o
Trust audit
CAUTIONgrade B · trust 89/100 Install with care. The audit found things worth knowing before you trust its output.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | WARN |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (3)
curl -LsSf https://astral.sh/uv/0.9.7/install.sh | sh
CLAUDE.md
docs/CLAUDE.md
Gates applied: no_behavioural_pass.
59ec1b8a02cffull audit observations/trust-audit/skill/harbor-framework__create-task.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 59ec1b8a02cf | CAUTION | B | 89 | first audit |
Questions
What does the Create Task skill do?
Framework for evaluating and improving agents
Is Create Task safe to install?
With care. The audit graded it B (89/100) and found 3 things worth knowing before you trust this skill, listed below with the exact line each was found on.
What can Create Task access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (59ec1b8a02cf), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.