Atlas / Skills / harbor-framework / Create Task

Create TaskCAUTION

skills/harbor-framework/create-task

Framework for evaluating and improving agents

Verdict
CAUTION
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
Apache-2.0
Stars
5,916
01

Overview

Framework for evaluating and improving agents

Read from source at commit 59ec1b8a02cfOBSERVED · 2026-10-08
02

Install

Commands as the repository documents them. They are shown, not run.

uvx --from 'harbor-rewardkit==0.2.*' rewardkit /tests
uvx harbor-rewardkit
uvx --with pytest==8.4.1 pytest /tests/test_outputs.py
03

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: create-task
description: Create a new Harbor task for evaluating agents. Use when the user wants to 
  scaffold, build, or design a new task, benchmark problem, or eval. Guides through 
  instruction writing, environment setup, verifier design (pytest vs Reward Kit vs 
  custom), and solution scripting.
---

Guide the user through creating a new Harbor task end-to-end. Don't just dump commands — 
walk them through each decision, especially around the verifier (which is usually the 
hardest part).

## Step 1: Scaffold the task

```bash
harbor task init "<org>/<task-name>"
```

Useful flags:
- `--description "..."`
- `--author "Jane Doe <[email protected]>"` (repeat for multiple authors)
- `--no-pytest` — skip the pytest test template (use if planning Reward Kit or custom verifier)
- `--no-solution` — skip solution/ directory
- `--metadata-template path.toml` — pre-populate task.toml

Produces:
```
<task-name>/
├── instruction.md         # Task prompt for the agent
├── task.toml              # Config and metadata
├── environment/Dockerfile # Container definition
├── solution/solve.sh      # Reference solution (optional)
└── tests/test.sh          # Verifier script
```

If the user wants a **multi-step task** (ordered steps with per-step
instructions, tests, and early stopping against a shared container), scaffold
the single-step layout first, then convert to the `steps/` layout described in
the *Multi-step tasks* section below.

## Step 2: Write instruction.md

This is the prompt the agent receives. Help the user write it clearly:

- **State the goal concretely** — what file to create, what behavior to produce
- **Specify expected outputs** — paths, formats, content
- **Include constraints** — language, tools, approach
- **Don't leak the tests** — describe what "done" looks like, not how you'll check it

Example (from the ssh-key-pair tutorial):
```markdown
# SSH Key Pair Generation

Generate an SSH key pair in the files `~/.ssh/id_rsa` and `~/.ssh/id_rsa.pub`.

Don't make them password protected.
```

## Step 3: Build the environment

Edit `environment/Dockerfile` to install dependencies the task needs. The agent works 
inside this container.

```dockerfile
FROM ubuntu:24.04
WORKDIR /app

# Install what the task requires — NOT the solution
RUN apt-get update && apt-get install -y openssh-client && rm -rf /var/lib/apt/lists/*
```

For multi-container setups, use `environment/docker-compose.yaml` instead (note: most 
cloud sandbox providers only support Dockerfile).

**Test the environment interactively** before writing the solution or tests:
```bash
harbor task start-env -p "<task-path>" -e docker -a -i
```

This is usually where task authors realize something is missing from the Dockerfile.

## Step 4: Decide how to verify

**This is the most important decision.** Ask the user: *"How do you want to grade this 
task?"* Then help them pick:

Also ask: *"Should the verifier run in the same environment as the agent, or in a
separate verifier environment?"*

- Use the default shared environment when tests need to inspect the agent's full
  workspace, installed tools, or services.
- Use a separate verifier environment when grading code, dependencies, API keys,
  or OS requirements should stay hidden from the agent, or when verification
  should run from a clean image.

For a separate verifier container, `tests/` is the verifier image build context
and the image must provide `/tests/test.sh` (Linux) or `/tests/test.bat`
(Windows). Harbor copies `/logs/artifacts` and configured artifacts into the
verifier environment, not the agent's whole workspace.

```toml
[verifier]
environment_mode = "separate"

[verifier.environment]
docker_image = "ubuntu:24.04"
```

### Option A: Reward Kit (recommended for most cases)

Use when the verifier has multiple criteria, needs partial credit, uses an LLM/agent 
judge, or would benefit from composable reusable checks. See the `rewardkit` skill.

Good fit signals:
- Multiple things to check (file exists + content correct + command works)
- Subjective quality dimensions (readability, correctness of prose)
- Want partial credit rather than pass/fail
- Want to compose built-ins like `file_contains`, `command_succeeds`, `json_key_equals`

`tests/test.sh`:
```bash
#!/bin/bash
uvx --from 'harbor-rewardkit==0.2.*' rewardkit /tests
```

Note: the package is named `harbor-rewardkit` but the executable is `rewardkit`,
hence `--from 'harbor-rewardkit==0.2.*' rewardkit`. Running
`uvx harbor-rewardkit` directly will fail.

Then add `tests/checks.py` and/or `tests/judge.toml`. Invoke the `rewardkit` skill to 
design the criteria.

### Option B: pytest (good for deterministic unit-style checks)

Use when the verification is straightforward assertion-style Python. Default template if 
`--no-pytest` wasn't passed.

`tests/test.sh`:
```bash
#!/bin/bash
apt-get update && apt-get install -y curl
curl -LsSf https://astral.sh/uv/0.9.7/install.sh | sh
source $HOME/.local/bin/env

uvx --with pytest==8.4.1 pytest /tests/test_outputs.py

if [ $? -eq 0 ]; then
  echo 1 > /logs/verifier/reward.txt
else
  echo 0 > /logs/verifier/reward.txt
fi
```

Example `tests/test_outputs.py`:
```python
from pathlib import Path

def test_file_exists():
    assert (Path.home() / ".ssh" / "id_rsa").exists()
```

### Option C: Custom shell

For simple single-command checks (e.g. a binary pass/fail from one command):
```bash
#!/bin/bash
if diff -q /app/output.txt /tests/expected.txt; then
  echo 1 > /logs/verifier/reward.txt
else
  echo 0 > /logs/verifier/reward.txt
fi
```

### Reward file format (all options)

- `/logs/verifier/reward.txt` — single number (usually `0` or `1`)
- `/logs/verifier/reward.json` — `{"accuracy": 0.95, "runtime_sec": 1.2}` for multiple metrics

**Always use absolute paths in `test.sh`.**

## Step 5: Write the solution

Write `solution/solve.sh` — a script that actually solves the task. The Oracle agent runs 
this to sanity-check that the task is solvable and the tests pass o
04

Trust audit

CAUTIONgrade B · trust 89/100 Install with care. The audit found things worth knowing before you trust its output.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)WARN
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (3)

MEDIUMSupply chain · prompt.pipe_to_shell · CWE-829, CWE-1357
SKILL.md:142
curl -LsSf https://astral.sh/uv/0.9.7/install.sh | sh
LOWInventory / provenance · inv.symlink · CWE-1104
CLAUDE.md
CLAUDE.md
Why it matters. link not followed
LOWInventory / provenance · inv.symlink · CWE-1104
docs/CLAUDE.md
docs/CLAUDE.md
Why it matters. link not followed

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha 59ec1b8a02cffull audit observations/trust-audit/skill/harbor-framework__create-task.json · Report an issue / request a re-scan
05

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-0859ec1b8a02cfCAUTIONB89first audit
06

Questions

What does the Create Task skill do?

Framework for evaluating and improving agents

Is Create Task safe to install?

With care. The audit graded it B (89/100) and found 3 things worth knowing before you trust this skill, listed below with the exact line each was found on.

What can Create Task access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (59ec1b8a02cf), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement