RewardkitSAFE
Framework for evaluating and improving agents
Overview
Framework for evaluating and improving agents
59ec1b8a02cfOBSERVED · 2026-10-08Install
Commands as the repository documents them. They are shown, not run.
uvx --from 'harbor-rewardkit==0.2.*' rewardkit /tests
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: rewardkit
description: Write Harbor task verifiers using Reward Kit. Use when creating or editing a
task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing
verifiers that produce a reward score.
---
Help the user write task verifiers with Reward Kit. Reward Kit is a lightweight Python
package that turns a directory of criteria files into a reward score. Each criterion is a
Python function call or a TOML judge file; folders become separate rewards.
## Setup in a Harbor task
Put criteria alongside `test.sh` in the task's `tests/` directory:
```
tests/
├── test.sh
├── checks.py # programmatic criteria
└── judge.toml # optional LLM/agent judge
```
`tests/test.sh`:
```bash
#!/bin/bash
uvx --from 'harbor-rewardkit==0.2.*' rewardkit /tests
```
This runs all criteria in `/tests/` against the workspace at `/app` and writes
`/logs/verifier/reward.json`. Defaults match Harbor's conventions — no extra config needed.
If judge criteria need API keys, pass them through `task.toml`:
```toml
[verifier.env]
ANTHROPIC_API_KEY = "${ANTHROPIC_API_KEY}"
```
Ask whether Reward Kit should run in the agent's shared environment or in a
separate verifier environment. Prefer a separate verifier environment when judge
prompts, grading dependencies, API keys, or clean-room checks should not be
available to the agent:
```toml
[environment]
network_mode = "no-network" # Agent env baseline — offline during agent.run()
[verifier]
environment_mode = "separate"
[verifier.environment]
network_mode = "public" # Verifier env baseline — LLM judge API calls
docker_image = "python:3.12-slim"
```
In shared mode, the verifier runs in the agent container and inherits
`[environment].network_mode`. Put `[verifier].network_mode` only when verify()
needs different network access than the agent phase (a phase override, not a
baseline). If agent and verifier need different baselines without runtime
switching, use `environment_mode = "separate"` and set
`[verifier.environment].network_mode`.
Judge criteria that call external APIs need a `public` baseline or allowlist on
the verifier environment. Programmatic checks that only read local files can use
`no-network`.
In separate mode, `tests/` is the verifier image build context and must provide
`/tests/test.sh` at runtime; Harbor does not upload `tests/` into the running
verifier container.
## Programmatic criteria
Call built-ins from any `.py` file in `tests/`:
```python
import rewardkit as rk
rk.file_exists("output.txt")
rk.file_contains("output.txt", "hello")
rk.command_succeeds("python main.py", weight=2.0)
rk.json_key_equals("result.json", "status", "ok")
```
All criteria accept `weight` (default `1.0`) and `isolated` (default `False`, runs in
overlayfs so side effects don't leak).
### Available built-ins
- **Files**: `file_exists`, `file_not_exists`, `file_contains`, `file_contains_regex`,
`file_matches`, `files_equal`, `diff_ratio`
- **Commands**: `command_succeeds`, `command_output_contains`, `command_output_matches`,
`command_output_matches_regex` (30s default timeout, optional `cwd`)
- **Data**: `json_key_equals`, `json_path_equals`, `csv_cell_equals`, `xlsx_cell_equals`
(needs `[office]` extra), `sqlite_query_equals`
- **HTTP**: `http_status_equals`, `http_response_contains`
- **Images**: `image_similarity`, `image_size_equals` (needs `[image]` extra)
- **Trajectory**: `trajectory_tool_used`, `trajectory_tool_not_used`, `trajectory_turn_count`
For extras, install with `uv tool install harbor-rewardkit[all]`.
## Custom criteria
Use the `@criterion` decorator. First parameter is always `workspace: Path`. Returns
`bool`, `float`, or a dict with `score` plus optional `reasoning`, `confidence`, and `model`:
```python
from pathlib import Path
from rewardkit import criterion
@criterion
def has_valid_output(workspace: Path) -> bool:
return (workspace / "output.txt").read_text().strip() != ""
```
Zero-parameter criteria auto-register. Criteria with extra args must be called via `rk`:
```python
@criterion(description="output has at least {n} lines")
def has_n_lines(workspace: Path, n: int) -> bool:
return len((workspace / "output.txt").read_text().splitlines()) >= n
rk.has_n_lines(10, weight=2.0)
rk.has_n_lines(50, weight=1.0)
```
For criteria shared across reward subdirs, define with `shared=True` in a root-level file
and call from subdirs.
## Judge criteria (LLM or agent-as-a-judge)
For subjective checks (quality, readability, edge cases), create a TOML file:
```toml
[judge]
judge = "anthropic/claude-opus-5-5" # LiteLLM model string
files = ["/app/main.py"]
[[criterion]]
description = "Is the code correct?"
type = "binary"
[[criterion]]
description = "How readable is the code?"
type = "likert"
points = 5
weight = 2.0
```
Criterion types:
- `binary` — yes/no → 1.0 or 0.0
- `likert` — 1..points, normalized to [0, 1]
- `numeric` — min..max, normalized to [0, 1]
- `rubric` — 2 to 10 described `levels` forming a scale from worst to best; position sets the score
### Agent judges
Agent judges shell out to a CLI and can explore the filesystem:
```toml
[judge]
judge = "claude-code"
model = "anthropic/claude-opus-5-5"
isolated = true
[[criterion]]
description = "Does the solution handle edge cases?"
type = "binary"
```
Slower and more expensive than LLM judges, but they can run commands and inspect files.
### JEV judge
JEV is a new type of language model from TypeSafe. It answers each criterion with a probability or a rubric
score and returns no reasoning, so it is fast and cheap. It needs the `jev` extra
(`harbor-rewardkit[jev]`) and `TYPESAFE_API_KEY`.
```toml
[judge]
judge = "jev"
files = ["/app/answer.md"]
[[criterion]]
description = "Does the answer address the requested task?"
[[criterion]]
description = "How complete is the answer?"
type = "rubric"
levels = ["Omits the information", "Covers part of it", "Covers all of it"]
```
Binary criteria pass aTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (2)
CLAUDE.md
docs/CLAUDE.md
Gates applied: no_behavioural_pass.
59ec1b8a02cffull audit observations/trust-audit/skill/harbor-framework__rewardkit.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 59ec1b8a02cf | SAFE | B | 89 | first audit |
Questions
What does the Rewardkit skill do?
Framework for evaluating and improving agents
Is Rewardkit safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Rewardkit access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
What do I need installed to use Rewardkit?
Its own instructions reference public and sqlite_query_equals. Dependencies are pinned to exact versions.
How current is this page?
The grade is for one exact copy of the source (59ec1b8a02cf), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.