Cost BenchmarkCAUTION
🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Overview
🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
ef7d4f0535e5OBSERVED · 2026-09-25What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: cost-benchmark
description: Run the corpus benchmark — booster locally, optional Gemini/Sonnet/Opus baselines — and persist a verifiable measured-vs-claimed table
argument-hint: "[--llm] [--anthropic]"
allowed-tools: Bash
---
# Cost Benchmark
Runs `scripts/bench.mjs` against the structural+adversarial corpus and writes per-case + summary results to `docs/benchmarks/runs/`. This is the verification gate that backs every measurable claim in `cost-booster-edit` / `cost-booster-route`.
## When to use
- Before publishing a release — verify booster win rate didn't regress.
- After expanding `bench/booster-corpus.json` — confirm new cases route correctly.
- When auditing a "claimed upstream" tag — flip it to "verified" once the bench supports it.
- On a cost question ("is Sonnet 4.6 cheaper than Opus 4.7 for these tasks?") — re-run with `BENCH_ANTHROPIC=1`.
## Steps
1. **Run the bench from `v3/`** (where `agent-booster` resolves):
```bash
( cd v3 && node ../plugins/ruflo-cost-tracker/scripts/bench.mjs ) # booster only — free, ~85 ms
( cd v3 && BENCH_LLM_BASELINE=1 node ../plugins/ruflo-cost-tracker/scripts/bench.mjs ) # + Gemini 2.0 Flash (cheap)
( cd v3 && BENCH_LLM_BASELINE=1 BENCH_ANTHROPIC=1 \
node ../plugins/ruflo-cost-tracker/scripts/bench.mjs ) # + Sonnet 4.6 + Opus 4.7
```
2. **Inspect the markdown summary** printed to stdout. The gate metric is `winRate` (Tier 1 cases). Adversarial cases are tracked separately as `escalationRate`.
3. **Persisted output** lands at:
- `docs/benchmarks/runs/latest.json` — pointer to the most recent run
- `docs/benchmarks/runs/<ISO-timestamp>.json` — historical record
4. **Read it back** in subsequent skills (e.g. `cost-report` step 2 reads `latest.json` for live tier-spend numbers).
## Smoke gates
- `winRate ≥ 0.80` on Tier 1 cases (smoke step 23). Lower the threshold by editing `scripts/smoke.sh`.
- `escalationRate` is reported but ungated — adversarial cases are diagnostic.
## Env overrides
| Env var | Default | Purpose |
|---|---|---|
| `BENCH_LLM_BASELINE` | unset | `=1` runs the OpenAI-compat baseline |
| `BENCH_LLM_MODEL` | `models/gemini-2.0-flash` | Override the OpenAI-compat model |
| `BENCH_LLM_BASE_URL` | Gemini OpenAI shim | Override endpoint |
| `BENCH_ANTHROPIC` | unset | `=1` runs Anthropic baseline (Sonnet 4.6 + Opus 4.7) |
| `BENCH_ANTHROPIC_MODELS` | `claude-sonnet-4-6,claude-opus-4-7` | Comma-separated Claude IDs |
| `BENCH_OUT` | timestamped file | Override output path |
| `BENCH_QUIET=1` | unset | Suppress markdown summary |
API keys auto-pulled from `gcloud secrets` (`GOOGLE_AI_API_KEY`, `ANTHROPIC_API_KEY`); override with `BENCH_LLM_API_KEY` / `BENCH_ANTHROPIC_API_KEY`.
## Cross-references
ADR-0002 §"Decision 1" / §"Riskiest assumption" · `cost-booster-edit/SKILL.md` (verification table consumes this skill's output) · `cost-report/SKILL.md` step 2 (reads `runs/latest.json`).Trust audit
CAUTIONgrade B · trust 89/100 Install with care. The audit found things worth knowing before you trust its output.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | WARN |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (4)
crates
plugin/agents
plugin/commands
plugin/skills
Gates applied: no_behavioural_pass.
ef7d4f0535e5full audit observations/trust-audit/skill/ruvnet__cost-benchmark.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-09-25 | ef7d4f0535e5 | CAUTION | B | 89 | first audit |
Questions
What does the Cost Benchmark skill do?
🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
Is Cost Benchmark safe to install?
With care. The audit graded it B (89/100) and found 4 things worth knowing before you trust this skill, listed below with the exact line each was found on.
What can Cost Benchmark access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (ef7d4f0535e5), read on 2026-09-25. The repository is watched, and a new audit runs when it changes — this is the first audit.