AutocontextSAFE
a recursive self-improving harness designed to help your agents (and future iterations of those agents) succeed on any task
Overview
a recursive self-improving harness designed to help your agents (and future iterations of those agents) succeed on any task
27843e93564fOBSERVED · 2026-10-09Install
Commands as the repository documents them. They are shown, not run.
npm install -g autoctx
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: autocontext description: > Iterative strategy generation and evaluation system. Use when the user wants to evaluate agent output quality, run improvement loops, queue tasks for background evaluation, check run status, inspect runtime artifacts and session branch lineage, or discover available scenarios. Provides LLM-based judging with rubric-driven scoring. allowed-tools: autocontext_judge autocontext_improve autocontext_status autocontext_scenarios autocontext_queue autocontext_runtime_snapshot --- # autocontext autocontext is an iterative strategy generation and evaluation system that uses LLM-based judging to score and improve agent outputs. ## Available Tools - **autocontext_judge** — Evaluate agent output against a rubric. Returns a 0–1 score with reasoning and per-dimension breakdowns. - **autocontext_improve** — Run a multi-round improvement loop. The agent output is judged, revised based on feedback, and re-evaluated until the quality threshold is met or max rounds are exhausted. - **autocontext_queue** — Enqueue a task for background evaluation by the task runner daemon. - **autocontext_status** — Check the status of runs and queued tasks. - **autocontext_scenarios** — List available evaluation scenarios and their families. - **autocontext_runtime_snapshot** — Inspect run artifacts, package provenance, branchable session lineage, and recent event-stream entries. ## Quick Start ### 1. Evaluate output quality Use `autocontext_judge` with a task prompt, the agent's output, and a rubric: ``` autocontext_judge( task_prompt="Write a Python function to parse CSV files", agent_output="def parse_csv(path): ...", rubric="Correctness, error handling, edge cases, documentation" ) ``` ### 2. Improve output iteratively Use `autocontext_improve` to automatically revise output through judge-guided feedback loops: ``` autocontext_improve( task_prompt="Write a Python function to parse CSV files", initial_output="def parse_csv(path): ...", rubric="Correctness, error handling, edge cases, documentation", max_rounds=5, quality_threshold=0.85 ) ``` ### 3. Queue background tasks Use `autocontext_queue` with a scenario name to enqueue evaluation tasks for asynchronous processing: ``` autocontext_queue(spec_name="my_scenario") ``` Check results later with `autocontext_status`. For deeper context, use `autocontext_runtime_snapshot` with the run ID. Add `session_id` when you need the active branch path before continuing work: ``` autocontext_runtime_snapshot(run_id="run_123", session_id="sess_123") ``` ### 4. Discover scenarios Use `autocontext_scenarios` to see what evaluation scenarios are available: ``` autocontext_scenarios() autocontext_scenarios(family="agent_task") ``` ## Configuration The extension auto-detects configuration from these sources: 1. **Project config** — `.autoctx.json` in the working directory (created via `autoctx init`) 2. **Environment variables:** - `AUTOCONTEXT_AGENT_PROVIDER` or `AUTOCONTEXT_PROVIDER` — Provider type - `AUTOCONTEXT_AGENT_API_KEY` or `AUTOCONTEXT_API_KEY` — Provider API key - `AUTOCONTEXT_AGENT_DEFAULT_MODEL` or `AUTOCONTEXT_MODEL` — Model override - `AUTOCONTEXT_DB_PATH` — SQLite database path override 3. **Pi provider** — Falls back to Pi's configured LLM provider ## CLI Companion For standalone usage outside Pi, install the `autoctx` CLI: ```bash npm install -g autoctx autoctx init autoctx solve "your problem" --iterations 5 autoctx simulate --description "your simulation" --runs 3 ```
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
27843e93564ffull audit observations/trust-audit/skill/greyhaven-ai__autocontext.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-09 | 27843e93564f | SAFE | B | 89 | first audit |
Questions
What does the Autocontext skill do?
a recursive self-improving harness designed to help your agents (and future iterations of those agents) succeed on any task
Is Autocontext safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Autocontext access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
What do I need installed to use Autocontext?
Its own instructions reference autoctx. Dependencies are pinned to exact versions.
How current is this page?
The grade is for one exact copy of the source (27843e93564f), read on 2026-10-09. The repository is watched, and a new audit runs when it changes — this is the first audit.