Aeon AutoresearchSAFE
Bankr Skills equip builders with plug-and-play tools to build more powerful agents.
Overview
Bankr Skills equip builders with plug-and-play tools to build more powerful agents.
4029e336cef5OBSERVED · 2026-10-09What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: aeon-autoresearch
description: |
Evolve any installed skill by generating four variations along separate theses (better inputs /
sharper output / more robust / rethink), scoring them on a weighted rubric, and applying the
winner. Never downgrades a working skill — aborts cleanly if no variation improves the original.
Use when an installed skill is producing low-signal output, hitting deprecated APIs, or feels
stale.
Triggers: "improve this skill", "evolve $skill", "auto-research my X", "regenerate variations".
---
# aeon-autoresearch
Self-improvement loop. Given a target SKILL.md, generates four parallel improved variations, scores each, applies the winner.
## Inputs
| Param | Description |
|---|---|
| `target` | Skill name or path to SKILL.md. Required. |
| `mode` | `evolve` (default) writes the diff. `dry-run` scores and prints, writes nothing. |
## The four variations
- **A — Better inputs**: replace deprecated APIs, add fallbacks, fix broken endpoints.
- **B — Sharper output**: tighter format, signal over noise, explicit verdicts, banned filler.
- **C — More robust**: empty-data handling, retries, dedup state, rate-limit awareness.
- **D — Rethink**: fundamentally different methodology for the same goal.
Each is a complete runnable SKILL.md. Frontmatter shape preserved.
## Scoring
1-5 per axis, weighted total max 50:
| Axis | Weight |
|---|---|
| Improvement vs original | 3× |
| Output value | 2× |
| Clarity, data quality, robustness | 1.5× each |
| Conventions | 1× |
Tie-break (within 2 points): prefer the variation making the biggest single improvement over many small ones.
## Safety guarantee
If every variation scores ≤ original on **Improvement**, the skill aborts with `AUTORESEARCH_NO_IMPROVEMENT`. No file written. Working skills are never downgraded.
Preserves the original's core purpose, frontmatter shape, and declared env vars.
## Versioning
Inside a git repo, changes land in a branch (`autoresearch/${target}`) — operator reviews the diff before merging. Outside a repo, the original is preserved at `${target}/SKILL.md.before-autoresearch` for rollback.
## Output
The diff, plus a report with the scoring table for all four variations and a one-paragraph rationale for the winner.
Pairs with `aeon-skill-evals` (surfaces what's underperforming) and `aeon-skill-repair` (deterministic bugs; autoresearch handles quality lifts).Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
4029e336cef5full audit observations/trust-audit/skill/bankrbot__aeon-autoresearch.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-09 | 4029e336cef5 | SAFE | B | 89 | first audit |
Questions
What does the Aeon Autoresearch skill do?
Bankr Skills equip builders with plug-and-play tools to build more powerful agents.
Is Aeon Autoresearch safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Aeon Autoresearch access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (4029e336cef5), read on 2026-10-09. The repository is watched, and a new audit runs when it changes — this is the first audit.