Atlas / Skills / florianbruniaux / Autoresearch

AutoresearchCAUTION

skills/florianbruniaux/autoresearch

The most comprehensive Claude Code guide: agentic workflows, hooks, skills, MCP servers, quizzes, and production-ready templates. 430K+ lines.

Verdict
CAUTION
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
CC-BY-SA-4.0
Stars
6,123
01

Overview

The most comprehensive Claude Code guide: agentic workflows, hooks, skills, MCP servers, quizzes, and production-ready templates. 430K+ lines.

Read from source at commit d90170da4369OBSERVED · 2026-10-07
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: autoresearch
description: "Autonomous improvement loop: scan codebase metrics, scaffold experiment files, run agent-driven iterations until metric improves"
argument-hint: "[--scaffold <loop-name>] [--run <loop-name>] [--status]"
effort: high
disable-model-invocation: true
---

# Autoresearch: Autonomous improvement loop

Scan codebase quality metrics, propose improvement loops, and run autonomous agent iterations. Inspired by [karpathy/autoresearch](https://github.com/karpathy/autoresearch), adapted from ML research to code quality.

**Concept**: The agent proposes a code change, runs the measurement, keeps the change if the metric improved, reverts via `git reset` if not, and repeats until manually stopped.

**Time**: Scan ~30s | Per iteration: depends on scope | Loop: runs indefinitely until you stop it

---

## Mode 1: Scan (default)

Measure current state, detect existing loops, propose next actions.

### Instructions

Run the following metrics and display a prioritized proposal table.

**Step 1: Measure codebase metrics**

Adapt grep patterns to your project's conventions. These are TypeScript defaults, adjust for your stack.

```bash
# M1: Function declarations (prefer arrow functions)
M1=$(grep -r "export function " src/ --include="*.ts" --include="*.tsx" -l 2>/dev/null | wc -l | tr -d ' ')

# M2: Interface declarations (prefer type aliases)
M2=$(grep -r "export interface " src/ --include="*.ts" --include="*.tsx" -l 2>/dev/null | wc -l | tr -d ' ')

# M3: ESLint disables
M3=$(grep -r "eslint-disable" src/ --include="*.ts" --include="*.tsx" -l 2>/dev/null | wc -l | tr -d ' ')

# M4: Type casts to any
M4=$(grep -r " as any" src/ --include="*.ts" --include="*.tsx" -l 2>/dev/null | wc -l | tr -d ' ')

# M5: TODO comments
M5=$(grep -r "// TODO" src/ --include="*.ts" --include="*.tsx" -l 2>/dev/null | wc -l | tr -d ' ')
```

**Step 2: Detect existing loops**

```bash
for dir in scripts/autoresearch/loop-*/; do
  [ -d "$dir" ] || continue
  LOOP_NAME=$(basename "$dir")
  # Check if loop has results
  if [[ -f "$dir/results.tsv" ]]; then
    ITERS=$(wc -l < "$dir/results.tsv" | tr -d ' ')
    BEST=$(sort -t$'\t' -k2 -n "$dir/results.tsv" | head -1 | cut -f2)
    echo "ACTIVE:$LOOP_NAME:iterations=$ITERS:best=$BEST"
  else
    echo "SCAFFOLDED:$LOOP_NAME"
  fi
done
```

**Step 3: Display**

```
Autoresearch Scan: {date}

Codebase metrics:

| # | Loop              | Metric            | Current | Target | Priority | Risk |
|---|-------------------|-------------------|---------|--------|----------|------|
| A | loop-remove-as-any| `as any` casts    | {M4}    | 0      | P1       | LOW  |
| B | loop-eslint-disable| eslint-disable   | {M3}    | 0      | P2       | MED  |
| C | loop-export-fn    | export function   | {M1}    | 0      | P1       | LOW  |
| D | loop-interface-type| export interface | {M2}    | 0      | P1       | LOW  |
| E | loop-todo-comments| TODO comments     | {M5}    | 0      | P3       | LOW  |

Existing loops: {detected loops or "none yet"}

Recommended next step (P1, LOW risk):
  /autoresearch --scaffold loop-remove-as-any
  Then write program.md, create a worktree, and run the loop.
```

---

## Mode 2: `--scaffold <loop-name>`

Generate the 3 mechanical files for a loop. **Does not generate `program.md`**: write that yourself to encode project-specific constraints.

### Instructions

Create the following files under `scripts/autoresearch/{loop-name}/`:

**`measure.sh`**: the evaluation harness (single metric, returns an integer):

```bash
#!/usr/bin/env bash
# measure.sh: {loop-name}
# Returns an integer. Direction: lower = better (unless loop targets coverage/score).
set -euo pipefail
grep -r "PATTERN" src/ --include="*.ts" --include="*.tsx" 2>/dev/null | wc -l | tr -d ' '
```

**`direction.txt`**: improvement direction:

```
lower
```

(Use `higher` for metrics like test coverage or quality score.)

**`files.txt`**: scope the agent should operate on:

```
src/
```

After creating the files, display:

```
Loop scaffolded: scripts/autoresearch/{loop-name}/

  measure.sh  : {pattern} in {scope} -> {N} occurrences today
  direction   : lower (fewer = better)
  files.txt   : src/

Current metric: {N} (target: 0)

Next steps:
  1. Write program.md -- agent behavior, constraints, what it can/cannot touch
     Reference: scripts/autoresearch/loop-remove-as-any/program.md
  2. Create a worktree: /worktree feature/autoresearch-{loop-name}
  3. cd into the worktree
  4. bash scripts/autoresearch/runner.sh {loop-name} 0 15
```

---

## Mode 3: `--run <loop-name>`

Execute the autonomous loop. The agent runs indefinitely: stop it manually when satisfied.

### Instructions

**Verify prerequisites:**

```bash
[ -f "scripts/autoresearch/{loop-name}/measure.sh" ] || { echo "ERROR: measure.sh missing. Run --scaffold first."; exit 1; }
[ -f "scripts/autoresearch/{loop-name}/program.md" ] || { echo "ERROR: program.md missing. Write it first, this encodes your constraints."; exit 1; }
```

**Run the loop:**

Read `scripts/autoresearch/{loop-name}/program.md` fully before starting. Then enter the following cycle, repeat until stopped:

```
LOOP ITERATION #{N}

1. Current metric: bash scripts/autoresearch/{loop-name}/measure.sh
2. Read program.md constraints
3. Propose ONE targeted change to files in files.txt
4. Apply the change
5. Re-measure: bash scripts/autoresearch/{loop-name}/measure.sh
6. Evaluate:
   - direction=lower AND new < previous -> KEEP (git add -p && git commit -m "autoresearch: {description}")
   - otherwise -> REVERT (git checkout -- .)
7. Log to results.tsv: {timestamp}\t{metric}\t{status}\t{description}
8. Continue to iteration #{N+1}
```

**Stopping criteria** (from program.md):
- Metric reaches target (e.g., 0)
- No more mechanical changes possible
- User manually stops the process

**Display each iteration:**

```
[iter #{N}] metric: {before} -> {after} | {KEPT/REVERTED} | {change description}
```

---

## Mode 4: `--status`

Show status of all loops in the
03

Trust audit

CAUTIONgrade B · trust 89/100 Install with care. The audit found things worth knowing before you trust its output.

LayerWhat it checksResult
L0Provenance & inventoryWARN
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (2)

MEDIUMInventory / provenance · inv.symlink · CWE-1104
whitepapers/recap-cards/en/_extensions
whitepapers/recap-cards/en/_extensions
Why it matters. link not followed
MEDIUMInventory / provenance · inv.symlink · CWE-1104
whitepapers/recap-cards/fr/_extensions
whitepapers/recap-cards/fr/_extensions
Why it matters. link not followed

Gates applied: no_behavioural_pass.

Audited 2026-10-07 · audit v0.4.1 · source sha d90170da4369full audit observations/trust-audit/skill/florianbruniaux__autoresearch.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-07d90170da4369CAUTIONB89first audit
05

Questions

What does the Autoresearch skill do?

The most comprehensive Claude Code guide: agentic workflows, hooks, skills, MCP servers, quizzes, and production-ready templates. 430K+ lines.

Is Autoresearch safe to install?

With care. The audit graded it B (89/100) and found 2 things worth knowing before you trust this skill, listed below with the exact line each was found on.

What can Autoresearch access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (d90170da4369), read on 2026-10-07. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement