Test Evidence ReviewSAFE
Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.
Overview
Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.
42a36917b8beOBSERVED · 2026-10-06What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: test-evidence-review
description: "Quality review of test files and evidence — goes beyond existence, evaluates assertion coverage. ADEQUATE/INCOMPLETE/MISSING/NOT ASSESSED per story."
argument-hint: "[story-path | sprint | system-name]"
user-invocable: true
allowed-tools: Read, Glob, Grep, Write, Bash(bash "*/.claude/skills/test-evidence-review/../../hooks/yaml-helper.sh" resolve_config *)
model: sonnet
---
!`bash "${CLAUDE_SKILL_DIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation,workflow,qa.level,testing.strict`
**Automation mode**: Resolve `modes.automation` (`project.local.yaml` →
`project.yaml` → default `collaborative`). Every `AskUserQuestion` call and
every file write follows `.claude/docs/automation-modes.md`
(collaborative asks always · guided major-only · autonomous logs and proceeds;
`automation_always_ask` categories always prompt).
**`qa.level` and `testing.strict`** (resolved above) decide what this review may
call a blocker — the same rules `/story-done` closes stories by, so a BLOCKING
item here is one that would stop `/story-done`:
- **`qa.level: minimal`** waives tests. A Logic, Integration or Config/Data story
with no test evidence (for Config/Data, no smoke report) is
`WAIVED (qa.level: minimal)` — not MISSING — and raises no BLOCKING item; a
test that does exist is still reviewed, but its gaps are ADVISORY. Visual/Feel and UI evidence (the retained screenshots, and the
Visual/Feel sign-off) is never waived.
- **Gate level per story type**: take `testing.strict.<type>` from the resolved
block (Logic→`logic`, Integration→`integration`, Visual/Feel→`visual`,
UI→`ui`, Config/Data→`config`): `true` → BLOCKING, `false` → ADVISORY, `unset`
→ a plain-boolean `testing.strict: true|false` (legacy form) applies to every
type; else the Default Gate Level in `.claude/docs/coding-standards.md`
(BLOCKING for Logic, Integration, Visual/Feel and UI; ADVISORY for
Config/Data). Every "BLOCKING" below means that story type's gate level.
# Test Evidence Review
`/smoke-check` verifies that test files **exist** and **pass**. This skill
goes further — it reviews the **quality** of those tests and evidence documents.
A test file that exists and passes may still leave critical behaviour uncovered.
A manual evidence doc that exists may lack the sign-offs required for closure.
**Output:** Summary report (in conversation) + optional `production/qa/evidence-review-[date].md`
**When to run:**
- Before QA hand-off sign-off (`/team-qa` Phase 5)
- On any story where test quality is in question
- As part of milestone review for Logic and Integration story quality audit
---
## 1. Parse Arguments
**Modes:**
- `/test-evidence-review [story-path]` — review a single story's evidence
- `/test-evidence-review sprint` — review all stories in the current sprint
- `/test-evidence-review [system-name]` — review all stories in an epic/system
- No argument — ask which scope: "Single story", "Current sprint", "A system"
---
## 2. Load Stories in Scope
Based on the argument:
**Single story**: Read the story file directly. Extract: Story Type, Test
Evidence section, story slug, system name.
**Sprint**: Read the most recently modified file in `production/sprints/`; extract
the list of story file paths from the sprint plan.
**System**: Glob `production/epics/[system-name]/story-*.md`.
> **If the resolved scope contains ZERO stories, stop here.** Report
> `NOT ASSESSED — no stories in scope`, name which scope was searched and which
> path was empty, and route: no sprint file → `/sprint-plan new` (at `workflow: minimal`,
> which has no sprints, `/test-evidence-review [epic-slug]` instead); a sprint plan
> listing no stories → `/create-stories [epic-slug]`; a `[system-name]` glob that
> matched nothing → name the glob. Do not continue to Section 3.
>
> **Guard the empty scope, not just the per-story unknown.** The verdict
> vocabulary here — ADEQUATE / INCOMPLETE / MISSING — needs a "could not check"
> value, or an unverifiable story acquires a verdict claiming somebody verified
> it; that is what `NOT ASSESSED` is for. **But giving the per-story unknown a
> home does nothing for the empty-scope unknown.** With no stories
> the Section 6 report renders an empty Summary table and ends
> `BLOCKING items: 0 / ADVISORY items: 0` — which reads as *everything reviewed,
> all fine*. This skill gates story closure, and `coding-standards.md` marks Logic,
> Integration, Visual/Feel and UI evidence BLOCKING, so a false-clean closes stories nobody
> reviewed.
>
> This is a recurring shape: the sophisticated inner rule present, the outer
> boundary unguarded. Ask it of any skill that aggregates —
> **what does this emit when the set is empty?**
For the resulting story set, collect the fields below with **targeted section
greps, not a full read of each story**:
```
Grep pattern="## Test Evidence" glob="production/epics/**/story-*.md" output_mode="content" -A 8
Grep pattern="## Acceptance Criteria" glob="production/epics/**/story-*.md" output_mode="content" -A 15
```
- **Story Type** (Logic / Integration / Visual/Feel / UI / Config/Data) and the
stated evidence path — both live under `## Test Evidence`, so the first grep's
`-A 8` captures them.
- Acceptance Criteria list — the `## Acceptance Criteria` block from the second grep.
- Story slug (from the file name) and System (from the directory path) — no read.
Full-read a story only when its Test Evidence section is missing or ambiguous.
(In Sprint mode, scope the globs to the sprint plan's story paths.)
---
## 3. Locate Evidence Files
For each story, find the evidence. **Check the story's stated evidence path first**
— the one Section 2 captured from its `## Test Evidence` section. If it exists, use
it. Fall back to searching only when no path is stated or it is not found:
**Logic stories**: search the engine's unit-test root — `tests/unit/[system]/`
(Godot), `Assets/Tests/EditMode/` (Unity), `Source/<Module>/PrivatTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
42a36917b8befull audit observations/trust-audit/skill/donchitos__test-evidence-review.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-06 | 42a36917b8be | SAFE | B | 89 | first audit |
Questions
What does the Test Evidence Review skill do?
Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.
Is Test Evidence Review safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Test Evidence Review access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (42a36917b8be), read on 2026-10-06. The repository is watched, and a new audit runs when it changes — this is the first audit.