Soak TestSAFE
Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.
Overview
Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.
42a36917b8beOBSERVED · 2026-10-06What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: soak-test
description: "Soak test protocol for extended play — what to observe and log for slow leaks, fatigue, late-appearing edge cases."
argument-hint: "[duration: 30m | 1h | 2h | 4h] [focus: memory | stability | balance | all]"
user-invocable: true
allowed-tools: Read, Glob, Grep, Write, Bash(bash "*/.claude/skills/soak-test/../../hooks/yaml-helper.sh" resolve_config *)
model: sonnet
---
!`bash "${CLAUDE_SKILL_DIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation`
**Automation mode**: Resolve `modes.automation` (`project.local.yaml` →
`project.yaml` → default `collaborative`). Every `AskUserQuestion` call and
every file write follows `.claude/docs/automation-modes.md`
(collaborative asks always · guided major-only · autonomous logs and proceeds;
`automation_always_ask` categories always prompt).
# Soak Test
A soak test (also called an endurance test) is an extended play session run
with specific observation goals. Unlike a smoke check (broad critical path,
~10 min) or a single-feature playtest (~30 min), a soak test runs for **30
minutes to several hours** to surface:
- **Memory leaks** — gradual heap growth that only appears after scene transitions
- **Performance drift** — frame time degradation that worsens over time
- **State accumulation bugs** — issues that only appear after N repetitions
of a mechanic (inventory full, score overflow, AI state corruption)
- **Fun fatigue** — mechanics that feel good in a first session but grow
repetitive over extended play
- **Content exhaustion** — the point where players run out of novel content
**This skill generates the observation protocol and analysis harness — the
human does the actual playing.**
**Output:** `production/qa/soak-test-[date]-[duration].md`
**When to run:**
- Polish phase — before `/gate-check release`
- After fixing a memory or stability issue (regression soak)
- When extended play has not been formally tracked
---
## 1. Parse Arguments
**Duration** (default: `1h`):
- `30m` — short soak; suitable for testing a single mechanic or scene
- `1h` — standard soak; covers most common leak categories
- `2h` — extended soak; recommended for first full Polish soak
- `4h` — deep soak; required for games with long session design (RPGs, sims)
**Focus** (default: `all`):
- `memory` — focus on heap size, object count, leak patterns
- `stability` — focus on crash/freeze/hang detection
- `balance` — focus on fun fatigue, content exhaustion, difficulty perception
- `all` — all of the above
---
## 2. Load Context
Read:
- `project.yaml` — `engine.name` (for engine-specific memory monitoring
guidance) and `performance.*` budgets (memory ceiling, target FPS); for any
key absent or empty (including when `project.yaml` has no `performance` or
`engine` block), fall back to `.claude/docs/technical-preferences.md`
- `design/gdd/game-concept.md` — intended session length (for comparison against
soak duration), core loop description (or `design/game-brief.md`, the one-page
brief that replaces it at `rigor: minimal` — core loop, and session length only
if its "Who it's for" line states one)
- Most recent file in `production/qa/playtests/` — prior playtest findings
(to avoid re-documenting known issues)
- Most recent file in `production/qa/qa-plan-*.md` — current sprint test coverage
(to understand what has been formally tested vs. what the soak covers)
Note any performance budget targets (`performance.*` from `project.yaml`, else `.claude/docs/technical-preferences.md`):
- Memory ceiling: [N MB, or "not set"]
- Target FPS: [N, or "not set"]
- Frame budget: [N ms, or "not set"]
---
## 3. Define Observation Checkpoints
Based on duration, generate timed checkpoints:
**30m soak**: T+0, T+10, T+20, T+30
**1h soak**: T+0, T+15, T+30, T+45, T+60
**2h soak**: T+0, T+20, T+40, T+60, T+80, T+100, T+120
**4h soak**: T+0, T+30, T+60, T+90, T+120, T+180, T+240
At each checkpoint, the observer records the observation items defined in
Phase 4.
---
## 4. Generate the Soak Test Protocol
### Memory / Stability observation items (if focus = memory or all)
Engine-specific monitoring guidance.
> **Record the unit the tool shows; never convert, and never assume one.**
> A soak test looks for **growth**, so every threshold below is a *ratio or a
> delta against this session's own T+0 baseline* — which is unit-agnostic and
> stays correct however the editor reports the number. Write the unit down at
> T+0 exactly as displayed and use it consistently for the rest of the run.
>
> This replaces a note asserting the return units of
> `Performance.get_monitor` — **NOT SOURCEABLE from `docs/engine-reference/`**,
> the identifier appears nowhere in it — a claim that sat one line under a row asking
> the tester to record "Static Memory (**KB**)". A wrong units claim in a leak
> detector is off by 1024× in the one measurement the protocol exists to take,
> and it would read as a plausible instruction throughout. Deltas need no such
> claim, so the safest fix was to stop needing it.
> **NOT SOURCEABLE — the tool, panel and counter names below are not covered by
> `docs/engine-reference/`**: Godot's Debugger → Monitors and its memory
> counters, Unity's Memory Profiler and its fields, Unreal's `stat memory`. Mark
> them so in the protocol: they are pointers for the tester to confirm in their
> own editor at T+0, not verified paths. If no memory tool can be found, nothing
> was measured, and the Verdict section's NOT ASSESSED applies.
**Godot 4:**
- Open Debugger → Monitors tab; track `Memory → Static Memory` and
`Object Count → Objects` across checkpoints
- Record: Static Memory (**unit as displayed**), Object Count, Orphan Nodes count
- Alert threshold: Memory growth > 20% from T+0 after the first 15 minutes
(some growth on load is expected; sustained growth indicates a leak)
- **Orphan Nodes is the one absolute number worth watching**: it should return to
its T+0 value after a scene unload. A ratio hides that;Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
42a36917b8befull audit observations/trust-audit/skill/donchitos__soak-test.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-06 | 42a36917b8be | SAFE | B | 89 | first audit |
Questions
What does the Soak Test skill do?
Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.
Is Soak Test safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Soak Test access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (42a36917b8be), read on 2026-10-06. The repository is watched, and a new audit runs when it changes — this is the first audit.