Atlas / Skills / donchitos / Test Flakiness

Test FlakinessSAFE

skills/donchitos/test-flakiness

Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
MIT
Stars
25,788
01

Overview

Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.

Read from source at commit 42a36917b8beOBSERVED · 2026-10-06
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: test-flakiness
description: "Find flaky tests from CI logs — aggregates pass rates, spots intermittent failures, recommends quarantine. After multiple runs."
argument-hint: "[ci-log-path | scan | registry]"
user-invocable: true
allowed-tools: Read, Glob, Grep, Write, Edit, Bash, Bash(bash "*/.claude/skills/test-flakiness/../../hooks/yaml-helper.sh" resolve_config *)
model: sonnet
---

!`bash "${CLAUDE_SKILL_DIR}/../../hooks/yaml-helper.sh" resolve_config --keys automation`

**Automation mode**: Resolve `modes.automation` (`project.local.yaml` →
`project.yaml` → default `collaborative`). Every `AskUserQuestion` call and
every file write follows `.claude/docs/automation-modes.md`
(collaborative asks always · guided major-only · autonomous logs and proceeds;
`automation_always_ask` categories always prompt).

# Test Flakiness Detection

A flaky test is one that sometimes passes and sometimes fails without any code
change. Flaky tests are worse than no tests in some ways — they train the team
to ignore red CI runs, masking genuine failures. This skill identifies them,
explains likely causes, and recommends whether to quarantine or fix each one.

**Output:** Updated `tests/regression-suite.md` quarantine section + optional
`production/qa/flakiness-report-[date].md`

**When to run:**
- Polish phase (tests have had many runs; statistical signal is reliable)
- When developers start dismissing CI failures as "probably flaky"
- After `/regression-suite` identifies quarantined tests that need diagnosis

---

## 1. Parse Arguments

**Modes:**
- `/test-flakiness [ci-log-path]` — analyse a specific CI run log file
- `/test-flakiness scan` — scan all available CI logs in `.github/` or
  standard log output directories
- `/test-flakiness registry` — read existing regression-suite.md quarantine
  section and provide remediation guidance for already-known flaky tests
- No argument — auto-detect: run `scan` if CI logs are accessible, else
  `registry`

---

## 2. Locate CI Log Data

### Option A — GitHub Actions (preferred)

Check for test result artifacts:
```bash
ls -t .github/ 2>/dev/null
ls -t test-results/ 2>/dev/null
```

For Godot projects: GdUnit4 outputs XML results compatible with JUnit format,
under `reports/` (its default report folder, `res://reports/`). Check `reports/`,
and any `test-results/` you saved runs into, for `.xml` files.

For Unity projects: game-ci test runner outputs NUnit XML to `test-results/`
by default.

For Unreal projects: automation logs go to `Saved/Logs/`. Grep for
`Result={Success}` and `Result={Fail}` — each test prints
`Test Completed. Result={<status>}`
(`docs/engine-reference/unreal/current-best-practices.md`, "Command Line").

### Option B — Local log files

If a path argument is provided, read that file directly.

### Option C — No log data available

If no logs found:
> "No CI log data found. To detect flaky tests, this skill needs test result
> history from multiple runs. Options:
> 1. Run the test suite at least 3 times and collect the output logs
> 2. Check CI pipeline output and save a log to `test-results/`
> 3. Run `/test-flakiness registry` to review tests already flagged as flaky
>    in `tests/regression-suite.md`"

Stop and ask the user which option to pursue.

---

## 3. Parse Test Results

For each CI log or result file found, parse:

**JUnit XML format** (GdUnit4):
- Grep for `<testcase name=` to get test names
- Grep for `<failure` or `<error` to identify failures
- Parse `classname` and `name` attributes for full test identifiers

**NUnit XML format** (Unity — the file whose `<test-run>` element `/smoke-check`
reads):
- Each test is a `<test-case` element; its `fullname` attribute is the identifier
- Its `result` attribute is `Passed`, `Failed`, `Inconclusive` or `Skipped`
  (`docs/engine-reference/unity/current-best-practices.md`, "Command Line");
  only `Passed` and `Failed` enter the history

**Plain text logs**:
- Grep for pass/fail patterns:
  - Godot: `PASSED` / `FAILED` adjacent to test names
  - Unreal: `Result={Success}` / `Result={Fail}`
  - Unity: `Test passed` / `Test failed`

Build a table: `test_id → [run1_result, run2_result, run3_result, ...]`

---

## 4. Identify Flaky Tests

A test is **flaky** if it appears in the result history with both PASS and
FAIL outcomes across runs with no code changes between them.

Flakiness thresholds:
- **High flakiness**: Fails in >25% of runs — quarantine immediately
- **Moderate flakiness**: Fails in 5–25% of runs — investigate and fix soon
- **Low/suspected flakiness**: Fails in 1–5% of runs — monitor; may be
  genuinely rare failure

**With fewer than 3 runs, every finding is *suspected*, whatever its fail rate.**
One failure in two runs reads as 50%, but it is one data point: do not
quarantine it, label it suspected, and ask whether more run data is available.
The tiers above, and quarantine, apply from 3 runs up.

For each flaky test, classify the likely cause:

### Cause classification

| Cause | Symptoms | Fix direction |
|-------|----------|---------------|
| **Timing / async** | Fails after awaiting signals or timers; pass rate correlates with system load | Add explicit await/synchronisation; avoid time-based delays |
| **Order dependency** | Fails when run after specific other tests; passes in isolation | Add proper setup/teardown; ensure test isolation |
| **Random seed** | Fails intermittently with no pattern; involves RNG | Pass explicit seed; don't use `randf()` in tests |
| **Resource leak** | Fails more often later in a test run | Fix cleanup in teardown; check orphan nodes (Godot) or object disposal (Unity) |
| **External state** | Fails when a file, scene, or global exists from a prior test | Isolate test from file system; use in-memory mocks |
| **Floating point** | Fails on comparisons like `== 0.5` | Use epsilon comparison (`is_equal_approx`, `Assert.AreApproximately`) |
| **Scene/prefab load race** | Fails when scenes are not yet ready | Await one frame after instant
03

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-10-06 · audit v0.4.1 · source sha 42a36917b8befull audit observations/trust-audit/skill/donchitos__test-flakiness.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-0642a36917b8beSAFEB89first audit
05

Questions

What does the Test Flakiness skill do?

Turn Claude Code into a full game dev studio — 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.

Is Test Flakiness safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Test Flakiness access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (42a36917b8be), read on 2026-10-06. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement