Atlas / Skills / galaxy-dawn / Results Analysis

Results AnalysisCAUTION

skills/galaxy-dawn/results-analysis

Semi-automated research assistant for academic research and software development. Supports Claude Code, Codex CLI, Kimi Code CLI, and OpenCode across ideation, coding, experiments, writing, and publication.

Verdict
CAUTION
Grade
B
Trust score
89 /100
Version
0.2.0
Hosts
—
License
MIT
Stars
5,693
01

Overview

Semi-automated research assistant for academic research and software development. Supports Claude Code, Codex CLI, Kimi Code CLI, and OpenCode across ideation, coding, experiments, writing, and publication.

Read from source at commit 29ad4d4206fbOBSERVED · 2026-10-07
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: results-analysis
description: This skill should be used when the user asks to "analyze experimental results", "run strict statistical analysis", "compare model performance", "generate scientific figures", "check significance", "do ablation analysis", or mentions interpreting experiment data with rigorous statistics and visualization. It focuses on strict analysis bundles, not Results-section prose.
tags: [Research, Analysis, Statistics, Visualization, Scientific Reporting]
version: 0.2.0
---

# Results Analysis

Run **strict, evidence-first experimental analysis** for ML/AI research.

Use this skill to produce a **strict analysis bundle**:
- `analysis-report.md`
- `stats-appendix.md`
- `figure-catalog.md`
- `figures/`

When the user asks for review, audit, no-write, dry-run, or when inputs are incomplete, use **read-only audit mode** instead of producing files or figures. In that mode, output only valid/invalid statistics, blockers, claim candidates, and what evidence is missing. If invoked by `/analyze-results`, the command layer may write a blocker summary, but this skill should not create figures, reports, or polished conclusions from incomplete evidence.

Do **not** use this skill to draft a paper `Results` section or a full experiment wrap-up report. Those belong to `ml-paper-writing` or `results-report`.

## Core contract

### This skill is responsible for
- validating experiment artifacts and comparison units,
- running rigorous descriptive and inferential statistics,
- generating **real scientific figures** when data/logs are available,
- writing figure purposes, caption requirements, and interpretation checklists,
- surfacing limits, blockers, and missing evidence explicitly.

### This skill is not responsible for
- paper-ready `Results` prose,
- manuscript narrative polishing,
- paper-ready figure/table packaging with `pubfig` / `pubtab`,
- project-level experiment retrospectives.

If the user wants the complete post-experiment summary report, hand off to `results-report` after this bundle is ready. If the user wants publication-grade figures/tables, export parameters, publication QA, or figure/table redesign, hand off to `publication-chart-skill`.

## Non-negotiable quality bar

1. **Prefer real figures over figure specs.**
   If the data can be read, generate real figures. Do not stop at “recommended visualization”.
   Exception: in read-only audit mode, do not generate figures; describe what figure would be valid after evidence is complete.
2. **Never fabricate statistics.**
   If sample size, seeds, or raw metrics are missing, state the blocker clearly.
3. **Report complete statistics.**
   Do not report only best scores or only p-values.
4. **Interpret every main figure.**
   Every major figure must have purpose, caption requirements, and post-figure interpretation notes.
5. **Separate evidence from prose.**
   This skill produces analysis artifacts; it does not write manuscript sections.

## Standard workflow

### 1. Inventory and validate artifacts

Start by identifying:
- metric tables (`csv`, `json`, `tsv`, logs),
- training curves and checkpoints,
- seeds / repeated runs,
- baselines, ablations, and comparison families,
- evaluation protocol metadata.

Validate:
- metric direction (higher/lower is better),
- unit of analysis (run, subject, fold, dataset, seed),
- number of runs / seeds,
- missing values or silent failures,
- comparability across methods.

If the comparison is not statistically valid, say so before continuing. Do not treat repeated `subject × task` rows, folds, windows, trials, or seeds as independent units unless the design justifies it.
Common blocker: a `subject × task` summary table is usually a repeated-measure summary, not an independent subject-level sample. If subjects have multiple task rows or missing task cells, state that before any significance or winner claim.

### 2. Lock the comparison questions

Before running statistics, define the exact comparison questions:
- Which method is compared to which baseline?
- What is the primary metric?
- What is the repeated-measure unit?
- Which ablation or robustness questions matter?
- Which findings are decision-changing?

Do not mix unrelated comparisons into one undifferentiated table.

### 3. Run strict statistics

Always produce:
- descriptive statistics: `mean ± std` when appropriate,
- `95% CI` or another clearly justified interval,
- run/seed counts,
- significance tests with assumptions stated,
- effect sizes,
- multiple-comparison handling when several contrasts are reported.

Default expectation:
- check parametric assumptions first,
- use non-parametric fallback when assumptions fail,
- state exactly what was tested and on what samples.

See:
- `references/statistical-methods.md`
- `references/statistical-reporting.md`

### 4. Generate real scientific figures

Produce actual figures whenever artifacts are available.

Minimum expectation for a non-trivial analysis bundle:
- **one main comparison figure**,
- **one supporting figure** (training dynamics / ablation / breakdown / error analysis),
- **one exact numeric summary table** in markdown.

Every main figure must define:
- figure purpose,
- plotted variables,
- error bar meaning,
- caption requirements,
- interpretation checklist.

See:
- `references/visualization-best-practices.md`
- `references/figure-interpretation.md`

### 5. Write analysis artifacts

#### `analysis-report.md`
Summarize:
- the analysis question,
- key findings,
- strongest supported comparisons,
- main caveats,
- what changed in the experimental understanding,
- claim candidates that may later be used in reports or manuscript writing.

Each claim candidate should use this shape:

```md
## Claim Candidates

- Claim:
  - Source evidence:
  - Allowed wording:
  - Forbidden stronger wording:
  - Uncertainty:
  - Next check:
  - Decision: keep | weaken | revise | discard
```

#### `stats-appendix.md`
Record:
- descriptive statistics,
- test choices,
- assumptio
03

Trust audit

CAUTIONgrade B · trust 89/100 Install with care. The audit found things worth knowing before you trust its output.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)WARN
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (1)

MEDIUMPrompt injection · prompt.conditional_escalation · CWE-94, CWE-1427
SKILL.md:18
When the user asks for review, audit, no-write, dry-run, or when inputs are incomplete, use **read-only audit mode** instead of producing files or figures. In that mode, output only valid/invalid stat

Gates applied: no_behavioural_pass.

Audited 2026-10-07 · audit v0.4.1 · source sha 29ad4d4206fbfull audit observations/trust-audit/skill/galaxy-dawn__results-analysis.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-0729ad4d4206fbCAUTIONB89first audit
05

Questions

What does the Results Analysis skill do?

Semi-automated research assistant for academic research and software development. Supports Claude Code, Codex CLI, Kimi Code CLI, and OpenCode across ideation, coding, experiments, writing, and publication.

Is Results Analysis safe to install?

With care. The audit graded it B (89/100) and found 1 thing worth knowing before you trust this skill, listed below with the exact line each was found on.

What can Results Analysis access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (29ad4d4206fb), read on 2026-10-07. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement