Update Golden ValuesCAUTION
Ongoing research training transformer models at scale
Overview
Ongoing research training transformer models at scale
8ca032502ee6OBSERVED · 2026-10-06What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: update-golden-values
description: Refresh golden values from a GitHub Actions workflow run (failing-only or all jobs), calculate signed per-model percentage changes, and produce a PR-ready summary. Use when the user asks to update goldens for a CI run, refresh golden values from a workflow ID, or generate a golden-value diff summary for a PR description.
---
# Update golden values + signed per-model percentage summary
End-to-end workflow for refreshing golden values from a GitHub Actions workflow run, reporting signed percentage changes per test/model environment, and writing a PR-ready summary.
The skill orchestrates two scripts that already live in the repo:
- `tests/test_utils/python_scripts/download_golden_values.py` — pulls artifacts from a workflow run and overwrites `tests/functional_tests/test_cases/**/golden_values_*.json`.
- `tests/test_utils/python_scripts/compare_golden_values_kl.py` — diffs the working-tree goldens against `git HEAD` and reports per-metric `avg_rel_diff = mean((old − new) / old)`. (Filename keeps the legacy `_kl` suffix; the script no longer computes KL divergence.)
## Inputs to gather from the user
1. **GitHub Actions workflow run ID** (e.g. `25341543542`). It's the numeric ID in the run URL.
2. **Source**: should be `github` for this workflow. (`gitlab` is supported by the download script but uses a different env path.)
3. **Scope** — accept one of:
- `only-failing` → run with `--only-failing` (download from failing/cancelled jobs only). Use this for "fix the broken tests" workflows.
- `all` → run without `--only-failing` (download from every job that produced golden values). Use this when the user wants a full refresh.
If the user doesn't specify, ask. Don't silently default.
## Workflow
```
- [ ] Step 1: Set up env (token + venv with deps)
- [ ] Step 2: Reset prior golden-value edits
- [ ] Step 3: Download goldens (scope = only-failing | all)
- [ ] Step 4: Run relative-diff comparison + generate per-model percentage table
- [ ] Step 5: Produce PR-ready summary
```
### Step 1 — Environment
The download script needs `GITHUB_TOKEN`. If the user has the `gh` CLI authenticated, derive it; do NOT export the token into a long-lived shell or commit it.
```bash
# token (one-shot, scoped to the command)
export GITHUB_TOKEN="$(gh auth token)"
# python deps (the script imports click, gitlab, requests)
python3 -m venv /tmp/gv_venv
/tmp/gv_venv/bin/pip install --quiet click python-gitlab requests
```
Reuse `/tmp/gv_venv` if it already exists. The comparison script only depends on `click` (also in the venv).
### Step 2 — Reset prior edits (only if user re-runs)
If the working tree already has prior golden-value modifications you want to discard before re-downloading:
```bash
git checkout -- tests/functional_tests/test_cases/
git ls-files --others --exclude-standard tests/functional_tests/test_cases/ \
| while IFS= read -r f; do rm -f "$f"; done
```
Skip this step when the user explicitly wants to layer a new download on top of an in-progress branch.
### Step 3 — Download
Build the command from the user-provided scope:
```bash
# scope = only-failing (default for "fix broken tests")
/tmp/gv_venv/bin/python tests/test_utils/python_scripts/download_golden_values.py \
--source github --pipeline-id <WORKFLOW_RUN_ID> --only-failing
# scope = all (full refresh; omit the flag)
/tmp/gv_venv/bin/python tests/test_utils/python_scripts/download_golden_values.py \
--source github --pipeline-id <WORKFLOW_RUN_ID>
```
When `--only-failing` is set, the GitHub path filters at `_fetch_and_filter_artifacts` on `matched_job["conclusion"] == "success"`, so only failing/cancelled jobs contribute artifacts. Without the flag, every job's golden-value artifact is pulled.
Capture the final two log lines for the summary; they look like:
```
INFO:__main__:Total tests with golden values: <N>
INFO:__main__:Total golden values found: <M>
```
### Step 4 — Relative-diff comparison
```bash
/tmp/gv_venv/bin/python tests/test_utils/python_scripts/compare_golden_values_kl.py \
--top 20 --csv /tmp/reldiff_summary.csv
```
The CSV holds one row per `(file, metric)` with four columns:
`file, metric, n_steps, avg_rel_diff`
- `n_steps` — count of shared steps that contributed (steps where `|old| < 1e-12` are skipped to avoid div-by-zero; NaN/inf are dropped).
- `avg_rel_diff` — `mean((old − new) / old)`. **Signed**: positive = the new run is smaller than the old run at the typical step (e.g. loss decreased), negative = larger.
Convert the raw ratio to a percentage for the report:
`avg_rel_diff_pct = 100 × avg_rel_diff`
Always include the `%` symbol and preserve the sign. Do not take the absolute value, produce magnitude-only statistics, or combine models into magnitude buckets. Generate one row per test/model environment instead:
```python
import collections
import csv
from pathlib import Path
rows = list(csv.DictReader(open('/tmp/reldiff_summary.csv')))
for r in rows:
r['n_steps'] = int(r['n_steps'])
r['avg_rel_diff_pct'] = 100 * float(r['avg_rel_diff'])
by_file = collections.defaultdict(dict)
for r in rows:
by_file[r['file']][r['metric']] = {
'n_steps': r['n_steps'],
'pct': r['avg_rel_diff_pct'],
}
preferred_metrics = [
'lm loss',
'mtp_1 loss',
'num-zeros',
'iteration-time',
'mem-allocated-bytes',
'mem-max-allocated-bytes',
]
present_metrics = {r['metric'] for r in rows}
metrics = [m for m in preferred_metrics if m in present_metrics]
metrics.extend(sorted(present_metrics - set(metrics)))
print('| Test / environment | Steps | ' + ' | '.join(f'`{m}` (%)' for m in metrics) + ' |')
print('| --- | --: | ' + ' | '.join('--:' for _ in metrics) + ' |')
for file_name in sorted(by_file):
path = Path(file_name)
test_name = path.parent.name
environment = path.stem.removeprefix('golden_values_')
values = by_file[file_name]
steps = max(v['n_steps'] for v in values.values())
cellsTrust audit
CAUTIONgrade B · trust 89/100 Install with care. The audit found things worth knowing before you trust its output.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | WARN |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (5)
.agents/skills
.claude/skills
CLAUDE.md
examples/post_training/modelopt/conf/nvidia/NVIDIA-Nemotron-Nano-9B-v2-Base.sh
megatron/core/models/hybrid/CLAUDE.md
Gates applied: no_behavioural_pass.
8ca032502ee6full audit observations/trust-audit/skill/nvidia__update-golden-values.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-06 | 8ca032502ee6 | CAUTION | B | 89 | first audit |
Questions
What does the Update Golden Values skill do?
Ongoing research training transformer models at scale
Is Update Golden Values safe to install?
With care. The audit graded it B (89/100) and found 5 things worth knowing before you trust this skill, listed below with the exact line each was found on.
What can Update Golden Values access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
What do I need installed to use Update Golden Values?
Its own instructions reference click. Dependencies are pinned to exact versions.
How current is this page?
The grade is for one exact copy of the source (8ca032502ee6), read on 2026-10-06. The repository is watched, and a new audit runs when it changes — this is the first audit.