Docs SyncSAFE
Agent Package Manager
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
This directory holds the eval suite for the docs-sync skill, per the genesis canonical evals doctrine (MODULE ENTRYPOINT primitive).
Files
trigger-evals.json-- 20 dispatch evals (10 should-trigger,
10 should-NOT-trigger), 60/40 train/val split. The validation split is the ship gate: rate >= 0.5 on should-trigger AND < 0.5 on should-not-trigger.
content-evals.json-- 3 content scenarios (E1 surgical CLI
fix, E2 new flag, E3 new package format) exercised withskill vs withoutskill to prove value-delta.
Ship gates
The skill is ready to graduate from rung 1 (label-gated) to rung 2 (default-on) when ALL of these pass:
- Trigger-eval val split: rate >= 0.5 on should-trigger AND
< 0.5 on should-not-trigger.
- Content evals E1, E2, E3 each produce a measurable value-delta
between with_skill and without_skill runs.
- Shadow-run on >= 5 recent real PRs in microsoft/apm with
no false-alarm advisories on test-only / CI-only PRs.
- Cost ceiling (15 LLM calls) not hit on any shadow-run case.
Notes
- Eval execution is currently manual. Future: tie into a CI job
similar to autopilot-pr-review-worker/evals/render_eval.py.
- The shadow-run phase is the most important. Synthetic evals
cannot fully predict classifier accuracy on real PR diffs.
2ea90c57fbfcOBSERVED · 2026-10-08What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: docs-sync
description: >-
Use this skill whenever a pull request is opened, reopened, or
synchronized in microsoft/apm to assess whether and how the
documentation corpus must change to stay truthful with the
proposed code change. Activate even when the PR title or body
says nothing about docs -- the skill must run on every PR to
detect silent drift between code and docs. Classifies impact
as no-change, in-place edit (one to a few paragraphs), or
structural change (new page or TOC reshape), then orchestrates
a CDO + doc-writer + python-architect + editorial-owner +
growth-hacker loop to produce a patch-ready advisory. Does NOT
review code quality, security, or test coverage. Does NOT
auto-merge or auto-push doc edits.
---
# docs-sync -- per-PR documentation impact panel
The docs corpus drifts silently and constantly. This skill catches
drift at PR-open time, classifies its impact, and orchestrates a
persona panel to produce a patch-ready advisory comment.
The pattern is **A1 PANEL + B1 FAN-OUT/SYNTHESIZER + A8 ALIGNMENT
LOOP**. The classifier is the cost gate (~70% of PRs short-circuit
to no-change with ~1 LLM call). When the panel does fan out, every
agent reads a bounded context (~10 KB) -- never the full corpus.
This skill is ADVISORY. It does not gate merge, apply verdict
labels, or push to the contributor's fork. The orchestrator is the
sole writer to the PR: exactly one comment per run (idempotent
edit-in-place), plus optional label sweeps.
## Architecture invariants
- **Cost ceiling: 15 LLM calls per run.** Hard-wired. The orchestrator refuses to spawn beyond. Header prints `N/15` for observability.
- **Single-writer interlock.** Only the orchestrator writes. Panelist subagents return JSON; they MUST NOT call any `gh` write command, post comments, or touch PR state.
- **Idempotent comment.** Exactly one comment per run, with a stable header `## Docs sync advisory`. Re-runs edit-in-place using `gh pr comment --edit-last`.
- **No fork-write.** Companion docs PRs require Step 7's fresh responsible-human issue-scope checkpoint and open in the BASE repo; never pushed to the contributor's fork. Labels request advice, not implementation.
- **Index-not-corpus reads.** Every classifier and architect agent reads `.apm/docs-index.yml`, NOT the corpus itself. The corpus is sampled only by the localizer (which reads the specific candidate pages) and by per-page panelists (which read one page each).
- **S7 deterministic tool bridge.** The python-architect panelist MUST run real `apm --help`, `grep`, and `python -c` commands to verify doc claims, never assert from prose.
## Roster
| Role | Agent | Always active? |
|---|---|---|
| Classifier | [doc-analyser](../../../.apm/agents/doc-analyser.agent.md) inside [docs-impact-classifier](../../../.apm/skills/docs-impact-classifier/SKILL.md) | Yes (every run) |
| Localizer | [docs-impact-localizer](../../../.apm/skills/docs-impact-localizer/SKILL.md) | Only on `in_place` verdict |
| Architect | [docs-impact-architect](../../../.apm/skills/docs-impact-architect/SKILL.md) | Only on `structural` verdict |
| Writer | [doc-writer](../../../.apm/agents/doc-writer.agent.md) | Per candidate page (fan-out) |
| Verifier | [python-architect](../../../.apm/agents/python-architect.agent.md) | Per candidate page (fan-out, S7) |
| Editorial | [editorial-owner](../../../.apm/agents/editorial-owner.agent.md) | Once across all redrafts |
| Growth | [oss-growth-hacker](../../../.apm/agents/oss-growth-hacker.agent.md) | Once across all redrafts |
| Synthesizer | [cdo](../../../.apm/agents/cdo.agent.md) | Once, with ALIGNMENT LOOP up to 3 redrafts |
## Topology
```
docs-sync SKILL (orchestrator thread)
|
Step 1: classify (1 LLM call, may exit here)
|
v
verdict?
/ | \
no-change in-place structural
| | |
EXIT | architect (TOC delta)
| |
+----<-----+
|
Step 2: localize (1 LLM call) -- per-page task brief
|
Step 3: FAN-OUT panel via task tool
|
+----+----+----+----+
v v v v v
writer verify edit growth
x N x N once once
(parallel; each <=10 KB context)
|
Step 4: schema-validate returns
|
Step 5: CDO synthesize (1 LLM call)
|
agree?
/ | \
revise (N<=3 redrafts) | agree
|
Step 6: emit ONE comment via safe-outputs.add-comment
Step 7: OPTIONAL companion docs PR (structural AND fresh
responsible-human issue-scope checkpoint)
```
## Execution checklist
### Step 1 -- Classify
Spawn ONE task: load the `docs-impact-classifier` skill, pass it the
PR number. It returns the classifier JSON.
Validate the JSON against `assets/classifier-return-schema.json`.
On schema failure, abort the run with a comment explaining the
internal error.
If verdict is `no_change`: skip to Step 6 with a brief advisory
("No docs impact detected. Reason: <one-line>. LLM calls: 1/15.")
### Step 2 -- Localize (in_place) or Architect (structural)
For `in_place`: spawn ONE task that loads the
`docs-impact-localizer` skill with the classifier output. Returns
per-page task briefs.
For `structural`: spawn ONE task that loads the
`docs-impact-architect` skill with the classifier output. Returns
TOC delta + new-page outlines + downstream in-place pages. THEN
spawn the localizer for those downstream pages.
### Step 3 -- Fan-out panel
**Cascade-size mitigation (PR 1244 class).** If `scope_pages[]` has
>8 entries, the per-page fan-out at one writer call per page would
approach the 15-call ceiling with no headroom for verifier redrafts.
BEFORE spawning, group `scope_pages[]` into SECTIONS:
- Pages under the same TOC section (e.g. all `consumer/**`) with the
SAME conTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
2ea90c57fbfcfull audit observations/trust-audit/skill/microsoft__docs-sync.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-08 | 2ea90c57fbfc | SAFE | B | 89 | first audit |
Questions
What does the Docs Sync skill do?
Agent Package Manager
Is Docs Sync safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Docs Sync access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (2ea90c57fbfc), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.