Atlas / Skills / microsoft / Docs Sync

Docs SyncSAFE

skills/microsoft/docs-sync

Agent Package Manager

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
—
License
MIT
Stars
3,981
01

Overview

From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.

This directory holds the eval suite for the docs-sync skill, per the genesis canonical evals doctrine (MODULE ENTRYPOINT primitive).

Files

  • trigger-evals.json -- 20 dispatch evals (10 should-trigger,

10 should-NOT-trigger), 60/40 train/val split. The validation split is the ship gate: rate >= 0.5 on should-trigger AND < 0.5 on should-not-trigger.

  • content-evals.json -- 3 content scenarios (E1 surgical CLI

fix, E2 new flag, E3 new package format) exercised withskill vs withoutskill to prove value-delta.

Ship gates

The skill is ready to graduate from rung 1 (label-gated) to rung 2 (default-on) when ALL of these pass:

  1. Trigger-eval val split: rate >= 0.5 on should-trigger AND

< 0.5 on should-not-trigger.

  1. Content evals E1, E2, E3 each produce a measurable value-delta

between with_skill and without_skill runs.

  1. Shadow-run on >= 5 recent real PRs in microsoft/apm with

no false-alarm advisories on test-only / CI-only PRs.

  1. Cost ceiling (15 LLM calls) not hit on any shadow-run case.

Notes

  • Eval execution is currently manual. Future: tie into a CI job

similar to autopilot-pr-review-worker/evals/render_eval.py.

  • The shadow-run phase is the most important. Synthetic evals

cannot fully predict classifier accuracy on real PR diffs.

Read from source at commit 2ea90c57fbfcOBSERVED · 2026-10-08
02

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: docs-sync
description: >-
  Use this skill whenever a pull request is opened, reopened, or
  synchronized in microsoft/apm to assess whether and how the
  documentation corpus must change to stay truthful with the
  proposed code change. Activate even when the PR title or body
  says nothing about docs -- the skill must run on every PR to
  detect silent drift between code and docs. Classifies impact
  as no-change, in-place edit (one to a few paragraphs), or
  structural change (new page or TOC reshape), then orchestrates
  a CDO + doc-writer + python-architect + editorial-owner +
  growth-hacker loop to produce a patch-ready advisory. Does NOT
  review code quality, security, or test coverage. Does NOT
  auto-merge or auto-push doc edits.
---

# docs-sync -- per-PR documentation impact panel

The docs corpus drifts silently and constantly. This skill catches
drift at PR-open time, classifies its impact, and orchestrates a
persona panel to produce a patch-ready advisory comment.

The pattern is **A1 PANEL + B1 FAN-OUT/SYNTHESIZER + A8 ALIGNMENT
LOOP**. The classifier is the cost gate (~70% of PRs short-circuit
to no-change with ~1 LLM call). When the panel does fan out, every
agent reads a bounded context (~10 KB) -- never the full corpus.

This skill is ADVISORY. It does not gate merge, apply verdict
labels, or push to the contributor's fork. The orchestrator is the
sole writer to the PR: exactly one comment per run (idempotent
edit-in-place), plus optional label sweeps.

## Architecture invariants

- **Cost ceiling: 15 LLM calls per run.** Hard-wired. The orchestrator refuses to spawn beyond. Header prints `N/15` for observability.
- **Single-writer interlock.** Only the orchestrator writes. Panelist subagents return JSON; they MUST NOT call any `gh` write command, post comments, or touch PR state.
- **Idempotent comment.** Exactly one comment per run, with a stable header `## Docs sync advisory`. Re-runs edit-in-place using `gh pr comment --edit-last`.
- **No fork-write.** Companion docs PRs require Step 7's fresh responsible-human issue-scope checkpoint and open in the BASE repo; never pushed to the contributor's fork. Labels request advice, not implementation.
- **Index-not-corpus reads.** Every classifier and architect agent reads `.apm/docs-index.yml`, NOT the corpus itself. The corpus is sampled only by the localizer (which reads the specific candidate pages) and by per-page panelists (which read one page each).
- **S7 deterministic tool bridge.** The python-architect panelist MUST run real `apm --help`, `grep`, and `python -c` commands to verify doc claims, never assert from prose.

## Roster

| Role | Agent | Always active? |
|---|---|---|
| Classifier | [doc-analyser](../../../.apm/agents/doc-analyser.agent.md) inside [docs-impact-classifier](../../../.apm/skills/docs-impact-classifier/SKILL.md) | Yes (every run) |
| Localizer | [docs-impact-localizer](../../../.apm/skills/docs-impact-localizer/SKILL.md) | Only on `in_place` verdict |
| Architect | [docs-impact-architect](../../../.apm/skills/docs-impact-architect/SKILL.md) | Only on `structural` verdict |
| Writer | [doc-writer](../../../.apm/agents/doc-writer.agent.md) | Per candidate page (fan-out) |
| Verifier | [python-architect](../../../.apm/agents/python-architect.agent.md) | Per candidate page (fan-out, S7) |
| Editorial | [editorial-owner](../../../.apm/agents/editorial-owner.agent.md) | Once across all redrafts |
| Growth | [oss-growth-hacker](../../../.apm/agents/oss-growth-hacker.agent.md) | Once across all redrafts |
| Synthesizer | [cdo](../../../.apm/agents/cdo.agent.md) | Once, with ALIGNMENT LOOP up to 3 redrafts |

## Topology

```
   docs-sync SKILL (orchestrator thread)
                 |
   Step 1: classify (1 LLM call, may exit here)
                 |
                 v
            verdict?
            /    |    \
   no-change  in-place  structural
       |        |          |
     EXIT       |       architect (TOC delta)
                |          |
                +----<-----+
                |
   Step 2: localize (1 LLM call) -- per-page task brief
                |
   Step 3: FAN-OUT panel via task tool
                |
       +----+----+----+----+
       v    v    v    v    v
     writer  verify edit growth
     x N    x N   once  once
       (parallel; each <=10 KB context)
                |
   Step 4: schema-validate returns
                |
   Step 5: CDO synthesize (1 LLM call)
                |
            agree?
            / | \
        revise (N<=3 redrafts) | agree
                                  |
   Step 6: emit ONE comment via safe-outputs.add-comment
   Step 7: OPTIONAL companion docs PR (structural AND fresh
           responsible-human issue-scope checkpoint)
```

## Execution checklist

### Step 1 -- Classify

Spawn ONE task: load the `docs-impact-classifier` skill, pass it the
PR number. It returns the classifier JSON.

Validate the JSON against `assets/classifier-return-schema.json`.
On schema failure, abort the run with a comment explaining the
internal error.

If verdict is `no_change`: skip to Step 6 with a brief advisory
("No docs impact detected. Reason: <one-line>. LLM calls: 1/15.")

### Step 2 -- Localize (in_place) or Architect (structural)

For `in_place`: spawn ONE task that loads the
`docs-impact-localizer` skill with the classifier output. Returns
per-page task briefs.

For `structural`: spawn ONE task that loads the
`docs-impact-architect` skill with the classifier output. Returns
TOC delta + new-page outlines + downstream in-place pages. THEN
spawn the localizer for those downstream pages.

### Step 3 -- Fan-out panel

**Cascade-size mitigation (PR 1244 class).** If `scope_pages[]` has
>8 entries, the per-page fan-out at one writer call per page would
approach the 15-call ceiling with no headroom for verifier redrafts.
BEFORE spawning, group `scope_pages[]` into SECTIONS:

- Pages under the same TOC section (e.g. all `consumer/**`) with the
  SAME con
03

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codeNA
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
none-observed
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (0)

No findings outside the package's declared scope.

Gates applied: no_behavioural_pass.

Audited 2026-10-08 · audit v0.4.1 · source sha 2ea90c57fbfcfull audit observations/trust-audit/skill/microsoft__docs-sync.json · Report an issue / request a re-scan
04

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-082ea90c57fbfcSAFEB89first audit
05

Questions

What does the Docs Sync skill do?

Agent Package Manager

Is Docs Sync safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Docs Sync access on my machine?

The audit observed no filesystem, network or shell use at all in its source.

How current is this page?

The grade is for one exact copy of the source (2ea90c57fbfc), read on 2026-10-08. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement