Warden SweepSAFE
A Model Context Protocol (MCP) server and CLI that provides tools for agent use when working on iOS and macOS projects.
Overview
A Model Context Protocol (MCP) server and CLI that provides tools for agent use when working on iOS and macOS projects.
a8e6e2fdb267OBSERVED · 2026-10-07Install
Commands as the repository documents them. They are shown, not run.
uv run ${CLAUDE_SKILL_ROOT}/scripts/scan.py [file ...]uv run ${CLAUDE_SKILL_ROOT}/scripts/index_prs.py <sweep-dir>uv run ${CLAUDE_SKILL_ROOT}/scripts/create_issue.py <sweep-dir>uv run ${CLAUDE_SKILL_ROOT}/scripts/organize.py <sweep-dir>uv run ${CLAUDE_SKILL_ROOT}/scripts/extract_findings.py <log-path-or-directory> -o <output.jsonl>uv run ${CLAUDE_SKILL_ROOT}/scripts/generate_report.py <sweep-dir>What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: warden-sweep
description: Full-repository code sweep. Scans every file with warden, verifies findings via deep tracing, creates draft PRs for validated issues. Use when asked to "sweep the repo", "scan everything", "find all bugs", "full codebase review", "batch code analysis", or run warden across the entire repository.
disable-model-invocation: true
---
# Warden Sweep
Full-repository code sweep: scan every file, verify findings with deep tracing, create draft PRs for validated issues.
**Requires**: `warden`, `gh`, `git`, `jq`, `uv`
**Important**: Run all scripts from the repository root using `${CLAUDE_SKILL_ROOT}`. Output goes to `.warden/sweeps/<run-id>/`.
## Bundled Scripts
### `scripts/scan.py`
Runs setup and scan in one call: generates run ID, creates sweep dir, checks deps, creates `warden` label, enumerates files, runs warden per file, extracts findings.
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/scan.py [file ...]
--sweep-dir DIR # Resume into existing sweep dir
```
### `scripts/index_prs.py`
Fetches open warden-labeled PRs, builds file-to-PR dedup index, caches diffs for overlapping PRs.
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/index_prs.py <sweep-dir>
```
### `scripts/create_issue.py`
Creates a GitHub tracking issue summarizing sweep results. Run after verification, before patching.
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/create_issue.py <sweep-dir>
```
### `scripts/organize.py`
Tags security findings, labels security PRs, updates finding reports with PR links, posts final results to tracking issue, generates summary report, finalizes manifest.
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/organize.py <sweep-dir>
```
### `scripts/extract_findings.py`
Parses warden JSONL log files and extracts normalized findings. Called automatically by `scan.py`.
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/extract_findings.py <log-path-or-directory> -o <output.jsonl>
```
### `scripts/generate_report.py`
Builds `summary.md` and `report.json` from sweep data. Called automatically by `organize.py`.
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/generate_report.py <sweep-dir>
```
### `scripts/find_reviewers.py`
Finds top 2 git contributors for a file (last 12 months).
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/find_reviewers.py <file-path>
```
Returns JSON: `{"reviewers": ["user1", "user2"]}`
---
## Phase 1: Scan
**Run** (1 tool call):
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/scan.py
```
To resume a partial scan:
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/scan.py --sweep-dir .warden/sweeps/<run-id>
```
Parse the JSON stdout. Save `runId` and `sweepDir` for subsequent phases.
**Report** to user:
```
## Scan Complete
Scanned **{filesScanned}** files, **{filesTimedOut}** timed out, **{filesErrored}** errors.
### Findings ({totalFindings} total)
| # | Severity | Skill | File | Title |
|---|----------|-------|------|-------|
| 1 | **HIGH** | security-review | `src/db/query.ts:42` | SQL injection in query builder |
...
```
Render every finding from the `findings` array. Bold severity for high and above.
**On failure**: If exit code 1, show the error JSON and stop. If exit code 2, show the partial results. List timed-out files separately from errored files so users know which can be retried.
---
## Phase 2: Verify
Deep-trace each finding using Task subagents to qualify or disqualify.
**For each finding in `data/all-findings.jsonl`:**
Check if `data/verify/<finding-id>.json` already exists (incrementality). If it does, skip.
Launch a Task subagent (`subagent_type: "general-purpose"`) for each finding. Process findings in parallel batches of up to 8 to improve throughput.
**Task prompt for each finding:**
Read `${CLAUDE_SKILL_ROOT}/references/verify-prompt.md` for the prompt template. Substitute the finding's values into the `${...}` placeholders.
**Process results:**
Parse the JSON from the subagent response and:
- Write result to `data/verify/<finding-id>.json`
- Append to `data/verified.jsonl` or `data/rejected.jsonl`
- For verified findings, generate `findings/<finding-id>.md`:
```markdown
# ${TITLE}
**ID**: ${FINDING_ID} | **Severity**: ${SEVERITY} | **Confidence**: ${CONFIDENCE}
**Skill**: ${SKILL} | **File**: ${FILE_PATH}:${START_LINE}
## Description
${DESCRIPTION}
## Verification
**Verdict**: Verified (${VERIFICATION_CONFIDENCE})
**Reasoning**: ${REASONING}
**Code trace**: ${TRACE_NOTES}
## Suggested Fix
${FIX_DESCRIPTION}
```diff
${FIX_DIFF}
```
```
Update manifest: set `phases.verify` to `"complete"`.
**Report** to user after all verifications:
```
## Verification Complete
**{verified}** verified, **{rejected}** rejected.
### Verified Findings
| # | Severity | Confidence | File | Title | Reasoning |
|---|----------|------------|------|-------|-----------|
| 1 | **HIGH** | high | `src/db/query.ts:42` | SQL injection in query builder | User input flows directly into... |
...
### Rejected ({rejected_count})
- `{findingId}` {file}: {reasoning}
...
```
---
## Phase 3: Issue
Create a tracking issue that ties all PRs together and gives reviewers a single overview.
**Run** (1 tool call):
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/create_issue.py ${SWEEP_DIR}
```
Parse the JSON stdout. Save `issueUrl` and `issueNumber` for Phase 4.
**Report** to user:
```
## Tracking Issue Created
{issueUrl}
```
**On failure**: Show the error. Continue to Phase 4 (PRs can still be created without a tracking issue).
---
## Phase 4: Patch
For each verified finding, create a worktree, fix the code, and open a draft PR. Process findings **sequentially** (one at a time) since parallel subagents cross-contaminate worktrees.
**Severity triage**: Patch HIGH and above. For MEDIUM, only patch findings from bug-detection skills (e.g., `code-review`, `security-review`). Skip LOW and INFO findings.
**Step 0: Setup** (run once before the loop):
```bash
uv run ${CLAUDE_SKILL_ROOT}/scripts/index_prs.py ${SWEEP_DIR}
```
Parse the JSTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | PASS |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
a8e6e2fdb267full audit observations/trust-audit/skill/getsentry__warden-sweep.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-07 | a8e6e2fdb267 | SAFE | B | 89 | first audit |
Questions
What does the Warden Sweep skill do?
A Model Context Protocol (MCP) server and CLI that provides tools for agent use when working on iOS and macOS projects.
Is Warden Sweep safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Warden Sweep access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
What do I need installed to use Warden Sweep?
Its own instructions reference gh, git, jq, uv and warden. Dependencies are pinned to exact versions.
How current is this page?
The grade is for one exact copy of the source (a8e6e2fdb267), read on 2026-10-07. The repository is watched, and a new audit runs when it changes — this is the first audit.