BenchmarkSAFE
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Overview
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
e9ca581a4f44OBSERVED · 2026-09-19What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: benchmark description: このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します。 origin: ECC --- # ベンチマーク — パフォーマンスベースラインと回帰検出 ## 使用時期 - PR前後にパフォーマンスへの影響を測定 - プロジェクトのパフォーマンスベースラインを設定 - ユーザーが「遅く感じる」と報告したとき - ローンチ前 — パフォーマンスターゲットを満たしていることを確認 - スタックを代替案と比較 ## 動作方法 ### モード1:ページパフォーマンス ブラウザMCPを介してリアルブラウザメトリクスを測定: ``` 1. 各ターゲットURLに移動 2. Core Web Vitalsを測定: - LCP (Largest Contentful Paint) — ターゲット < 2.5s - CLS (Cumulative Layout Shift) — ターゲット < 0.1 - INP (Interaction to Next Paint) — ターゲット < 200ms - FCP (First Contentful Paint) — ターゲット < 1.8s - TTFB (Time to First Byte) — ターゲット < 800ms 3. リソースサイズを測定: - 合計ページウェイト(ターゲット < 1MB) - JSバンドルサイズ(ターゲット < 200KBgzipped) - CSSサイズ - 画像ウェイト - サードパーティスクリプトウェイト 4. ネットワークリクエストをカウント 5. レンダリングブロッキングリソースをチェック ``` ### モード2:APIパフォーマンス APIエンドポイントをベンチマーク: ``` 1. 各エンドポイントに100回ヒット 2. 測定:p50、p95、p99レイテンシ 3. トラック:レスポンスサイズ、ステータスコード 4. ロード下でテスト:10個の同時リクエスト 5. SLAターゲットと比較 ``` ### モード3:ビルドパフォーマンス 開発フィードバックループを測定: ``` 1. コールドビルド時間 2. ホットリロード時間(HMR) 3. テストスイート期間 4. TypeScriptチェック時間 5. Lint時間 6. Dockerビルド時間 ``` ### モード4:前後の比較 変更前後に実行して影響を測定: ``` /benchmark baseline # 現在のメトリクスを保存 # ... 変更を加える ... /benchmark compare # ベースラインと比較 ``` 出力: ``` | Metric | Before | After | Delta | Verdict | |--------|--------|-------|-------|---------| | LCP | 1.2s | 1.4s | +200ms | WARNING: WARN | | Bundle | 180KB | 175KB | -5KB | ✓ BETTER | | Build | 12s | 14s | +2s | WARNING: WARN | ``` ## 出力 `.ecc/benchmarks/`にJSONとしてベースラインを保存。Gitで追跡されるため、チームはベースラインを共有します。 ## 統合 - CI:すべてのPRで`/benchmark compare`を実行 - `/canary-watch`とペアリングしてデプロイ後の監視 - `/browser-qa`とペアリングして完全な出荷前チェックリスト
Trust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
e9ca581a4f44full audit observations/trust-audit/skill/affaan-m__benchmark.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-09-19 | e9ca581a4f44 | SAFE | B | 89 | first audit |
Questions
What does the Benchmark skill do?
The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Is Benchmark safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Benchmark access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (e9ca581a4f44), read on 2026-09-19. The repository is watched, and a new audit runs when it changes — this is the first audit.