Eval SkillsCAUTION
The most comprehensive Claude Code guide: agentic workflows, hooks, skills, MCP servers, quizzes, and production-ready templates. 430K+ lines.
Overview
The most comprehensive Claude Code guide: agentic workflows, hooks, skills, MCP servers, quizzes, and production-ready templates. 430K+ lines.
d90170da4369OBSERVED · 2026-10-07Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| claude-code | mentioned | |
| codex | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: eval-skills description: "Audit a project or library of skills for metadata, trigger boundaries, workflow completion criteria, local resource closure, instruction economy, tool scope, effort, and routing evidence. Use before shipping skills, after importing a skill collection, or when a skill triggers too often or not at all." when_to_use: "Trigger phrases: 'audit my skills', 'check skill quality', 'review skills', 'score skills', 'eval skills'. Also use when a user asks why a skill does not trigger automatically, or whether a skill can ship to Claude Code and Codex." allowed-tools: Read Glob Grep Bash(find *) Bash(claude plugin validate *) argument-hint: "[path (default: every discovered skill root)]" effort: medium --- # Skill evaluator Discover every skill in scope, classify its distribution profile, score it on eight criteria, and keep structural validity separate from routing evidence. Every score states the host and profile it applies to. ## When to use - Before committing or publishing a skill - After bulk-importing skills from another project or library - When a skill triggers too often, too rarely, or only on one host - When deciding whether a skill can ship to Claude Code, Codex, or an Agent Skills upload Read every skill in the selected scope. Do not infer collection-wide quality from a sample. When an argument names a path, audit only that path. --- ## Discovery ### Claude Code locations | Location | Path | Audit note | |---|---|---| | Personal | `~/.claude/skills/<name>/SKILL.md` | Skip `.trash/`. `synced/` holds skills downloaded from claude.ai: report them, never edit them | | Project | `.claude/skills/` in the start directory and every parent up to the repository root | Commit-shared | | Nested | `<subdir>/.claude/skills/<name>/SKILL.md` | Loads once Claude works on files in `<subdir>`. On a name clash both load and the nested one is `/<subdir>:<name>` | | Additional directory | `.claude/skills/` in a directory passed with `--add-dir` | Session-scoped | | Command file | `.claude/commands/**/*.md` | Legacy format. Same frontmatter except `name` and `paths` | | Plugin | `<plugin>/skills/<name>/SKILL.md` or a plugin-root `SKILL.md` | Namespaced as `/<plugin>:<name>` | | Enterprise | `.claude/skills/` in the managed settings directory | Report only | Reserved folder names: a folder named `synced` (any case) is skipped in the enterprise, personal, and project locations. Outside a plugin, a folder or command file named `anthropic-skills` or starting with `anthropic-skills:` does not load. ### Codex locations | Scope | Path | |---|---| | Repository | `.agents/skills` in every directory from the working directory up to the repository root | | User | `~/.agents/skills` | | Admin | `/etc/codex/skills` | | System | Bundled with Codex | Codex follows symlinked skill folders. A skill disabled with a `[[skills.config]]` entry (`enabled = false`) in `~/.codex/config.toml` is reported as disabled, not scored as broken. ```bash find .claude/skills .agents/skills -name SKILL.md -not -path '*/.trash/*' 2>/dev/null find .claude/commands -name '*.md' ! -name 'README*' 2>/dev/null ``` **Done when:** every file in scope is listed with its host, location, and the discovery rule that loads it. --- ## Distribution profiles Classify each skill before scoring it. The profile decides which fields are valid. | Profile | Where it runs | Frontmatter it may use | |---|---|---| | Claude Code native | Claude Code at any local level, including plugin skills | Every field in the Claude Code table below | | Agent Skills spec | claude.ai uploads, the Skills API, packaging with `package_skill.py` | `name`, `description`, `license`, `compatibility`, `metadata`, `allowed-tools`. Any other key fails the upload with a hard error | | Codex | Codex CLI, IDE extension, ChatGPT desktop app | `name` and `description` are required. Display metadata, invocation policy, and tool dependencies go in `agents/openai.yaml` | A skill published to several profiles follows their intersection. Keeping host-specific fields inside `metadata` to record provenance is valid for a portable skill and is not a defect. **Done when:** each skill has one declared or inferred profile, with the evidence for the inference. --- ## Claude Code frontmatter reference Claude Code reads frontmatter only when the opening `---` is the file's first line. It ignores an unknown field without reporting an error. When the YAML does not parse, the skill loads with no fields set: `/name` still works, but Claude cannot match the description. Detect parse failures with `claude plugin validate <skills-dir>` (Claude Code v2.1.233 or later). Boolean fields accept `yes`, `no`, `on`, `off`, `1`, and `0` in any case from v2.1.218. | Field | Notes | |---|---| | `name` | In a personal or project skill directory, sets the command shown in `/` and typed to invoke it; the directory name also invokes it. Defaults to the directory name. In a plugin it sets the last segment after the plugin prefix. Command files take their name from the file | | `description` | What the skill does and when to use it. If omitted, the first non-empty body line is used | | `when_to_use` | Appended to `description` in the listing and counted in the same cap | | `argument-hint` | Autocomplete hint such as `[issue-number]` | | `arguments` | Named positional arguments for `$name` substitution. Space-separated string or YAML list | | `disable-model-invocation` | `true` removes the description from context: only the user can invoke. Also blocks preloading into subagents and scheduled-task invocation | | `user-invocable` | `false` hides the skill from `/`; Claude can still invoke it | | `allowed-tools` | Pre-approves the listed tools for the turn that invokes the skill. It does not restrict tools: every tool stays callable. Space- or comma-separated string, or YAML list. Workspace trust does not gate it | | `disallowed-tools` | Removes tools from the pool while the sk
Trust audit
CAUTIONgrade B · trust 89/100 Install with care. The audit found things worth knowing before you trust its output.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | WARN |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (2)
whitepapers/recap-cards/en/_extensions
whitepapers/recap-cards/fr/_extensions
Gates applied: no_behavioural_pass.
d90170da4369full audit observations/trust-audit/skill/florianbruniaux__eval-skills.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-07 | d90170da4369 | CAUTION | B | 89 | first audit |
Questions
What does the Eval Skills skill do?
The most comprehensive Claude Code guide: agentic workflows, hooks, skills, MCP servers, quizzes, and production-ready templates. 430K+ lines.
Is Eval Skills safe to install?
With care. The audit graded it B (89/100) and found 2 things worth knowing before you trust this skill, listed below with the exact line each was found on.
What can Eval Skills access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
Which assistants does Eval Skills work with?
Its documentation mentions claude-code and codex. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (d90170da4369), read on 2026-10-07. The repository is watched, and a new audit runs when it changes — this is the first audit.