skill-upCAUTION
An evaluation and evolution tool for Agent Skills.
Overview
From the repository's own README, as read at the audited commit. Badges and raw HTML are left out.
skill-up
Evaluate Agent Skills, agents, and workspaces. Evolve Skills with evidence.
English | 中文
📖 User Manual · 用户手册
Overview
skill-up is a CLI for evaluating Agent Skills, agents, and their workspaces. It runs declarative cases against an Agent Engine, grades the response and workspace changes, and
13ba68522229OBSERVED · 2026-09-24Install
Commands as the repository documents them. They are shown, not run.
npx skills add https://github.com/alibaba/skill-up/tree/main/skills/skill-upper -g -a codex -y
npx skills add https://github.com/alibaba/skill-up/tree/main/skills/skill-upper -g -a claude-code -y
npx skills add https://github.com/alibaba/skill-up/tree/main/skills/skill-upper -g -a codex -y
npx skills add https://github.com/alibaba/skill-up/tree/main/skills/skill-upper -g -a claude-code -y
npx skills add https://github.com/alibaba/skill-up/tree/main/skills/skill-upper -g -a codex -y
npx skills add https://github.com/alibaba/skill-up/tree/main/skills/skill-upper -g -a claude-code -y
Host compatibility
What the documentation claims. We have not run a compatibility test.
| Host | Status | Notes |
|---|---|---|
| claude-code | mentioned | |
| codex | mentioned | |
| cursor | mentioned |
What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
--- name: code-stats description: Analyzes code files and reports statistics including line counts, file counts by extension, and total size. Use this skill whenever the user wants to understand the composition of a codebase - asking about "how many lines of code", "what file types exist", "code distribution", or needing a quick audit of project size and structure. Make sure to invoke this skill when users mention analyzing codebases, counting lines, checking file distributions, or auditing code. version: 1.0.0 --- # Code Stats Skill Analyzes code files and generates statistics about a codebase. ## Usage When invoked, this skill will: 1. Scan the specified directory (defaults to current directory) 2. Count total lines, files by extension 3. Report the largest files 4. Output results in a structured format 5. Sort extension summaries by total line count in descending order ## Output Format Always use this exact format for the output: ``` # Code Statistics ## Summary - Total Files: <count> - Total Lines: <count> - Total Size: <size_bytes> bytes ## Files by Extension | Extension | Files | Lines | |-----------|-------|-------| | .go | 12 | 1500 | | .md | 5 | 200 | ## Top 5 File Extensions by Line Count 1. .go — 1500 lines 2. .md — 200 lines ## Largest Files (top 5) 1. main.go (500 lines) 2. util.go (300 lines) ``` If a file has no extension, group it under `(no ext)` in both extension sections. ## Process 1. Use `Bash` with `find` to locate all code files 2. Use `Bash` with `wc -l` to count lines per file 3. Group files with no extension under `(no ext)` 4. Sort both extension summaries by total line count in descending order 5. Aggregate the results into the output format above 6. If no directory is specified, analyze the current working directory ## Examples - "Analyze the current directory" - "Run code-stats on ./src" - "Show me file type distribution" - "How many lines of Python do we have?"
Trust audit
CAUTIONgrade F · trust 58/100 Install with care. The audit found things worth knowing before you trust its output.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | WARN |
| L2 | Instruction surface (what it tells the agent) | WARN |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- declared (3 observation(s))
- Network
- declared (5 observation(s))
- Shell
- none-observed
- Dependencies
- not all pinned
- Secrets in source
- found
Findings (25)
Args: []string{"--token", "sk-ant-api03-AAAAAAAAAAAAAAAAAAAA"},APIKey: "dashscope-test-token",
APIKey: "sk-ant-should-not-appear",
const secret = "sk-super-secret-token-value"
const secret = "env-super-secret-token-value"
a := &CustomAgent{BaseAgent: BaseAgent{Cfg: Config{Name: "x", APIKey: "sk-real-secret-token"}}} //nolint:gosec // fake credential fixturecurl -fsSL https://raw.githubusercontent.com/alibaba/skill-up/main/install.sh | bash
.golangci.yml
.goreleaser.yaml
CLAUDE.md
func (r *probeMergeTestRuntime) Exec(_ context.Context, cmd string, _ ExecOptions) (ExecResult, error) {func (r *claudeCodeTestRuntime) Exec(_ context.Context, command string, opts runtime.ExecOptions) (runtime.ExecResult, error) {func (r *codexTestRuntime) Exec(_ context.Context, command string, opts runtime.ExecOptions) (runtime.ExecResult, error) {func (f fakeFindRuntime) Exec(_ context.Context, _ string, _ ExecOptions) (ExecResult, error) {func (r *nodeBootstrapTestRuntime) Exec(_ context.Context, command string, _ runtime.ExecOptions) (runtime.ExecResult, error) {OutputFile: "../../etc/important",
Kwargs: map[string]string{"target": "../../etc/important"},const testEvalPath = "../../examples/code-stats/evals/eval.yaml"
{name: "nested parent traversal", path: "fixtures/../../secret.txt"},{Path: "../../etc/passwd", Content: "root"},if curl -fsS http://127.0.0.1:8080/health 2>/dev/null | grep -q healthy; then
{ "name": "remote-report.html", "url": "http://127.0.0.1:8080/artifacts/report.html", "content_type": "text/html" },custom := httpEngine("http://127.0.0.1:0")custom := httpEngine("http://127.0.0.1:0")custom := httpEngine("http://127.0.0.1:0")Gates applied: no_behavioural_pass.
13ba68522229full audit observations/trust-audit/skill/alibaba__skill-up.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-09-24 | 13ba68522229 | CAUTION | F | 58 | first audit |
Questions
What does the skill-up skill do?
An evaluation and evolution tool for Agent Skills.
Is skill-up safe to install?
With care. The audit graded it F (58/100) and found 25 things worth knowing before you trust this skill, listed below with the exact line each was found on.
What can skill-up access on my machine?
The audit observed that it reaches the network and reads or writes files. Each of those is consistent with what it says it does. Secrets in the source: found — see the findings.
Which assistants does skill-up work with?
Its documentation mentions claude-code, codex and cursor. That is what the text claims, not a compatibility test we ran.
How current is this page?
The grade is for one exact copy of the source (13ba68522229), read on 2026-09-24. The repository is watched, and a new audit runs when it changes — this is the first audit.