Kiro ImplSAFE
Turn approved specs into long-running autonomous implementation. A minimal, adaptable SDD harness with Agent Skills for Claude Code, Codex, Cursor, Copilot, Windsurf, OpenCode, Gemini CLI, and Antigravity.
Overview
Turn approved specs into long-running autonomous implementation. A minimal, adaptable SDD harness with Agent Skills for Claude Code, Codex, Cursor, Copilot, Windsurf, OpenCode, Gemini CLI, and Antigravity.
4504485f1027OBSERVED · 2026-10-07What it tells the agent
The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.
---
name: kiro-impl
description: Implement approved tasks using TDD with subagent dispatch. Runs all pending tasks autonomously or selected tasks manually.
---
# kiro-impl Skill
<background_information>
You operate in two modes:
- **Autonomous mode** (no task numbers): Dispatch a fresh sub-agent per task, with independent review after each
- **Manual mode** (task numbers provided): Execute selected tasks directly in the main context
- **Success Criteria**:
- All tests written before implementation code
- Code passes all tests with no regressions
- Tasks marked as completed in tasks.md
- Implementation aligns with design and requirements
- Task completion follows the selected review mode
- **Review Mode**:
- Default is `required`
- Accept explicit forms: `--review required|inline|off`
- Also accept clear natural-language opt-outs such as `skip review` or `without review` as `off`
- If the request is ambiguous, keep `required`
</background_information>
<instructions>
## Step 1: Gather Context
Reuse steering/spec context already available from conversation; load missing context below.
Select skills for the current task even when steering/spec context is already available:
- `{{KIRO_DIR}}/specs/{feature}/spec.json` and `tasks.md` for approvals, task selection, and dependencies
- Referenced sections of `requirements.md` and `design.md`, expanding to related contracts as needed for the selected task
- Core steering context: `product.md`, `tech.md`, `structure.md`
- Additional steering files only when directly relevant to the selected task's boundary, runtime prerequisites, integrations, domain rules, security/performance constraints, or team conventions that affect implementation or validation
- Use explicitly requested skills and task-relevant local skills/playbooks, including design, accessibility, and UX. Select by description and read only needed guidance, even for small tasks; preserve required checks and host/project rules.
### Preflight
**Validate approvals**:
- Verify tasks are approved in spec.json (stop if not, see Safety & Fallback)
**Discover validation commands**:
- Inspect repository-local sources of truth in this order: project scripts/manifests (`package.json`, `pyproject.toml`, `go.mod`, `Cargo.toml`, app manifests), task runners (`Makefile`, `justfile`), CI/workflow files, existing e2e/integration configs, then `README*`
- Derive a canonical validation set for this repo: `TEST_COMMANDS`, `BUILD_COMMANDS`, and `SMOKE_COMMANDS`
- Prefer commands already used by repo automation over ad hoc shell pipelines
- For `SMOKE_COMMANDS`, choose the lightest trustworthy runtime-liveness check for the app shape (for example: root URL load, Electron launch, CLI `--help`, service health endpoint, mobile simulator/e2e harness if one already exists)
- Keep the full command set in the parent context, and pass only the task-relevant subset to implementer and reviewer sub-agents
**Establish repo baseline**:
- Run `git status --porcelain` and note any pre-existing uncommitted changes
## Step 2: Select Tasks & Determine Mode
**Parse arguments**:
- Extract feature name from `$1`
- If task numbers provided in `$2` (e.g., "1.1" or "1,2,3"): **manual mode**
- If no task numbers: **autonomous mode** (all pending tasks)
- Determine review mode from the invocation:
- `--review required` or omitted → `required`
- `--review inline` → `inline`
- `--review off`, `skip review`, or `without review` → `off`
**Build task queue**:
- Read tasks.md, identify actionable sub-tasks (X.Y numbering like 1.1, 2.3)
- Major tasks (1., 2.) are grouping headers, not execution units
- Skip tasks with `_Blocked:_` annotation
- For each selected task, check `_Depends:_` annotations -- verify referenced tasks are `[x]`
- If prerequisites incomplete, execute them first or warn the user
- Use `_Boundary:_` annotations to understand the task's component scope
## Step 3: Execute Implementation
### Autonomous Mode (sub-agent dispatch)
**Iteration discipline**: Process exactly ONE sub-task (e.g., 1.1) per iteration. Do NOT batch multiple sub-tasks into a single sub-agent dispatch. Each iteration follows the full cycle: dispatch implementer → review → commit → re-read tasks.md → next.
**Context management**: Re-read `tasks.md` before each iteration. Carry forward the task outcome, commit/evidence references, unresolved constraints, and relevant Implementation Notes; do not repeat completed worker transcripts in later handoffs.
**Delegation scope**: Use native subagents for these task-local steps, passing the focused inputs below rather than the full parent conversation. Independent feature/PR chats and their worktrees are managed through the host; do not create a standalone chat for each implementation or review step.
If multi-agent capability is available, for each task (one at a time):
**a) Dispatch implementer**:
- Read `templates/implementer-prompt.md` from this skill's directory
- Construct a prompt by combining the template with task-specific context:
- Task description and boundary scope
- Paths to spec files: requirements.md, design.md, tasks.md
- Exact requirement and design section numbers this task must satisfy (using source numbering, NOT invented `REQ-*` aliases)
- Task-relevant steering context and parent-discovered validation commands (tests/build/smoke as relevant)
- Selected skill/playbook paths and concise task-relevant guidance, including required checks; inline necessary guidance when the worker cannot access those paths
- Whether the task is behavioral (Feature Flag Protocol) or non-behavioral
- **Previous learnings**: Include any `## Implementation Notes` entries from tasks.md that are relevant to this task's boundary or dependencies (e.g., "better-sqlite3 requires separate rebuild for Electron"). This prevents the same mistakes from recurring.
- The implementer sub-agent will read the spec files and build its own Task Brief (acceptance criteria, completionTrust audit
SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.
| Layer | What it checks | Result |
|---|---|---|
| L0 | Provenance & inventory | PASS |
| L1 | Static analysis of the code | NA |
| L2 | Instruction surface (what it tells the agent) | PASS |
| L3 | Class-specific surface | PASS |
| L4 | Behavioural (sandbox) | SKIPPED |
What the source does
- Filesystem
- none-observed
- Network
- none-observed
- Shell
- none-observed
- Dependencies
- pinned
- Secrets in source
- none-found
Findings (0)
No findings outside the package's declared scope.
Gates applied: no_behavioural_pass.
4504485f1027full audit observations/trust-audit/skill/gotalab__kiro-impl.json · Report an issue / request a re-scanAudit history
Every audit this skill has had.
| Date | Source | Verdict | Grade | Score | Change |
|---|---|---|---|---|---|
| 2026-10-07 | 4504485f1027 | SAFE | B | 89 | first audit |
Questions
What does the Kiro Impl skill do?
Turn approved specs into long-running autonomous implementation. A minimal, adaptable SDD harness with Agent Skills for Claude Code, Codex, Cursor, Copilot, Windsurf, OpenCode, Gemini CLI, and Antigravity.
Is Kiro Impl safe to install?
The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.
What can Kiro Impl access on my machine?
The audit observed no filesystem, network or shell use at all in its source.
How current is this page?
The grade is for one exact copy of the source (4504485f1027), read on 2026-10-07. The repository is watched, and a new audit runs when it changes — this is the first audit.