Atlas / Skills / openai / Code Change Verification

Code Change VerificationSAFE

skills/openai/code-change-verification

A lightweight, powerful framework for multi-agent workflows

Verdict
SAFE
Grade
B
Trust score
89 /100
Version
—
Hosts
1 documented
License
MIT
Stars
29,851
01

Overview

A lightweight, powerful framework for multi-agent workflows

Read from source at commit 090ff821bbcdOBSERVED · 2026-10-06
02

Host compatibility

What the documentation claims. We have not run a compatibility test.

HostStatusNotes
codexmentioned
03

What it tells the agent

The instruction file, verbatim from the audited commit — this is the text the model reads, and the surface the audit's instruction layer examines. Quoted here so you can judge it without cloning anything.

---
name: code-change-verification
description: Run the required final formatting, lint, type, and test checks after eligible SDK changes pass review.
---

# Code Change Verification

## Overview

Ensure work is only marked complete after formatting, linting, type checking, and tests pass. Use this skill when changes affect runtime code, tests, or build/test configuration. You can skip it for docs-only or repository metadata unless a user asks for the full stack. This is a post-review final gate: when `$implementation-final-review` applies, do not invoke the broad stack until its clean-review condition applies to the stable task diff.

## Quick start

1. Keep this skill at `./.agents/skills/code-change-verification` so it loads automatically for the repository.
2. Codex on macOS/Linux: `/usr/bin/env -u OPENAI_API_KEY OPENAI_AGENTS_TEST_IN_CODEX_SANDBOX=1 UV_DEFAULT_INDEX=https://pypi.org/simple bash .agents/skills/code-change-verification/scripts/run.sh`.
3. Other macOS/Linux environments: `env UV_DEFAULT_INDEX=https://pypi.org/simple bash .agents/skills/code-change-verification/scripts/run.sh`.
4. Windows: `powershell -ExecutionPolicy Bypass -File .agents/skills/code-change-verification/scripts/run.ps1`.
5. On macOS/Linux, the script runs `make format`, `make lint`, `make typecheck`, and `make tests` sequentially and stops at the first failure. Parallelism inside each Make target, including pytest workers, is unchanged.
6. The Bash script streams each command's output directly. The Windows wrapper retains parallel lint, typecheck, and test steps with periodic heartbeat updates.
7. If any command fails, fix the issue, rerun the script, and report the failing output.
8. Confirm completion only when all commands succeed with no remaining issues.

## Start condition and host capacity

- During iterative review, use only focused tests and a narrowly targeted static check when the changed typing boundary requires one. Defer repository-wide `make typecheck` and the rest of this complete stack until review is clean.
- Immediately before starting the complete stack, use available read-only task or process evidence to check whether another repository-wide test, typecheck, build, examples runner, or integration command is already active on the same host.
- When concrete contention is visible, continue useful non-heavy work such as review, remediation, evidence preparation, or focused checks, then check again later. Do not create or wait on a repository lock, host-wide mutex, or sentinel file.
- Start automatically once review is clean, the diff is stable, and observable host capacity is available. Do not require a user-triggered `finalize` message. If host telemetry is unavailable, do not block solely because capacity cannot be measured.

## Codex execution policy

Repository verification and all child processes must remain in the normal Codex workspace sandbox. Never request elevated sandbox permissions for the verification wrapper, and never retry the wrapper with broader host access after a failure.

On macOS, tests marked `requires_native_macos_sandbox` need to start their own `sandbox-exec` process. The Codex command sets `OPENAI_AGENTS_TEST_IN_CODEX_SANDBOX=1`, which skips only that marker before nested sandbox creation. All other tests remain enabled. Ordinary local and CI runs do not set this variable and therefore keep the marked tests enabled.

The marked tests run separately on a disposable GitHub-hosted macOS runner. If that trusted runner is unavailable, report the missing native-macOS coverage; do not compensate by weakening the Codex sandbox boundary.

## Environment setup

The verification scripts assume repository dependencies are already installed. Do not run `make sync` as part of every verification pass; use it for a fresh checkout, after dependency files change, or when dependency resolution fails before the checks start.

On Linux, some Python packages with native extensions may require system packages such as `libffi-dev`, Python development headers, or build tools. If verification cannot start because one of these packages is missing, treat it as a local environment setup issue. Install the missing dependency when possible, or report the failing command and missing dependency in the PR test plan before rerunning verification in a prepared environment.

## Manual workflow

- For a fresh checkout, or if dependencies are not installed or have changed, run `make sync` first to install dev requirements via `uv`.
- Run from the repository root with `make format` first, then `make lint`, `make typecheck`, and `make tests`.
- Do not skip steps; stop and fix issues immediately when a command fails.
- Run the manual steps sequentially and stop at the first failure. Keep the parallelism provided by each Make target.
- Re-run the full stack after applying fixes so the commands execute in the required order.

## Resources

### scripts/run.sh

- Runs `make format`, `make lint`, `make typecheck`, and `make tests` sequentially from the repository root. It streams output, preserves the first failure or cancellation status, and cleans up the active step's process group before continuing or exiting.

### scripts/run.ps1

- Windows-friendly wrapper that runs the same sequence with `make format` first and the remaining steps in parallel with fail-fast semantics, plus periodic heartbeat updates while work is still running. Use from PowerShell with execution policy bypass if required by your environment.
04

Trust audit

SAFEgrade B · trust 89/100 Nothing in the source contradicts what it says it does. Grade A is reserved for packages that have also passed the behavioural sandbox.

LayerWhat it checksResult
L0Provenance & inventoryPASS
L1Static analysis of the codePASS
L2Instruction surface (what it tells the agent)PASS
L3Class-specific surfacePASS
L4Behavioural (sandbox)SKIPPED

What the source does

Filesystem
declared (1 observation(s))
Network
none-observed
Shell
none-observed
Dependencies
pinned
Secrets in source
none-found

Findings (2)

LOWInventory / provenance · inv.symlink · CWE-1104
CLAUDE.md
CLAUDE.md
Why it matters. link not followed
LOWFilesystem / path · fs.traversal · CWE-22, CWE-59
scripts/run.sh:10
REPO_ROOT="${REPO_ROOT:-$(cd "${SCRIPT_DIR}/../../../.." && pwd)}"

Gates applied: no_behavioural_pass.

Audited 2026-10-06 · audit v0.4.1 · source sha 090ff821bbcdfull audit observations/trust-audit/skill/openai__code-change-verification.json · Report an issue / request a re-scan
05

Audit history

Every audit this skill has had.

DateSourceVerdictGradeScoreChange
2026-10-06090ff821bbcdSAFEB89first audit
06

Questions

What does the Code Change Verification skill do?

A lightweight, powerful framework for multi-agent workflows

Is Code Change Verification safe to install?

The audit found nothing in the source that contradicts what it says it does, and graded it B (89/100). Grade A is held back for packages that have also passed a sandboxed behavioural run, which is why a clean skill reads B.

What can Code Change Verification access on my machine?

The audit observed that it reads or writes files. Each of those is consistent with what it says it does. Secrets in the source: none found.

What do I need installed to use Code Change Verification?

Its own instructions reference finalize and sandbox-exec. Dependencies are pinned to exact versions.

Which assistants does Code Change Verification work with?

Its documentation mentions codex. That is what the text claims, not a compatibility test we ran.

How current is this page?

The grade is for one exact copy of the source (090ff821bbcd), read on 2026-10-06. The repository is watched, and a new audit runs when it changes — this is the first audit.

Advertisement