Use Code Change Verification safely

OpenAI's Code Change Verification skill runs formatting, lint, type checks, and tests as a final gate after code review passes.

Published on 09.09.2026

What the Code Change Verification skill is

The Code Change Verification skill comes from OpenAI's official openai-agents-python repository, where it lives in the agent skills directory used by the OpenAI Agents SDK for Python. According to the provider, the skill ensures that work is only marked complete once formatting, linting, type checking, and tests have all passed. It applies to changes that affect runtime code, tests, or build/test configuration, but can be skipped for docs-only or repository-metadata-only changes unless the full check stack is explicitly requested. The skill explicitly presents itself as a final gate that runs after review has already happened, and it only becomes active once the preceding review step has concluded on a stable diff.

The verification flow in detail

According to the documentation, the actual check is executed via a startup script, with different invocations depending on the operating system: on macOS/Linux there is a specific invocation for use inside a Codex sandbox environment and a more general invocation for other environments, while Windows uses a PowerShell script. On macOS and Linux, the script runs the four steps make format, make lint, make typecheck, and make tests sequentially and stops at the first failure, while parallelism within individual steps, such as multiple test workers, remains unchanged. The Bash script streams each command's output directly and unaltered, while the Windows wrapper runs the lint, typecheck, and test steps in parallel with periodic heartbeat updates. If a command fails, the underlying issue must be fixed and the script rerun; only once every command succeeds with no remaining issues is the change considered fully verified.

When the full check kicks in

A notable feature of the skill is its tiered approach to resource usage. According to the documentation, during an ongoing, iterative review only focused individual tests and a narrowly targeted static check for the specific typing boundary affected should be used; the complete, repository-wide check run with make typecheck and the remaining steps is deliberately deferred until review is clean. Immediately before starting the complete check chain, available read-only process or task indicators should be used to check whether another repository-wide test, typecheck, build, or integration run is already active on the same host. If such concrete resource contention is detected, other, less resource-intensive work should continue instead, with a later re-check, without creating a repository lock, a host-wide mutex, or a sentinel file.

Prerequisites and setup

For the skill to load automatically, according to the documentation it must be placed in the .agents/skills/code-change-verification directory within the respective repository. It also requires a working Makefile with the format, lint, typecheck, and tests targets, since the script wrappers ultimately just call these Make targets in sequence. For use inside a Codex sandbox on macOS or Linux, a specific environment-variable invocation is also intended, which among other things sets an alternative Python package index to keep network access within the sandbox under control.

Security, limitations, and best practices

According to the documentation, verification runs and any child processes they start must remain strictly within the normal Codex sandbox; the verification wrapper must not request elevated sandbox permissions, nor should it retry with elevated privileges, which prevents an automated verification run from unintentionally exceeding its intended access rights. The limitation of the skill is that it does not itself perform any qualitative assessment of the code; it functions purely as execution and stop logic around already-existing tools such as formatters, linters, type checkers, and test runners. If one of those underlying tools is poorly configured, that gap carries through unchanged into the verification. For projects without a unified Makefile providing the four named targets, the skill is not directly usable in its current form and would first need to be adapted to fit the project's own build toolchain.

Published on 09.09.2026

Categories

Frequently asked questions

When is the complete stack appropriate?

After a clean implementation review and once the diff is stable. Focused checks are better for intermediate states.

Does a green run prove complete security?

No. It provides integration evidence but does not replace security analysis or product acceptance.