aoti-debug

Official PyTorch guidance for systematic debugging of AOTInductor failures.

aoti-debug is an official skill from the PyTorch source repository for investigating failures and crashes in AOTInductor, or AOTI. It targets developers working with ahead-of-time compiled PyTorch models who encounter unexpected behavior around aot_compile, aot_load, aoti_compile_and_package, or aoti_load_package. According to the provider, the guidance covers segmentation faults, device mismatches, constant-loading problems, incorrect outputs, CUDA failures, and Triton assertions. This catalog entry therefore presents a focused diagnostic guide for Claude Code in the PyTorch source tree, not a standalone runtime package, compiler, or automatic repair tool.

Purpose and classification

According to the provider, the first check is independent of the specific symptom: compilation device, loading device, input devices, and input shapes must agree. An AOTI artifact compiled for CUDA cannot simply be loaded on CPU, while a CPU artifact must remain on CPU. Different device indices may be possible with the current API path, but changing the device type is not supported. Inputs must also satisfy the shapes, dtypes, and strides expected at compilation time, unless explicitly modeled dynamic-shape constraints cover the runtime values. This order of operations prevents an obvious contract violation from being mistaken for a deeper compiler defect.

Diagnosing runtime failures

For input failures, the primary source recommends the AOTI_RUNTIME_CHECK_INPUTS runtime check. It can produce clearer diagnostics for mismatched device type, dtype, size, and stride. For CUDA illegal memory access, the provider suggests starting with additional checks such as TORCHINDUCTOR_NAN_ASSERTS and then using measures that make asynchronous or nondeterministic failures easier to localize. CUDA_LAUNCH_BLOCKING can synchronize kernel execution. PYTORCH_NO_CUDA_MEMORY_CACHING can remove the caching allocator from part of the failure pattern. These settings belong in a controlled development environment and must be evaluated against the selected PyTorch revision, backend, and hardware.

Kernels and intermediate values

To locate a problematic kernel, the skill points to the AOTInductor intermediate value debugger. AOT_INDUCTOR_DEBUG_INTERMEDIATE_VALUE_PRINTER can expose generated kernels step by step. A suitable filter can narrow inspection to selected kernels and their inputs. TORCH_LOGS with Inductor and output-code information, together with TORCH_SHOW_CPP_STACKTRACES, can provide additional context. According to the provider, dynamic shapes and custom C++ operators are common starting points; for custom operators, the meta function may need review for SymInt requirements. These are diagnostic paths, not promises that one environment variable will fix the underlying cause.

Triton and dynamic shapes

An included sub-guide addresses Triton assertions matching the index out of bounds pattern. It requires access to the AOTI package or extracted archive containing its wrapper file. The investigation connects the failed assertion to the generated kernel, its dynamic size parameter, the corresponding model input, and the source nodes. Empty tensors, missing lower bounds, and untested edge cases are common hypotheses. The guide recommends including empty and minimum-size shapes during export or explicitly handling empty inputs in the model path. Any change must then be confirmed with relevant PyTorch tests and technical review.

API status and boundaries

The primary source distinguishes deprecated APIs such as torch._export.aot_compile and torch._export.aot_load from the current torch._inductor.aoti_compile_and_package and torch._inductor.aoti_load_package functions. According to the provider, the newer package path stores device metadata and therefore selects the appropriate device type when loading; changing the device type remains unsupported. The skill does not replace release or backend review and guarantees neither numerical equivalence nor identical performance. It assumes a controlled checkout, reproducible inputs, knowledge of PyTorch compilation, and inspection of generated artifacts. Secrets, private models, and sensitive input data should not be placed in logs or examples.

Source, safety, and compatibility

The primary source is https://github.com/pytorch/pytorch/tree/main/.claude/skills/aoti-debug. The complementary official PyTorch documentation at https://docs.pytorch.org/docs/main/user_guide/torch_compiler/torch.compiler_aot_inductor.html provides the AOTInductor context for exported models. PyTorch publishes its source tree under BSD-3-Clause. This description was reviewed on September 9, 2026. GitHub stars are not stored because the Skill model has no github_stars field. Compatibility with Claude Code follows from publication under .claude/skills; technical results depend on the PyTorch commit, operating system, backend, and device. According to the provider, every diagnosis should use minimal reproducible inputs and proposed changes should be reviewed before adoption.

Free
Provider
PyTorch
License
BSD-3-Clause
Last reviewed
09.09.2026

Repository and documentation

Categories

Compatible with

Claude Code