aoti-debug
Official PyTorch guidance for systematic debugging of AOTInductor failures.
- Skill Road
- aoti-debug
Categories
aoti-debug is an official skill from the PyTorch source repository for investigating failures and crashes in AOTInductor, or AOTI. It targets developers working with ahead-of-time compiled PyTorch models who encounter unexpected behavior around aot_compile, aot_load, aoti_compile_and_package, or aoti_load_package. According to the provider, the guidance covers segmentation faults, device mismatches, constant-loading problems, incorrect outputs, CUDA failures, and Triton assertions. This catalog entry therefore presents a focused diagnostic guide for Claude Code in the PyTorch source tree, not a standalone runtime package, compiler, or automatic repair tool.
Purpose and classification
According to the provider, the first check is independent of the specific symptom: compilation device, loading device, input devices, and input shapes must agree. An AOTI artifact compiled for CUDA cannot simply be loaded on CPU, while a CPU artifact must remain on CPU. Different device indices may be possible with the current API path, but changing the device type is not supported. Inputs must also satisfy the shapes, dtypes, and strides expected at compilation time, unless explicitly modeled dynamic-shape constraints cover the runtime values. This order of operations prevents an obvious contract violation from being mistaken for a deeper compiler defect.
Diagnosing runtime failures
For input failures, the primary source recommends the AOTI_RUNTIME_CHECK_INPUTS runtime check. It can produce clearer diagnostics for mismatched device type, dtype, size, and stride. For CUDA illegal memory access, the provider suggests starting with additional checks such as TORCHINDUCTOR_NAN_ASSERTS and then using measures that make asynchronous or nondeterministic failures easier to localize. CUDA_LAUNCH_BLOCKING can synchronize kernel execution. PYTORCH_NO_CUDA_MEMORY_CACHING can remove the caching allocator from part of the failure pattern. These settings belong in a controlled development environment and must be evaluated against the selected PyTorch revision, backend, and hardware.
Kernels and intermediate values
To locate a problematic kernel, the skill points to the AOTInductor intermediate value debugger. AOT_INDUCTOR_DEBUG_INTERMEDIATE_VALUE_PRINTER can expose generated kernels step by step. A suitable filter can narrow inspection to selected kernels and their inputs. TORCH_LOGS with Inductor and output-code information, together with TORCH_SHOW_CPP_STACKTRACES, can provide additional context. According to the provider, dynamic shapes and custom C++ operators are common starting points; for custom operators, the meta function may need review for SymInt requirements. These are diagnostic paths, not promises that one environment variable will fix the underlying cause.
Triton and dynamic shapes
An included sub-guide addresses Triton assertions matching the index out of bounds pattern. It requires access to the AOTI package or extracted archive containing its wrapper file. The investigation connects the failed assertion to the generated kernel, its dynamic size parameter, the corresponding model input, and the source nodes. Empty tensors, missing lower bounds, and untested edge cases are common hypotheses. The guide recommends including empty and minimum-size shapes during export or explicitly handling empty inputs in the model path. Any change must then be confirmed with relevant PyTorch tests and technical review.
API status and boundaries
The primary source distinguishes deprecated APIs such as torch._export.aot_compile and torch._export.aot_load from the current torch._inductor.aoti_compile_and_package and torch._inductor.aoti_load_package functions. According to the provider, the newer package path stores device metadata and therefore selects the appropriate device type when loading; changing the device type remains unsupported. The skill does not replace release or backend review and guarantees neither numerical equivalence nor identical performance. It assumes a controlled checkout, reproducible inputs, knowledge of PyTorch compilation, and inspection of generated artifacts. Secrets, private models, and sensitive input data should not be placed in logs or examples.
Source, safety, and compatibility
The primary source is https://github.com/pytorch/pytorch/tree/main/.claude/skills/aoti-debug. The complementary official PyTorch documentation at https://docs.pytorch.org/docs/main/user_guide/torch_compiler/torch.compiler_aot_inductor.html provides the AOTInductor context for exported models. PyTorch publishes its source tree under BSD-3-Clause. This description was reviewed on September 9, 2026. GitHub stars are not stored because the Skill model has no github_stars field. Compatibility with Claude Code follows from publication under .claude/skills; technical results depend on the PyTorch commit, operating system, backend, and device. According to the provider, every diagnosis should use minimal reproducible inputs and proposed changes should be reviewed before adoption.
- Provider
- PyTorch
- License
- BSD-3-Clause
- Last reviewed
- 09.09.2026
Repository and documentation
Categories
Compatible with
Related guides
Guides and background related to this entry.
Set up the Fakechat plugin for Claude Code
Install the Fakechat plugin, start Claude Code with the channels flag, and test messages and files through a local browser interface.
30.09.2026
Setting up Laravel Boost
Install Laravel Boost in a Laravel application and connect it to Claude Code, Cursor, or Codex.
29.09.2026
Set up the Azure DevOps MCP Server
Start Set up the Azure DevOps MCP Server with verified links, minimal permissions, and a safe first test.
25.09.2026
Installing a Claude Code plugin
Installing a plugin from the official Anthropic marketplace – using the Code Review plugin as an example.
24.09.2026