Set up the PyTorch AOTI debugging skill
aoti-debug is an official PyTorch skill that systematically diagnoses AOTInductor errors such as segfaults, device mismatches, and loading failures.
- Skill Road
- Set up the PyTorch AOTI debugging skill
Published on 09.09.2026
What aoti-debug is and why it matters
aoti-debug is a Claude Code skill that lives directly inside the official PyTorch GitHub repository, at the path .claude/skills/aoti-debug/SKILL.md within pytorch/pytorch. It targets developers working with AOTInductor, commonly abbreviated AOTI. According to the maintainers, AOTInductor is PyTorch 2's ahead-of-time compilation component, which translates trained models into an executable form in advance rather than recompiling them on every program start. That speeds up the startup of inference workloads considerably, but it also introduces a class of errors that looks quite different from ordinary Python tracebacks: segmentation faults, silent crashes while loading constants, or confusing messages about mismatched devices. This is exactly where the skill steps in. It acts as a structured debugging guide that Claude Code pulls in automatically whenever the conversation mentions functions such as aot_compile, aot_load, aoti_compile_and_package, or aoti_load_package, or when a user describes an AOTI crash. For teams shipping PyTorch models to production through AOTI, a skill like this meaningfully shortens time-to-root-cause because it condenses the accumulated experience of PyTorch maintainers into a repeatable checklist instead of forcing every engineer to rediscover the same failure patterns from scratch.
Prerequisites
To use the skill effectively you need a local PyTorch installation with AOTInductor enabled, and ideally access to a GPU, since many of the described failure modes, such as illegal memory access errors or device mismatch problems, are CUDA-specific. The skill itself ships as part of the PyTorch repository and is loaded through the Claude Code skill mechanism, for instance via npx skills add pytorch/pytorch with the flag for the desired skill, or by manually copying the SKILL.md file into a project's local .claude/skills directory. Users should also know which device type their model was compiled on, since that information is a prerequisite for almost every debugging step described. Basic familiarity with the PyTorch compilation pipeline, such as what guards, kernels, or Torch Inductor codegen mean, makes the suggested steps easier to follow, but it is not strictly required because the skill explains the terminology in context as it goes.
Step-by-step setup
First, the skill file is placed in the project's or user profile's skill directory, after which Claude Code recognizes it automatically whenever a matching request comes in. The skill's own first content step is to check the error message and route based on pattern matching: if it recognizes a message of the form Assertion index out of bounds, it points to a specialized sub-guide for Triton index errors. For every other error type, the skill follows a fixed sequence that, according to the vendor, always checks first whether the compile device matches the load device, whether input devices match the model's device, and whether input shapes match the shapes used during compilation. These three checks are central because AOTInductor, per its documentation, does not support cross-device loading: a model compiled on GPU cannot be loaded on CPU and vice versa, while only the device index within the same device type may vary. Knowing this baseline rule lets you narrow down most segfaults and runtime errors in the very first debugging step, before ever having to dig into kernel logs.
Security and best practices
Because this is a pure debugging aid that does not modify production data, the security risk profile is fairly low, but a deliberate approach to the suggested environment variables still pays off. Flags such as AOTI_RUNTIME_CHECK_INPUTS or TORCHINDUCTOR_NAN_ASSERTS enable, according to the vendor, additional runtime checks that slow down execution, which is why they should only be set temporarily during troubleshooting rather than left on permanently in production. CUDA_LAUNCH_BLOCKING likewise forces synchronous instead of asynchronous kernel execution and should be turned off again once diagnosis is complete, to avoid unnecessary performance loss. It is also worth noting that the skill explicitly distinguishes between a deprecated and a current API: torch._export.aot_compile and aot_load are marked deprecated, while aoti_compile_and_package and aoti_load_package are presented as the recommended, current interface that automatically stores device metadata inside the package. Projects still relying on the old API should plan a migration in the medium term, since future PyTorch releases may eventually remove the older functions.
Practical example and limits
A typical scenario: a developer compiles a model on a CUDA GPU but accidentally loads it in a CPU-only container, which triggers a cryptic error about a pointer that is not registered with any device. In this case the skill would immediately point to the device-matching rule and hand over the fix, instead of hours spent digging through kernel traces. For nondeterministic CUDA illegal memory access errors, the skill proposes a multi-step approach: first sanity checks via compile-time flags, then deterministic reproduction using PYTORCH_NO_CUDA_MEMORY_CACHING and CUDA_LAUNCH_BLOCKING, and finally precise kernel identification through what it calls the Intermediate Value Debugger. The skill's limits show up where errors originate outside the AOTI stack itself, for example in custom C++ code for user-defined operators with dynamic shapes; here the skill only notes that the meta function may need to be made SymInt-aware, without offering an automated fix. Anyone working intensively with AOTInductor should therefore treat the skill as a structured starting point that covers many standard cases well, while more complex custom-op problems will still require independent deep-dive investigation.
Frequently asked questions
Can a CUDA-compiled package be loaded on CPU?
No. Compilation and loading must use the same device type; a different device index may be possible depending on the API.
Does this entry store GitHub stars?
No. The Skill model has no github_stars field, so that metric is neither invented nor stored.