metal-kernel
Official PyTorch guidance for native Metal kernels and MPS operators on Apple Silicon.
- Skill Road
- metal-kernel
Categories
metal-kernel is an official skill in the PyTorch repository. It is aimed at developers who need to add PyTorch operators to Apple's MPS backend, migrate existing MPSGraph implementations to native Metal kernels, or port CUDA-oriented logic to Apple Silicon. According to the provider, the intended approach is to use PyTorch's native Metal infrastructure under c10/metal rather than MPSGraph. This makes the skill a technical contribution guide for the PyTorch source tree, not a GPU driver, not a standalone Metal framework, and not a general introduction to machine learning.
Purpose and architecture
The workflow described by the provider connects three layers. First, native_functions.yaml must contain the correct dispatch mapping for the operator. During a migration, every relevant overload needs review, including functional, in-place, out, and Tensor or Scalar variants. The Metal kernel then belongs under aten/src/ATen/native/mps/kernels. For elementwise operations, the source describes functors and registration macros for unary, binary, and alpha-parameter cases. The third layer is the host-side stub under aten/src/ATen/native/mps/operations. It binds TensorIterator to the matching kernel library and registers MPS dispatch. This separation helps an engineer identify whether a failure belongs to YAML dispatch, Metal code, or host integration.
Practical implementation
According to the provider, REGISTER_UNARY_OP and REGISTER_BINARY_OP hide iterator plumbing for common elementwise operations. Depending on semantics, floating-point, integer, BFloat16, and complex data types may require separate registrations or functor overloads. The source explains operation-math and accumulation types, as well as precise mathematical helpers from c10/metal. Complex values deserve special attention because their Metal representations can be vector types. When migrating an existing MPSGraph operation, the old host implementation must be removed and every reference to the legacy function should be checked. Leaving one overload on the old path can make callers silently use MPSGraph even when another overload uses native Metal.
Testing and diagnosis
PyTorch's test infrastructure in test/test_mps.py checks basic output agreement. The skill also points to torch.mps.compile_shader, which can compile individual Metal shaders just in time and test them in isolation on an MPS device. The official documentation distinguishes threads as the total number of threads from group_size as the number of threads per threadgroup. This matters because dispatchThreads has different semantics from dispatchThreadgroups. For a multi-kernel pipeline, each intermediate result should be compared with a CPU or NumPy reference. Copying every GPU value to the CPU merely to perform an error check can introduce an avoidable synchronization cost.
Boundaries and safety
The skill assumes familiarity with C++, Objective-C or Objective-C++, Metal, PyTorch builds, and MPS. The referenced files are in the PyTorch source tree; they do not automatically change an installed PyTorch release. Apple hardware, a compatible operating system, compilers, SDKs, and a correctly configured build are required. The guide does not guarantee performance or numerical equality for every shape and data type. Empty and non-contiguous tensors, large tensors, index bounds, and asynchronous failures require particular care. According to the provider, errors should be reported through the supported MPS error buffer instead of forcing a GPU-to-CPU synchronization. Builds and tests belong in a controlled development workspace; unknown files, scripts, or generated changes must be reviewed before execution.
Source and classification
The primary source is https://github.com/pytorch/pytorch/tree/main/.claude/skills/metal-kernel. The complementary official PyTorch documentation at https://docs.pytorch.org/docs/stable/generated/torch.mps.compile_shader.html covers isolated shader testing. The PyTorch source tree is licensed under BSD-3-Clause. This catalog description was reviewed on September 9, 2026. GitHub stars are not stored for this entry because the Skill model has no github_stars field. The skill is compatible with Claude Code because it is published as a .claude/skills resource; the technical guidance still requires validation against the PyTorch version installed in the developer's environment.
- Provider
- PyTorch
- License
- BSD-3-Clause
- Last reviewed
- 09.09.2026
Repository and documentation
Categories
Compatible with
Related guides
Guides and background related to this entry.
Set up the Fakechat plugin for Claude Code
Install the Fakechat plugin, start Claude Code with the channels flag, and test messages and files through a local browser interface.
30.09.2026
Setting up Laravel Boost
Install Laravel Boost in a Laravel application and connect it to Claude Code, Cursor, or Codex.
29.09.2026
Set up the Azure DevOps MCP Server
Start Set up the Azure DevOps MCP Server with verified links, minimal permissions, and a safe first test.
25.09.2026
Installing a Claude Code plugin
Installing a plugin from the official Anthropic marketplace – using the Code Review plugin as an example.
24.09.2026