A Universal Kernel
The common belief is that GPU kernels belong to one vendor and must be rewritten for the next. A decode kernel built for an AMD RX 7900 XTX (24 GB VRAM) ran on an Apple M4 GPU (16 GB unified memory) and on an NVIDIA RTX 5090 (32 GB VRAM) after one line changed, and it was the same line both times. The kernel transfers. The numbers inside it do not.
Everyone who works on GPUs learns it early. Kernels are not portable. NVIDIA has CUDA, AMD has ROCm, Apple has Metal. A fast kernel on one is a rewrite on the next. I believed it too.
Then, on 30 July 2026, a decode kernel I had built for an AMD RX 7900 XTX (24 GB VRAM) ran on an Apple M4 GPU (16 GB unified memory), in a Mac mini. Decode went from 5.39 to 17.0 tokens per second, byte-identical. The next day it ran on an NVIDIA RTX 5090 (32 GB VRAM). 4.49 to 156.2. That is 34.8 times.
Both times I changed one line. It was the same line.
The line
The kernel splits one row of a matrix multiply across a small group of GPU threads that run in lockstep. Each thread owns a run of packed 4-bit weights, unpacks them in registers, and adds up its part. So the kernel needs to know how many threads are in a group.
AMD calls the group a wave. Apple calls it a simdgroup. NVIDIA calls it a warp. On all three the width is 32.
My compiler knew that for AMD. Nobody had told it for the other two. A capability check asked for the width, got no answer, and quietly sent every 4-bit multiply down a slow generic path. On the RTX 5090 (VRAM) that path used 1.3% of what its VRAM can move. After one declared fact, 253 generated kernels installed themselves and the RTX 5090 (VRAM) ran at 44%.
The kernel was never the problem. It had been portable all along. What was missing was a fact about the chip.
A kernel is two things
The mechanism. Threads in a group own neighboring packed words. Weights are unpacked in registers and never written to memory. Partial sums are reduced at the end. It follows from the math and from how every current GPU is built.
The facts. How wide is a thread group? What shape is the matrix unit? How much bandwidth can the memory sustain? These are numbers about one chip.
A hand-written kernel fuses the two, in one vendor's language, and the result looks like it belongs to that vendor. A generated kernel keeps them apart. The mechanism is written once. The facts are declared per chip, and the code is produced from both.
| Fact | AMD RX 7900 XTX, 24 GB VRAM | Apple M4 GPU, 16 GB unified memory | NVIDIA RTX 5090, 32 GB VRAM |
|---|---|---|---|
| Thread group width | 32 | 32 | 32 |
| Matrix unit shape | 16 x 16 x 16 | 8 x 8 x 8 | 16 x 8 x 16 |
| Arithmetic peak, measured | 105 trillion per second | 3.78 trillion per second | 255.4 trillion per second |
| Share of the published peak | 86% | not published | 61% |
| Memory bandwidth | 960 GB per second | not yet measured | 1,700 GB per second, measured |
| Memory budget for the model | 24 GB | 12.7 GB | 32 GB |
| What limits one decoded token | bytes | bytes | bytes |
The last row is why transfer works. On all three, one token is limited by moving weights. The crossover depends on the chip's arithmetic rate, its bandwidth, and the bits per weight. Model size cancels out. So the same mechanism is the right one everywhere.
How much is universal
The generated 4-bit kernel matched my hand-written AMD one within 0.13% slower to 0.41% faster, across contexts from 512 to 4,096, identical tokens. The generated 6-bit kernel reproduced its hand-written twin byte for byte, worst timing ratio 1.0106. The description loses nothing against the original.
On the RTX 5090 (VRAM) one token spends 5,980 microseconds in kernels. 4,501 of them are in generated kernels worked out on the RX 7900 XTX (VRAM). That is 75% of the token, on a different vendor, language, and driver.
The method moved too. Measure four facts. Find the regime. Check the multiply reaches the matrix unit. Decide the strategy by memory arithmetic. Prove correctness. Let the search pick the geometry. The last two steps did not change at all between targets.
What stays behind
Before starting on the M4 GPU (unified memory) I wrote down, for each AMD win, what could move and what could not. The verdict: the AMD work held portable questions and candidate axes, not a recipe. That held up. Four things stayed behind.
Numbers found by search. Tiled attention starts to win at a 512-token context on the RX 7900 XTX (VRAM) and at 128 on the RTX 5090 (VRAM). A threshold is a measurement of one chip.
Hardware that is not there. On the RX 7900 XTX (VRAM), reaching the matrix unit was worth about five times on prompts. On the M4 GPU (unified memory) both paths measure the same, within 0.91 to 1.03. The biggest lever on one GPU did not exist on another.
Special instructions. The attention kernel had an AMD instruction typed in as a literal string, so the M4 GPU (unified memory) could not use that route at all. Now each compiler backend declares what it provides. The M4 GPU has no such instruction, so it gets plain arithmetic that accumulates in 32 bits, pinned by a test. CUDA has none either, and there the compiler raises an error where it could emit an unverified substitute. A wrong fast answer is worse than a refusal.
The runtime around the kernel. On the RTX 5090 (VRAM) my kernels were faster than llama.cpp's by 496 microseconds per token and I still lost by 729. llama.cpp overlapped 1,125 microseconds of work. Mine overlapped none. That belongs to how each program drives the GPU's queues, not to any kernel.
One result cuts the other way. Wider vector loads were refuted on the RX 7900 XTX (VRAM), 398 against 405 GB per second. On the RTX 5090 (VRAM) they were refuted, then promoted once a hidden copy was removed, and recovered 67 to 88 microseconds per token. The mechanism was portable. The verdict was not.
The rule
Never branch on the name of the GPU. Declare the fact the difference rests on, and derive the behavior from it.
The default runs in opposite directions for two kinds of fact. An optimization defaults to off. A chip that declares the facts may turn it on, and a chip that declares nothing loses a little speed. AMD publishes the numbers for its fast local memory, so it gets that optimization. Apple publishes none, so it goes without.
A correctness requirement defaults to on. Only a chip that declares a guarantee, and cites the hardware property it rests on, may skip it.
Both fail safe. Get the direction backwards and you ship a fast wrong answer on every GPU you have not studied.
Why one mechanism wins everywhere
These three chips solve one problem one way. Memory is slow next to arithmetic, and the chip stays busy by keeping many threads in flight while each waits. Vasily Volkov's Understanding Latency Hiding on GPUs (2016) describes this as the defining trait of the GPU, across vendors. How much work must be in flight is Little's law (Little, 1961): work in flight equals latency times throughput.
A kernel that respects that law respects something true of all three chips. The vendor's language is syntax over the top.
Where the claim stops
Three GPUs and one family of dense models. Not proof for every chip.
It is strongest for decode, where all three are limited by bytes. It is weakest for prompts on the M4 GPU (unified memory), still at 54.2 tokens per second against llama.cpp's 221.2. No transferred idea has fixed that, because the lever that fixed it on the RX 7900 XTX (VRAM) is not in the M4 GPU.
Transfer was never free. Each target's facts were measured, not read from a spec sheet. The measured arithmetic peak of the RTX 5090 (VRAM) is 61% of its published figure. Each target's geometry was searched again.
So "universal" needs care. The kernel as a description is universal, among GPUs that hide latency with wide groups of threads. The kernel as a binary is not, and neither is any number inside it.
The work left per GPU is small and of one kind. It is not writing kernels. It is telling the compiler the truth about the chip.
References
Little, J. D. C. (1961). A Proof for the Queuing Formula: L = λW. Operations Research, 9(3), 383 to 387. doi.org/10.1287/opre.9.3.383
Volkov, V. (2016). Understanding Latency Hiding on GPUs. PhD dissertation, University of California, Berkeley. Technical Report UCB/EECS-2016-143. www2.eecs.berkeley.edu/Pubs/TechRpts/2016/EECS-2016-143.pdf