BoltBeam: Opening the Black Box of Kernel Engineering
Companies are spending fortunes on tokens and cannot say what they bought, while the people selling tokens say the price does not matter. A token is GPU time turned into output by software that nobody shows. I built BoltBeam to open that software up and audit it: where every microsecond of a token goes, and why. It took 4,907 commits.
Why BoltBeam
Why are companies paying this much for AI?
Enterprises spent $37 billion on generative AI in 2025, three times the year before (Menlo Ventures, 2025). $12.5 billion of it went straight to model APIs (Menlo Ventures, 2025). That is tokens. Uber burned its whole 2026 AI budget in four months (Fortune, 2026). Asked whether any of it reached the customer, its COO, Andrew Macdonald, said, "That link is not there yet." (Fortune, 2026) Then Uber capped every employee's spend (TechCrunch, 2026).
Uber is not alone. MIT looked at 300 corporate AI projects: 95% showed no measurable return, on $30 to $40 billion (MIT NANDA, 2025). Alex Karp, the CEO of Palantir, went on CNBC to cry about stolen alpha, and said what executives say in private. Enterprises "are paying for tokens that create no value." (Yahoo Finance, 2026)
That is one side. CEOs outside AI are saying: hey, we do not know how to use this. Why are we wasting so much money?
Here is the other side. At AI networking events, people from more than one company told me the same thing to my face. CEOs do not care about tokens. The industry is printing money. AI is a goldmine.
Which is it? It cannot be both. The CEOs want real world value out of these tokens, not reckless spend. The people selling the tokens say nobody cares what a token costs. So who's the liar?
So I asked the dumb question. What is inference? What is actually being sold?
A token is GPU work. That is all it is. Somebody gets GPUs, which cost money by the hour, runs a model, and sells what comes out. The price of a token is the GPU hour divided by the tokens that hour makes. That second number is not physics. It is software. It is set by the kernels, the small programs that run the model's math on the chip.
Start with the GPU, because everything happens on it. A GPU is two parts: memory that holds the model, and thousands of small cores that do the sums. Inference is the work of turning that model into words, and on the GPU it runs as a chain of three things. A tensor is data: a table of numbers the model learned, sitting in the GPU's memory. A kernel is a small program the cores run: it reads one tensor and does the model's math on it. A token is one full pass of the model: every kernel runs once, in order, each one feeding the next, and the last one picks the word. Inference is that pass, repeated for every word of the answer. So the time of a token is the sum of its kernels, and each kernel's time starts with the bytes of its tensor.
Here is that chain once, from the GPU down to one table and back up to the price, with real numbers. One model, Qwen3-8B.
The GPU. An NVIDIA RTX 5090 with 32 GB of VRAM. It has 170 streaming multiprocessors with 21,760 lanes between them that do the sums (BoltBeam target registry), and its VRAM moves 1,700 GB per second (fork records, new target bring-up). The whole model sits in that VRAM, and most of it has to come out again for every token.
A tensor. The last step of every token scores all 151,936 words the model knows. Its weight table is 151,936 rows of 4,096 numbers, stored at 6.5625 bits each (llama.cpp source, ggml-common.h). Multiplied out, the table is 510.5 MB:
A kernel. One kernel reads that table once per token, and the VRAM hands it over at 1,700 GB per second at best. So the kernel can never finish in less than 300 microseconds:
llama.cpp's kernel takes 0.87 microseconds more than that. Mine took 9.91 more (fork records, decode ledger).
A token. One token is 253 tables like that one, seven in each of the 36 layers plus the vocabulary, with hundreds of small operations between them. Together the tables are 4.676 GB (my records, M3 roofline check, 21 September 2026). Every kernel has a floor, and the floors add up to the floor of the token, 2.75 milliseconds:
Tokens per second. Tokens per second is one over the time a token takes, so the ceiling is 364:
llama.cpp makes 247.8 with 512 tokens of conversation loaded. That is 4.04 ms a token, 68% of the ceiling. The other 32% is kernels running slower than their floors, and the time between kernels.
The price. The price of a token is the GPU hour divided by the tokens that hour makes, as at the top of this article:
Every microsecond a kernel wastes lands in that bottom number. That is what BoltBeam audits.
Speed is a bottleneck, not a business case. A slow model wastes the hour you are paying for, and it wastes yours: you wait, you lose the thread, you get less done. That is worth measuring, and it is what this article measures. But tokens per second is not value. Inference is a tool, and a tool's speed sets how fast a skilled person works; it does not supply the skill. Google's DORA study of nearly 5,000 professionals calls AI an amplifier: strong teams get stronger with it, and teams with weak processes ship low-quality work faster (DORA, 2025). METR ran a randomized trial with 16 experienced developers on their own repositories: with early-2025 AI tools they took 19% longer, while believing they had been 20% faster (METR, 2025). With an engineering background these tools multiply what you can do. Without one, fast tokens are fast noise.
So a token is a currency. Cash goes in, compute comes out. Is cash compute? And if it is, why is nobody auditing the exchange rate?
The tell is sitting in public. OpenRouter lists 16 providers for one open model, DeepSeek. Same weights. Same math. Prices differ by 4 times. Speeds run from 4 to 57 tokens per second (Coworker AI, 2026). Some of that is hardware and precision. The rest is how good they are at inference, and not one of them shows you how.
Why not? They think of it as alpha, and nobody gives out alpha for free. But kernel engineering is software. It gets better in the open.
And the secret is worth less than they think. I have used these services. OpenRouter owns no GPUs. It forwards your request and takes 5.5% (OpenRouter, 2026; Saffari, 2026). It is valued at $1.3 billion (Menlo Ventures, 2026). Modal owns no GPUs either. It rents them from the big clouds, many for less than a day, and resells an H100 for about $3.95 an hour (Modal, 2025; Modal pricing). The providers behind OpenRouter do the same one level down: rent GPUs, load an open model, sell the tokens.
I am not knocking the engineering. I am describing the business. It is token reselling, stacked in layers. Every layer adds a margin. Only the bottom layer owns the thing that is scarce.
The economics of tokens do not make sense
Here is what nobody selling tokens will tell you. At scale, owning the GPU is cheaper than buying the tokens.
An H100 costs $30,000 to $40,000 (GMI Cloud, 2026). Renting one costs about $3 an hour, which is $26,000 a year (bex.co, 2026). Rent it for a year and a bit and you have paid for the card. Run an owned H100 full and an open model costs about $0.76 per million tokens. The hosted price for the same class of model is $1.40 to $1.60 (bex.co, 2026). Half.
The catch is load. At 10% use that same GPU costs $7.58 per million tokens. The line sits near 50 million tokens a day (bex.co, 2026). Below it, buy tokens. Above it, buy GPUs. And it only works for open models, because nobody can self host a closed one.
A company spending billions of tokens is far above that line. It is cheaper for that company to buy GPUs than to keep paying for tokens. And once you own GPUs you are not only a buyer. You can sell tokens too. That is the entire business of the providers on OpenRouter. You just have to set it up. That means doing the inference yourself. Which is the one piece of software nobody will show you.
And owning compounds. A token bill is rent. You pay it, it is gone, and next month you pay it again. A GPU is an asset. Take a $35,000 card against $3 an hour of rental. After about 16 months the card is paid for, and from then on the cost is electricity, roughly $570 a year (my arithmetic: 500 watts on average at $0.13 per kWh). Over three years that is $79,000 of rent against $37,000 of ownership, and the gap widens every month after.
It compounds a second way, and this is the one I care about. When you own the compute, every improvement in the software is yours. The card I started with made 15.77 tokens per second in June and 114 in July (fork records, 11 June; fork README). I did not buy anything. If you rent tokens, a faster kernel is the provider's margin. If you own the GPU, it is your tokens. A GPU does age, and a newer one will beat it. An aging asset still beats a bill.
Even the pricing gives it away. You can charge by the token, by the task, or by the GPU hour. By the hour is what it costs. By the token is what the market picked, and it hides the one number a buyer needs: how many tokens an hour really makes.
So it is clear who is who. The companies selling tokens are not the pickaxe makers. They sell a service. They have no real alpha. The alpha is the GPUs and the chips, and the biggest winners are the people who own them. NVIDIA's data center business took in $89.0 billion in one quarter, up 117% in a year, at a 75% margin (NVIDIA, 2026). The five biggest builders plan $660 to $690 billion of infrastructure in 2026 (Futurum, 2026). Enterprises paid $12.5 billion for model APIs. Not like for like, since that spending also feeds consumer products and training. Still about fifty times.
So why is nobody doing it?
If this makes perfect sense, why are enterprises not running their own inference? Two reasons, and neither is flattering.
The first is ignorance. They do not understand it, so they pay someone who says he does. Warren Buffett told a room of business students in 1993, "Risk comes from not knowing what you're doing." (Quote Investigator, 2018) That is how inference feels to a CEO. It is not risky. It is unknown. And the unknown gets outsourced at whatever price is asked.
The second is laziness. They know they could do it. It is a hassle, so they do not. Tech has had a saying for this for fifty years: "Nobody ever got fired for buying IBM." (Forbes, 2018) Nobody gets fired for buying tokens either. If the bill is insane, blame the vendor. If you build it yourself and it breaks, that one is on you.
Ignorance and laziness both have a price, and somebody is collecting it. Jeff Bezos is reported to have told his suppliers, "Your margin is my opportunity." (Quote Investigator, 2019) Every layer in that stack lives on a margin that exists because the buyer would rather not know.
That margin is the opportunity. And the cure for ignorance is not a vendor. It is an audit.
Everything between the chip and the token is software. Software can be read, measured, and improved by one person with a GPU. So I tried.
The question nobody could answer
In June 2026 I ran one model on one GPU through two programs. llama.cpp made 101.2 tokens per second. My fork of tinygrad made 15.77 (fork records, 11 June). Same weights, same chip. The same dollar bought six times more tokens in one program than in the other.
I asked why, and I meant the specific answer. Which kernel, by how much, for what reason, and what is the most that fixing it could return. Nobody and nothing answered at that level, including the code that was fast.
Three black boxes
The vendor library. The fastest kernels ship as compiled binaries from the GPU maker. You cannot read them.
The hand-written kernel. You can read llama.cpp's main decode kernel. Reading it does not explain it. It is four or five ideas woven into one function: packed weights, which is a memory argument. Rearranged quantization math, which is algebra. A special integer instruction, which is hardware. Unrolled loops and accumulators held in registers, which is Vasily Volkov's Understanding Latency Hiding on GPUs (Volkov, 2016). Nothing says which idea buys how much (fork records, principles).
The autotuner. A compiler tries thousands of variants and keeps the fastest. It returns a winner and a number. It never returns a reason. It cannot tell you whether the thing it timed is the thing your user waits for.
What I wanted was an audit
An accountant does not accept "the business is doing better." Every dollar goes on a line, and the lines sum to the total. I wanted that for a token. Every microsecond assigned to a cause. A residual of zero.
The causes are few. A kernel moves too many bytes. It feeds the wrong arithmetic unit. It leaves the chip idle while it waits on memory. Or there is a boundary around it: a launch, a copy, a conversion.
The third one has a formula. So does the floor. Both are in the next section, and you can check them on a phone.
What BoltBeam is
BoltBeam answers four questions about an AI model. Every one of them comes down to arithmetic.
1. How fast could it ever go?
Start with the roofline. It is a model from Berkeley. It says a program is held back by one of two things: how fast the chip does math, or how fast memory feeds it (Williams, Waterman and Patterson, 2009). Which one holds back a chatbot?
Memory. To write one token, the chip reads the model's weights out of memory. Nearly all of them. Every single token. The math on each byte is tiny. So the speed limit is a division:
the floor for one step = bytes it must read / bytes per second the memory can move
Try it on a real step. The last thing a model does for every token is score every word it knows. That table has 151,936 rows and 4,096 columns (fork records, kernel census).
- 151,936 x 4,096 = 622,329,856 weights.
- Each weight is stored in 6.5625 bits (llama.cpp source, ggml-common.h). That is 510,504,960 bytes. Call it 510 MB.
- The RTX 5090's memory moves 1,700 GB per second. Measured, not read off a spec sheet (fork records, new target bring-up).
- 510 MB divided by 1,700 GB per second is 300 microseconds.
That is the floor. No kernel, no compiler, no genius gets under it on this card. The ledger's exact figure is 300.06 (fork records, decode ledger). llama.cpp runs that step 0.87 microseconds above the floor. I run it 9.91 above (fork records, decode ledger).
Now "fast" means something. It is not a benchmark score. It is your distance from the floor.
2. Is the chip being fed?
Do the same division for the whole token. My first run on the RTX 5090 made 4.49 tokens per second. That moved 21.5 GB per second, 1.3% of what the memory can do (fork records, new target bring-up). The card was not slow. It was starving.
Why would a chip starve? This is Little's law: L = λW (Little, 1961). The number of things inside a system equals the arrival rate times the time each one spends inside. A coffee shop that gets 2 customers a minute, and keeps each one 5 minutes, has 10 people in it. Always. That example is mine. The proof is his.
Memory is the slow barista. Every request waits. The only way to keep the memory busy is to have enough requests in flight to cover the wait:
requests in flight = requests per second x seconds each one waits
Too few in flight and the memory sits idle. Volkov's dissertation is about exactly this. GPUs "may execute up to about a thousand of physical threads per chip to better utilize their numerous execution units and hide execution latencies" (Volkov, 2016). He asks "how the number of threads needed to hide latency depends on basic parameters of executed code such as arithmetic intensity" (Volkov, 2016). And he found that earlier GPU performance models "mispredict observed throughputs by factors of up to 1.7" (Volkov, 2016).
Read that last one again. The published models were off by up to 1.7 times. That is why BoltBeam measures and does not guess.
My compiler had not been told that this chip runs its threads in groups of 32. One declared fact. The same model then made 156.2 tokens per second and moved 755 GB per second, 44% of the memory (fork records, new target bring-up). Check the arithmetic. 44 divided by 1.3 is 33.8. 156.2 divided by 4.49 is 34.8. The speed followed the bytes.
Check it on a second GPU
Does the division hold on a different chip? I ran it on the laptop I am writing this on. A MacBook Air with an Apple M3. Different vendor. Different memory. No fan.
I wrote the prediction down first. Apple publishes 100 GB per second for this chip (Apple, MacBook Air M3 specifications). A token reads about 4.8 GB. So the ceiling is about 20.7 tokens per second, and a good program should land under it, and not far under.
Then I measured. The memory really moves 89.9 GB per second, 90% of the published figure. llama.cpp made 17.56 tokens per second. The same day I ran the same llama.cpp test on the RTX 5090. It made 252.6 (my records, M3 roofline check, 21 September 2026).
| RTX 5090 | Apple M3 | |
|---|---|---|
| Memory rate, measured | 1,700 GB per second | 89.9 GB per second |
| Bytes read per token | 4.68 GB | 4.68 GB |
| Floor of the vocabulary step | 300 microseconds | 5,680 microseconds |
| Ceiling for a whole token | 364 tokens per second | 19.2 tokens per second |
| llama.cpp, empty context | 252.6, 69% | 17.56, 91% |
| llama.cpp, 512 tokens of context | 247.8, 68% | not run |
| My tinygrad fork, 512 tokens of context | 237.8, 65% | not run |
The memory is 18.9 times slower, and the ceiling is 18.9 times lower. Same division, different chip. Every run landed under its ceiling. None went over. And it did not matter which program ran the model: on the big card llama.cpp and my fork both sit at about two thirds of the same ceiling.
They are 4.2% apart in this quick run, llama.cpp ahead. My ledger, with a stricter protocol, has my fork ahead at that depth. One repetition of 20 tokens does not overrule it, and it does not confirm it either.
Look at the first measured row. The laptop runs closer to its ceiling than the big card does. Why? I did not measure that today, so this is a reading and not a result. A token takes 57 milliseconds on the laptop and 4 on the card. At 4 milliseconds the small costs around each kernel, a launch, a wait, stop being small. That is the kind of residual the ledger exists to find.
The full record, in plain language with a glossary, the predictions as written and every command, is its own research entry: Chips and Speed: how to serve your addicted customers w/ AI psychosis.
The run also caught a bug in my own tool. BoltBeam's quick ceiling said a token reads 3.87 GB. Adding up the parts by hand said 4.67. The quick path had dropped the layers stored in a second format, 17% of the bytes. On the RTX 5090 that put the ceiling at 439 and my fork's 156.2 at 36% of it, when the fork's own record says 44% (fork records, new target bring-up). Fixed the same day, with a test (BoltBeam commit e141a82). The arithmetic audited the auditor.
3. Did my change actually help?
This is where the two tools split the work. tinygrad is the compiler. It writes the kernels and runs the model, and I can read every kernel it writes. BoltBeam never runs the model. It keeps the books.
- BoltBeam reads the model file and lists the parts (BoltBeam, README).
- It does the division above for every part. Each part gets a floor.
- tinygrad runs the model. The GPU's profiler says what each part really took.
- Actual minus floor is the gap. Sort by gap. The biggest gap is where the money is.
- tinygrad generates a new kernel for that part. BoltBeam judges the whole token, not the part. I have shipped changes that made one part faster and the whole model slower.
The arithmetic also kills ideas before they cost anything. I thought merging the small steps would help. The ceiling said 2.8 tokens per second at most. The run returned 0.2 (BoltBeam records, 14B aggregate fusion closeout). Closed the same day.
4. Have we been here before?
Every verdict goes in a ledger with its reason. One entry says: never unpack this weight format through 16-bit numbers, it runs 16% slower (BoltBeam candidate registry, decode_q4k_gemv_f16_dequant). The next person reads that in ten seconds. They do not burn a week of GPU time to learn it again.
So what is BoltBeam? The roofline says where the floor is. tinygrad lets me read and change what runs. BoltBeam keeps the books between the two.
Why tinygrad, what BEAM is, and what machine search means
I needed a compiler I could read. tinygrad calls itself "something between PyTorch and karpathy/micrograd" (tinygrad README). A new accelerator "only needs to support a total of ~25 low level ops" (tinygrad README). Twenty-five. That is the whole reason. tinygrad lowers a model to a handful of ops and then to a kernel I can print and read. It reduces the complexity until inference is something one person can understand end to end. Compare that with the usual way to run a model. PyTorch has more than 2,000 operators, and even its own reduced set for backend writers is about 250 (PyTorch 2.x). Under them sit vendor libraries such as cuBLAS, which ship as closed binaries (PyTorch 2.x). tinygrad does not rely on PyTorch or on those libraries. It generates the kernel itself. In my fork the NVIDIA route never calls the CUDA launch API at all. It drives the hardware through the driver directly (fork records, five lever test).
I did not pick it because it was fast. It was six times slower. I picked it because I could see inside it.
BEAM is tinygrad's search. Set BEAM to a number and you get that "number of beams in kernel beam search" (tinygrad docs). For each kernel it tries variants, times them on the GPU, keeps the best few, and tries again. It works. It is also the autotuner from the last section: a winner, a number, no reason. And it times one kernel alone, which is the exact test that shipped slowdowns for me.
Marlin is the other end. It is a hand-built 4-bit kernel. It holds close to the full 4 times speedup from 4-bit weights up to batch sizes of 16 to 32, and gives up to 2.8 times end to end inside vLLM (Frantar et al., 2024). It gets there with asynchronous memory access, task scheduling, pipelining, and quantization support built for the job (Frantar et al., 2024). That is the Volkov playbook, done by experts, for one vendor.
Here is the biggest issue, and neither side names it. The search space of kernel engineering is enormous, and no GPU can time it. BEAM's own list is small: 334 actions for one kernel, the upcasts, unrolls, locals, group reduces, tensor cores and axis swaps (tinygrad source, search.py). It applies them one after another and keeps the best few. After steps, the number of sequences it could have tried is 334 to the power :
That is 111,556 at two steps and about 12 billion at four. BEAM does not try them. It keeps a beam, times each candidate on the GPU, and stops when a step gains less than 0.01 microseconds (tinygrad source, search.py). That is a walk through the space with a stopwatch, not a search of it.
And it is the space of one fixed list. Kernel engineering has more knobs than the list. One flash attention tile alone has six, and multiplying their choices gives 432 tile geometries:
On the RTX 5090 (VRAM) every one of the 432 was legal and matched the production tile, and timing them all took one capture (fork records, flash geometry search). Now add the weight format, the data type, the layout, and whether two kernels fuse. Then multiply across the kernels of one token, because the verdict is the whole token and never the kernel. The space of one token is the product of its kernels' spaces:
Nobody times that. Marlin's answer is to not search: one expert picks one point in the space, for one chip. BEAM's answer is to walk a short list with a stopwatch. Mine is to cut the space before the stopwatch runs: the chip's facts say which geometries are legal, the roofline says which are worth timing, and the ledger says which were already refuted.
So BEAM searches without understanding, and Marlin understands without searching. I wanted both, with a reason attached. BoltBeam keeps Marlin as a yardstick. When a 4-bit kernel turns out to be bound by its scale metadata, the fix it names is a Marlin-style layout (BoltBeam source, quant_gemv.py).
That is what I mean by machine search, and the rule in my fork is strict. The kernel that actually runs must be generated from a description, selected by the search, and verified by the compiler's gates (fork records, pure machine search). A hand-written kernel may stay around as a rollback, or as a ceiling to compare against. It may not be the default. Every route carries a provenance label, and only two labels are allowed to ship: generated by the search, or generated by tinygrad's own scheduler (fork records, pure machine search). The label goes on what executes, "not the benchmark story around it" (fork records, pure machine search).
Why so strict? A hand-written kernel is a recipe. It works on the chip it was written for. A generated kernel is a description plus the chip's facts, which is why mine moved to two new GPUs. And the rule cost nothing. The generated 4-bit kernel tracked my hand-written one within 0.13% slower to 0.41% faster, with identical tokens (fork records, transfer matrix).
How it works
BoltBeam never runs a model. A compiler does that. BoltBeam reads, bounds, judges, and records.
Read. It reads the model file, not the model's name: tensors, shapes, weight formats, roles. A model it cannot classify gets no search space. It refuses where it could guess.
Bound. Before anyone tunes anything, it derives the ceiling for each role and says which limit binds it. That is question 1, done for every part of the model and written down first.
Judge. A verdict is one of eight words: diagnostic, candidate, promote, refute, defer, inconclusive, search-space-incomplete, adapter-incomplete. To earn promote the output must stay correct, the intended route must have run, the whole model must be faster, memory must fit, and one switch must turn it off.
Record. Every verdict goes in a ledger that is only added to. A refuted idea leaves the search, so nobody pays for it twice. Each refutation keeps its reason. The reason behind the 16-bit entry above: the GPU converts through 32 bits first, so the conversions double.
What is legal lives in three data files. A new GPU is a data edit, not a code branch. A guard rejects any rule written for a specific model.
What I use it for
BoltBeam is one of three tools, and they go together.
BubbleBeam proposes. It reads the facts a chip declares, such as the width of a thread group, the size of its fast local memory, and the shape of its matrix unit. From those it lists the legal values for every setting. It runs on a CPU. It never touches a GPU.
FutureSight predicts. It takes the candidates and rejects or orders them before anything runs. This one does not fit in local memory. This one's reduction does not cover its threads. This one goes first. Still no GPU.
BoltBeam decides. It alone owns the candidates (fork source). It spends GPU time only on what survived, judges the whole token, and records the verdict.
Each one replaces a step of BEAM. Here is what BEAM does, from its source (tinygrad source, search.py), and what I put in its place.
| The step | tinygrad's BEAM | Mine |
|---|---|---|
| What to try | One fixed list of actions, the same for every chip, cut off by hard caps | BubbleBeam: legal values worked out from the facts the chip declares |
| What to skip | Nothing. Every candidate that compiles gets timed on the GPU | FutureSight: rejects and orders candidates before anything runs |
| How to judge | The fastest of 3 runs, for one kernel alone | BoltBeam: the whole token, one of eight verdicts |
| What is kept | The winning options, in a cache. No reason | BoltBeam: a ledger entry, with the reason, for wins and losses |
Propose, predict, measure. GPU time is the expensive part, so the two cheap tools run first and the expensive one runs last.
I use the three for the same jobs every time. Bringing up a new GPU: the Apple M4 and the RTX 5090 each started as a list of declared facts. Moving to a bigger model, where the winning route for 8B is the wrong route for 14B. Picking geometry, the tile sizes and split counts that nobody can derive by hand. And closing a question for good, so the next model inherits the answer and not the argument.
The names are Pokémon moves, a pun on BEAM, tinygrad's kernel search. Just for fun.
The audit that closed to zero
On an RTX 5090 the fork decoded at 208.8 tokens per second and llama.cpp at 246.4. That is 729 microseconds per token, and the audit assigned all of it.
My kernels were faster, by 496 microseconds. I lost anyway. llama.cpp overlapped 1,125 microseconds of work per token and my route overlapped none (fork records, wall account). Little's law, one level up: they kept more work in flight.
No kernel benchmark finds that. No autotuner reports it. It is a plain sentence, it is checkable, and it says what to build next.
Two weeks later the same ledger read the other way. On 1 September the fork decoded at 247.80 tokens per second against llama.cpp's 246.41, which is 0.56% ahead (fork records, decode ledger). This is the account, region by region, in microseconds per token.
| Region of one token. NVIDIA RTX 5090, my tinygrad fork against llama.cpp, 512 tokens of context, 1 September 2026 | Microseconds ahead of llama.cpp |
|---|---|
| norms | +92.9 |
| FFN down | +41.2 |
| gate and up | +26.2 |
| Q, K, V | +16.4 |
| attention score | +4.6 |
| O projection | -3.0 |
| vocabulary | -8.4 |
| net, in the kernels | +169.9 |
Five regions won and two lost. Then the second measure, the one a benchmark never gives you: the roofline. The floor for the vocabulary projection is 300.06 microseconds, because the bytes cannot move faster than that. llama.cpp sits 0.87 above the floor. The fork sits 9.91 above it, which is 3.3% (fork records, decode ledger). llama.cpp is on the floor there. That is the row still to win, and the ledger says so by name.
The ledger also states its own limit. This is one protocol at a 512-token context. A fresh sweep did not repeat the lead at short context. At 4,096 tokens the fork was 4.56% ahead (fork records, depth sweep). An audit that hides its weak rows is not an audit.
Where it stops
BoltBeam measures nothing itself. It needs a compiler that exposes its settings. The bounded space is a bet: if the best route is outside the listed families, it finds the best of what it was told is legal. Many ledger entries are one comparison on one machine. The accounting is for dense transformers. And it costs more than a recipe. It pays with the second model.
Why an auditor built this
I did not come to this from GPUs. I came from audit. I spent more than eight years at KPMG and PwC, working where audit meets tech. I did completeness testing and revenue recognition for tech companies. Then I was a solution architect implementing Workiva, where the whole job is knowing how a number flows end to end, from the source system to the line in the filing.
It is the same job.
Completeness testing asks whether every transaction made it into the ledger. The census asks whether every microsecond of a token made it into the account.
Revenue recognition says revenue counts when it is earned, not when it is claimed. The promote gate says a speedup counts when the token gets faster, not when a benchmark says so.
A reconciliation takes two records of the same thing and explains the difference until nothing is left. That is the RTX 5090 account: 729 microseconds, every one assigned, residual zero.
The roofline is the control total. It tells you what the number should be before you look. Then you trace, from the model file through every kernel to the wall clock, the way you trace a figure from the general ledger to the 10-K.
And every audit needs an independent source to tie out to. Mine was llama.cpp. I used it as a reference. When my findings said a kernel was slow, llama.cpp told me whether that was true of the hardware or only of my code.
The industry is spending hundreds of billions of dollars on a number nobody has audited. Cash goes in, tokens come out, and the rate between them is set by software nobody will show you. I know how to audit a number. So I did.
References
Inline citations link to the source of each claim. Claims about my own work cite the project's records. The tinygrad fork is public and its records are linked. BoltBeam has a public mirror, BoltBeam-Public, MIT licensed, holding the tool, its tests, its docs and the measurements cited here; development happens in a private repository and lands there in snapshots. What people told me at events, and my own career, are my account.
Frantar, E., Castro, R. L., Chen, J., Hoefler, T., and Alistarh, D. (2024). MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models. arXiv:2408.11743
tinygrad. README and environment variables. github.com/tinygrad/tinygrad, docs.tinygrad.org/env_vars
tinygrad-arkey, the public fork this work was done in. github.com/JulianAbeleda/tinygrad-arkey
Fortune (26 May 2026). Uber burned through its entire 2026 AI budget in four months. Now its COO is questioning whether it's worth it. fortune.com
Yahoo Finance (2 July 2026). Palantir CEO Alex Karp claims AI companies are stealing customers' data while charging them for unproductive tokens. Reporting on his CNBC Squawk Box interview. finance.yahoo.com
Coworker AI (17 July 2026, updated 14 September 2026). OpenRouter Pricing. Source of the provider count, price spread and throughput range for DeepSeek on OpenRouter. coworker.ai
Menlo Ventures (9 December 2025). 2025: The State of Generative AI in the Enterprise. Enterprise spend of $37 billion, of which $12.5 billion on foundation model APIs. menlovc.com
MIT Project NANDA (July 2025). The GenAI Divide: State of AI in Business 2025. Reported in Virtualization Review.
Menlo Ventures (26 May 2026). OpenRouter Now Processes More Than a Quadrillion Tokens a Year. menlovc.com
Futurum Group (12 February 2026). AI Capex 2026: The $690B Infrastructure Sprint. futurumgroup.com
Quote Investigator (18 March 2018). Risk Comes from Not Knowing What You're Doing. Traces the line to Warren Buffett at Columbia Business School in 1993, reported by the Omaha World-Herald in January 1994. quoteinvestigator.com
Quote Investigator (13 January 2019). Your Margin Is My Opportunity. Traces the line to Fortune's 2012 profile of Jeff Bezos by Adam Lashinsky. quoteinvestigator.com
bex.co (8 July 2026). The Self-Hosted GPU Breakeven Point. Source of the owned and hosted cost per million tokens and the break-even volume. bex.co
GMI Cloud (14 April 2026). NVIDIA H100 GPU Pricing: 2026 Rent vs. Buy Cost Analysis. gmicloud.ai
NVIDIA (26 August 2026). NVIDIA Announces Financial Results for Second Quarter Fiscal 2027. nvidianews.nvidia.com
OpenRouter. Frequently asked questions: pricing and provider routing. openrouter.ai/docs/faq
OpenRouter (2026). How OpenRouter model routing works: providers, fallbacks and the auto router. openrouter.ai/blog/insights/model-routing
Modal. Plan pricing, per-second GPU rates. modal.com/pricing
Modal. Keeping 20,000 GPUs healthy. On running across rented capacity from several clouds. modal.com/blog/gpu-health
Williams, S., Waterman, A., and Patterson, D. (2009). Roofline: An Insightful Visual Performance Model for Multicore Architectures. Communications of the ACM, 52(4), 65 to 76. doi.org/10.1145/1498765.1498785
llama.cpp. ggml-common.h, the Q6_K block: "Effectively 6.5625 bits per weight". github.com/ggml-org/llama.cpp
Apple. MacBook Air (13-inch, M3, 2024), technical specifications: "100GB/s memory bandwidth". support.apple.com/en-us/118551
Little, J. D. C. (1961). A Proof for the Queuing Formula: L = λW. Operations Research, 9(3), 383 to 387. doi.org/10.1287/opre.9.3.383
Volkov, V. (2016). Understanding Latency Hiding on GPUs. PhD dissertation, University of California, Berkeley. Technical Report UCB/EECS-2016-143. www2.eecs.berkeley.edu/Pubs/TechRpts/2016/EECS-2016-143.pdf