arkey researchRSS
All posts / Journal
note2026-10-07inference

How do I make money in AI?
The answer is hard…ware

I spent four months learning the inference stack on three GPUs and asked the auditor's question. Who holds the constraint? Not the model companies. Not the data. The hardware. So my next business is the whole stack: modify the GPU, build the box, send the first prompt, and run a local inference network for the office. Nobody in the US sells that as one package, set up and ready to use.

gpuinfrastructureeconomicsinferenceBy Julian AbeledaProject generalCreated 2026-10-06Edited 2026-10-07

How do I make money in AI?

I spent the last four months learning the AI inference stack. I took three different GPUs and tried to understand them from first principles. An Apple M3 GPU (16 GB unified memory). An AMD Radeon RX 7900 XTX (24 GB VRAM). An NVIDIA RTX 5090 (32 GB VRAM). My repos are public.

I came to it as an auditor. I worked at PwC and KPMG. An auditor does not ask who is famous. An auditor asks one simple question.

Who has power?

The auditor's question

Every business has a limiting factor. One thing that, if you take it away, stops everything. Find that thing and you find who holds the power.

So what is the limiting factor for the AI companies?

If you guessed Sam Altman, you are wrong. If you guessed Dario, you are wrong. If you guessed data, you are wrong too. Data can be bought, scraped and generated.

Try a different test. Take away OpenAI's data centers. Can it survive? No. Take away Anthropic's. Same answer. They die.

Marketers in this space say AGI is here and coding is dead. People who say that do not understand the stack. If AI made software free, why am I paying for ChatGPT? Why are the model companies paying for anything? Ask what they spend their money on. The answer is chips, power and buildings.

The industry wants you to think the product is tokens. Tokens are a bill you pay every month. Hardware is something you own. Token pricing hides the hardware underneath it.

Who has the pickaxe?

The old line from Silicon Valley still holds. Sell the tools. Be the casino. In a gold rush you do not have to pick the miner who wins. You sell pickaxes to all of them. The house always wins.

I wrote about who prints the money in The Real Winners of the AI Wars. This paper asks the next question. If software is cheap and easy to copy, what about hardware?

That question stuck with me, because the answer is clear.

What I measured

I did not want to argue this from opinion. So I measured it.

When a model writes a reply, it writes one token at a time. For each token, the GPU reads the whole model out of memory. So the speed limit is simple. Memory speed divided by model size. The chip's math is mostly waiting. Little (1961) gives the rule for how much work must be in flight to keep a pipe full. Volkov (2016) shows why GPUs hide that wait. Both point at the same thing. Memory moves the work.

Here is what that looks like on two of my GPUs, from Chips and Speed. Model: Qwen3-8B, 4-bit, 4.68 GB of weights.

GPU and memoryMemory speed, publishedMemory speed, measuredLimit, tokens per secondTokens per second, llama.cppShare of the limit
Apple M3 GPU (16 GB unified memory)100 GB per second89.9 GB per second19.217.5691%
NVIDIA RTX 5090 (32 GB VRAM)1,792 GB per second1,700 GB per second364252.669%

The limit is the measured memory speed divided by 4.68 GB. The published speed is the number on the box. No chip reaches it.

Same model. Same program. About 14 times the speed. The code did not change. The hardware did.

Two bars on the same model and program. The Apple M3 GPU with 16 GB unified memory makes 17.56 tokens per second, 91% of its limit. The NVIDIA RTX 5090 with 32 GB VRAM makes 252.6, 69% of its limit.
The code did not change. The hardware did.

Nothing I ran went over its limit. Software can get you close to the wall. It cannot move the wall.

The bottleneck at every level

The wall moves as the buyer gets bigger. It stays hardware.

One GPU. Slow replies have two causes. The model is bigger than the GPU memory, so part of it sits in system RAM and is read from there, many times slower. Or the model fits and the memory is just slow. The table above is this case. The fix is more memory, or faster memory.

Two GPUs, or two machines. Now the model can be bigger than one card. How you split it decides what limits you.

  • Split by layers. Each GPU holds half the layers and runs them in turn. Only a few kilobytes cross the link per token, so the link is rarely the wall. But the GPUs take turns, so one user gets about the speed of one card.
  • Split every layer. Both GPUs work on every token at once. Now they must talk many times per token, and the wait on each message is the wall. A slow link loses the gain.

Thunderbolt used to be too slow for the second way. That changed in December 2025. macOS 26.2 lets Macs read each other's memory over Thunderbolt 5 (Gigazine, 2025). The wait per message fell from about 300 microseconds to under 10. Four Mac Studios ran Qwen3 235B at about 32 tokens per second, against 19 on one (Hardware Corner). Four machines, 1.7 times the speed. The link still sets the ceiling. A better link raised it.

A data center. Hundreds of GPUs. Now the wall is the network and the power. NVIDIA's answer is a rack of 72 GPUs wired so each one talks to the others at 1.8 TB per second, over 5,000 copper cables (NVIDIA, GB200 NVL72). That is more than 14 times a PCIe 5 slot.

At every size the fix is something you buy and wire. Memory, then the link, then the network and the power.

The economics

The second question is about the moat. If your software can be one shot by a model, how long does your business last? How hard is your product to copy?

Hardware is hard to copy. That is the point.

Here is the business, from my planning file. It is the whole stack, in one company:

  1. Modify the GPU. Give a consumer card the memory its chip can already use.
  2. Build the box. Several modified cards in one machine, tested before the buyer pays.
  3. The first prompt. The model is installed and answering before the box leaves.
  4. The network. Every computer in the office points at the box. Local inference, on site.

Each step exists somewhere. In the US, nobody sells all four as one package: modified GPUs that arrive as local inference, already set up. The first product is the RTX 4090. Shops in China already rebuild it from 24 GB to 48 GB of VRAM. They move the GPU chip onto a new board with memory on both sides, and add 12 more memory chips. A US shop, GPU Lab, sells them with a 90 day warranty (gpulab.net). Many US companies, labs and universities will not buy from China. They need a US invoice, a warranty and fast delivery.

Four of those cards in one box gives 192 GB of GPU memory. That is enough to run DeepSeek-V4-Flash, a 284 billion parameter model, on site (DeepSeek-V4-Flash-0731). A community vLLM fork documents exactly this four card setup (github.com/yhfgyyf/vllm-deepseek-v4-sm89).

So what do I sell? Four cards, or one box?

What I sellRevenue, illustrativeHardware cost, illustrativeGross dollars
Four rebuilt RTX 4090 48 GB cards, sold one at a time$6k to $8k eachabout $4k to $5k eachabout $4k to $16k total
One tested box with four RTX 4090 48 GB cardsabout $45kabout $20k to $24kabout $21k to $25k per box

The box wins. A buyer next to eBay compares four odd graphics cards on price. Nobody buys four odd graphics cards. They buy a private AI system that works. So the box is the product. Single cards are the cash flow on the side.

The real margin sits on top of the box. Installation. Model setup. Remote support. Wiring it into the software the business already uses. That is where the hard skill gets paid.

And I do not gamble on stock. The buyer pays a deposit. I build to a signed parts list. The box passes tests we agreed on first: every card uses its full 48 GB, the model runs across all four, the temperatures hold. Then the buyer pays the rest.

The closest product on the market is TinyBox, from Tiny Corp, the company behind tinygrad. Its shop sells the RTX 5090 box (Tiny Corp shop, tinybox green v2) and the RTX PRO 6000 box (Tiny Corp shop, tinybox green v2 with RTX PRO 6000). The 5090 box is an older listing. It is still for sale, but the tinygrad.org home page no longer shows it.

SystemGPUsGPU memoryPricePrice per GB of GPU memory
TinyBox Green v24 × NVIDIA RTX 5090128 GB$50,000about $391
TinyBox Green v24 × NVIDIA RTX PRO 6000 Blackwell384 GB$91,000about $237
My appliance, target4 × NVIDIA RTX 4090, rebuilt to 48 GB192 GB$44,995about $234

Five ways to run DeepSeek-V4-Flash on site, against its 167 GB of weights. TinyBox with four RTX 5090 cards, 128 GB for $50,000, does not fit. Renting three H100s on Modal, 240 GB, fits for $103,780 a year, every year. TinyBox with four RTX PRO 6000 cards, 384 GB, fits for $91,000. Four stock RTX 4090 cards, 96 GB, do not fit. The same four chips upgraded to 48 GB each, 192 GB, fit for a $44,995 target.
The upgrade moves the same chips past the line.

Price per GB is not the question a buyer asks. The buyer asks: what is the cheapest way to run this model? The model needs 167 GB for its weights alone (DeepSeek-V4-Flash-0731). Four stock RTX 4090 cards have 96 GB. They cannot run it. The same four chips with 48 GB each have 192 GB. They can. The upgrade is the whole difference.

That upgrade is the skill I sell. NVIDIA already puts this chip, AD102, on a 48 GB card. It is the RTX 6000 Ada, a workstation card sold at a workstation price (NVIDIA, RTX 6000 Ada). The gaming version, the RTX 4090, ships with 24 GB (NVIDIA, RTX 4090). The chip is the same. The memory decides the price tier. NVIDIA and AMD are not paid to give a careful buyer the most memory per dollar. They are paid to sell the next tier up. A shop that moves the chip to a 48 GB board sells the buyer what the tiers hold back. The RTX 4090 is the first card, not the only one. Any card whose chip can use more memory than it ships with is a candidate. The skill is the same. Selling it as the full package is the business.

The hardware and parts for one box cost $23,967.70 in my working case. With labor, reserves and delivery it is $29,839.70. At $44,995 that leaves about $15,000 per box before overhead and tax. That is $5,005 less than the TinyBox with four RTX 5090 cards, which cannot run the model. It is under half the price of the TinyBox that can, and cheaper per GB of GPU memory: about $234 against $237.

Tokens are the service. Hardware is the tool.

Now the other side of the ledger. How much are people spending on tokens?

A lot. Enterprises spent $37 billion on generative AI in 2025. $12.5 billion of it went straight to model APIs (Menlo Ventures, 2025). Uber burned its whole 2026 AI budget in four months (Fortune, 2026). I covered this in How BoltBeam Works.

Look at who sells those tokens. OpenRouter runs about 1.5 quadrillion tokens a year for more than 8 million developers (Menlo Ventures, May 2026). OpenRouter owns no GPUs. It forwards the request and takes a cut. Modal owns no GPUs either. It rents an NVIDIA H100 for $3.949 an hour (Modal pricing, October 6, 2026). Run that all year and the rent is $34,593. For one GPU. Every year.

That is the tell. Tokens are the service. The obvious service of AI. Hardware is the tool. Everyone selling the service is standing on somebody else's tool, and paying rent on it.

Four stacked layers. A router, OpenRouter, owns no GPUs. GPU rental, Modal, owns no GPUs and rents an H100 at $3.949 an hour. Tokens, spent once. At the bottom, in blue, the tool: the GPU and the box, the only thing that is owned.
Every layer above the tool pays rent on it.

So why rent when you can own?

Rent never stops. A tool gets paid off. And the moment you own the tool, you are not only a buyer. You can sell the service too. That is the entire business of every provider behind OpenRouter. Rent GPUs, load an open model, sell the tokens. Own the GPUs and you skip the first step.

And here is the question nobody is asking. What is a token worth after you spend it?

Nothing. A token is used the moment you buy it. It is gone. And the next token is better. Every few months a new model comes out that is smarter and cheaper per token. So the token you bought today was worth the most on the day you bought it. It only loses value from there.

That is the opposite of Bitcoin. Bitcoin is scarce, so holding it is the bet. Tokens get more plentiful and better every year. Better tokens only exist in the future. Every dollar spent on tokens today compounds negatively.

Now flip it. The box does not care which model it runs. When a better open model comes out next year, it runs on the same box. The tool gets better tokens for free. The token buyer pays again.

Two charts side by side. Left: Bitcoin on November 1 each year, $60,950 in 2021, $20,480 in 2022, $35,440 in 2023, $69,467 in 2024, $110,052 in 2025. Right: the price of the same model answer, $60 per million tokens in 2021 and $0.06 in 2024, ten times cheaper every year.
Bitcoin goes up over time. A token is always beaten by a better model.

Bitcoin has a fixed supply, so over the years it has gone up. It went from $60,950 to $110,052 between November 2021 and November 2025, with a 66% fall on the way (Coinbase, BTC-USD daily closes). A token goes the other way. There is always a better model. In November 2021 a model at GPT-3's level cost $60 per million tokens. In November 2024 a model at the same level cost $0.06. That is 10 times cheaper every year (a16z, LLMflation). Epoch AI finds the same thing, 9 to 900 times a year depending on the task (Epoch AI). Money spent on tokens buys an answer that will soon be cheap.

Owning buys things a token price cannot. The data never leaves the building. The price cannot change next quarter. Nobody can switch the model off.

One honest check, because I am an auditor. Owning only pays when the tool is busy. So how busy is busy?

Here is my own usage. I am one person. I use AI to build deterministic software: kernels, a measurement tool, a training pipeline, this site. In 48 active days Claude Code processed 35.2 billion tokens for me, over 89 sessions (Claude Code stats, October 6, 2026).

Claude Code stats, all time. Favorite model Opus 5. Total tokens 35.2 billion. 89 sessions. 48 active days. Longest session 4 days 12 hours. Longest streak 17 days.
One person. 35.2 billion tokens in 48 active days.

That is one tool. I use Codex too. It shows another 34.2 billion tokens over 1,651 chats (Codex stats, October 6, 2026).

Codex stats. Lifetime tokens 34.2 billion. Peak tokens 1.2 billion. Longest chat 11 hours 37 minutes. Longest streak 39 days.
The second tool. Another 34.2 billion tokens.

Together that is about 69 billion tokens, from one person, in a few months. That is the number that matters. AI is worth this much to me because it builds things that keep working, and my usage only goes up.

The models are getting more efficient. The same answer gets about 10 times cheaper every year (a16z, LLMflation). DeepSeek is the clearest case. DeepSeek-V4-Flash is a 284 billion parameter model that fits in the box's 192 GB, and on OpenRouter it costs $0.018 per million input tokens and $1.28 per million output tokens (OpenRouter models list, October 6, 2026).

That is why I think 192 GB is the sweet spot for value per token. It is the smallest box that fits this model: 167 GB of weights, with about 25 GB left for the conversation cache. Below it, 128 GB cannot run the model. Above it, 384 GB costs $91,000.

What is the 25 GB for? It is the working memory. Every conversation the model is in has a cache: what it has already read, kept so it does not read it again. The cache grows with the context, the length of the conversation. The 25 GB also has to hold the model's scratch space while it runs, so the real room for cache is less.

DeepSeek's newest models compress that cache hard. vLLM measured 9.62 GiB per conversation at a full one million tokens, and about half that when the cache is stored in 8 and 4 bits (vLLM, DeepSeek V4). That works out to at most about 4.9 KB per token. V4-Flash is the smaller model, so its real number is lower.

Context of one conversationCache it needs, at most
32 thousand tokens0.16 GB
128 thousand tokens0.65 GB
1 million tokens5.2 GB

So the 25 GB holds about five conversations at a full million tokens, or about 39 at 128 thousand, before the scratch space. The fork for this four card setup tested 8, 32 and 128 thousand token inputs, and four 8 thousand token requests at once (vLLM DeepSeek-V4 fork).

That is enough for the conversations happening right now. Offloading is for all the others. My sessions run for hours, and each turn rereads the whole conversation. When a cache is thrown away, the next turn has to read the entire conversation again from scratch. Offloading keeps old caches on cheaper memory, so a conversation that comes back reloads its cache instead.

Where the cache livesExample sizeContext it can hold, at most 4.9 KB per token
GPU memory left after the model25 GBabout 5 million tokens
System RAM256 GBabout 52 million tokens
SSD4 TBabout 812 million tokens

The RAM and SSD sizes are examples, not the box's final parts list. Three standard techniques do this:

  • A smaller cache. Store it in 8 bits instead of 16, which halves it (vLLM, quantized KV cache). The four card fork already does this.
  • Offload to system RAM. Move caches that are not in use to the CPU's memory. It is slower than GPU memory and far faster than reading the conversation again.
  • Offload to SSD. LMCache moves caches through GPU memory, CPU memory, then local SSD (LMCache). SGLang's HiCache does the same (SGLang HiCache).

Total context is what matters for usage like mine. A box that can keep hundreds of millions of tokens of conversation warm on a cheap SSD can bring back yesterday's session without reading it all again.

Cheaper tokens make renting cheaper too. I am not hiding that. On price alone, one person can still rent today. But the box gets better with time, and rent does not. Every year a better model fits in the same 192 GB, and my token count keeps climbing. At some point I would rather own the hardware than keep paying rent on it. An office that works like me gets there first.

Why not just use ChatGPT?

The obvious question. ChatGPT already does this. Claude already does this. Why own a box?

Because I see the bigger picture. The model companies price their apps below what the hardware costs them. That buys two things. Your conversations, to train the next model. And your habits, so you reach for their app first.

The training part is in their own terms. On ChatGPT's consumer plans, your chats can train OpenAI's models unless you turn off "Improve the model for everyone" (OpenAI help). Claude's Free, Pro and Max plans, Claude Code included, ask you to choose. If you allow it, Anthropic keeps that data for five years (Anthropic, 2025).

The habit part is my view. They want you to depend on Claude or ChatGPT, not to learn. Each answer you get from them is one more thing you did not learn to do, and one more industry their models move into.

Here is the best way I know to think about it. Treat AI as a wish. The best wish is not to solve one problem. The best wish is more wishes.

A subscription gives you wishes on someone else's terms, at someone else's price, with someone else keeping the record. A box you own gives you unlimited wishes. You can go slower, make mistakes, and learn, and none of it is metered.

Who is doing this?

Every part of my plan already exists. I could not find anyone who sells all of it together.

Modding the cards. GPU Lab in Michigan turns a 24 GB RTX 4090 into a 48 GB card (GPU Lab). It sells single cards, not systems. Shops in Shenzhen move the 4090 chip onto new boards with 48 GB (Igor's Lab). One seller lists a single workstation with a modded 48 GB card (Jawa listing).

Selling boxes with several GPUs. Tiny Corp sells the TinyBox, from the table above. NVIDIA sells the DGX Spark and the DGX Station (NVIDIA, DGX Station). Lambda, Bizon, Puget Systems, Exxact and Comino sell workstations and servers (Lambda, Bizon, Puget Systems, Exxact, Comino). All of them use stock cards.

Private AI for small businesses. VRLA Tech ships servers with Ollama installed, ready when you power them on, with lifetime US support. It aims at teams of 2 to 15 people (VRLA Tech). Zylon sells private AI software to regulated industries (Zylon). Both use stock hardware.

LayerWho does itWhat is missing
Modified GPUsShenzhen shops, GPU Lab in the USSold as single cards, no tested multi card system, no software
Multi GPU boxesTiny Corp, NVIDIA, Lambda, Bizon, Puget, Exxact, CominoStock cards, so more money per GB
Ready to use private AIVRLA Tech, ZylonStock cards, smaller models
All of it, for a small or medium businessnobody I found

That is the gap. The whole stack as one package, in the US: modified GPUs, four in one tested box, a large open model already installed, the office's computers already pointed at it, guardrails, and support. The open question is trust. Nobody selling boxes has yet asked a business to trust a modded card. That is what the tests before final payment are for.

Putting it all together

This is the end result of the business plan. The box is one central place for local inference, serving every computer in the office. It arrives ready. Plug it into the network, turn it on, and the inference is already set up for an enterprise, with its own guardrails.

Four computers point at one box. Two laptops and a desktop run Arkey on the Server route, with Codex, Crush or Kimi Code. A browser runs DeepSeek Harness web chat. The box has four RTX 4090 48 GB cards, 192 GB.
One box. Every computer in the office.

The hardware is hard. The software should not be.

I can promise that because I have worked on every layer of this stack myself.

  • Kernels. In my tinygrad fork I wrote the GPU code that runs the model (github.com/JulianAbeleda/tinygrad-arkey). One decode kernel ran on AMD, Apple and NVIDIA (A Universal Kernel).
  • Measurement. BoltBeam checks each kernel against the speed limit of its chip (How BoltBeam Works).
  • Training. DayCare teaches a small model one skill, without the usual black box (DayCare).
  • The client. Arkey, below, is what a person actually touches.

The box boots into one menu. Arkey. It is open source (github.com/JulianAbeleda/arkey_v3). Pick the TUI and you are talking to the model in a terminal coding agent.

The agent is swappable. Arkey runs Codex, Kimi Code or Crush, and Claude Code is next. Here it is Crush, Charm's open source terminal agent (github.com/charmbracelet/crush). Arkey does not ship any of them. It copies the one you already installed into its own folder, then points it at the model you chose. Your normal install is left alone. Whichever agent your team already knows is the one it uses on the box.

Not everyone wants a terminal. For a chat window in the browser, the box runs DeepSeek Harness, DeepSeek's open source agent harness with a web UI (github.com/deepseek-ai/deepseek-harness). Pointing it at the box is one form: a custom provider whose base URL is the box's address on the office network.

The DeepSeek Harness web chat. A sidebar with New Session, Plugins and workspaces. In the middle, 'Into the Unknown' above a message box that says 'Describe what you want to build', with the model set to DeepSeek-V4.1-Flash.
DeepSeek Harness in the browser. Point its provider at the box.

The Arkey boot menu in a terminal. Three choices: TUI, choose the modded harness, Arkey Crush. Config, routes and hardware. Exit, close Arkey.
Arkey boot. One menu.

Config has four choices. That is the whole setup.

The Arkey config menu. Local, runtime to installed model. Server, a llama.cpp server on your network. Frontier, a hosted AI provider. GPU Auto-scan, detect and align llama.cpp.
Arkey config. Four routes.

  • Local. Run the model on this machine.
  • Server. Use a llama.cpp server on your network. That is the box. Every computer in the office runs Arkey and picks the same server.
  • Frontier. Use a hosted provider when you need one.
  • GPU Auto-scan. Find the GPU and set llama.cpp up for it.

Own the box. Point Arkey at it. Done.

Why most people will not do this

I am not hiding the plan. Anyone can do it. Most people will not, because hardware is hard.

Hard means three things. The skill only comes from hours at the bench. Every mistake costs a real card. And no model can do it for you.

Reballing memory on a GPU is hard. Moving a GPU chip onto a new clamshell board is hard. Building a four card machine that does not cook itself is hard. Networking and kernel engineering are hard, even though they are software.

This is my partner and me, at the bench. Before we touch a 4090, we practice on an AMD board, to test that we have the skill.

Hot air from above and below. Practice on an AMD board.

Lifting the memory chips off the board. First 30 seconds.

Reballing the memory under the microscope. First 30 seconds.

Watch it done. Gamers Nexus filmed an RTX 4090 being rebuilt from 24 GB to 48 GB, by hand, in Brother Zhang's repair shop in China.

Gamers Nexus, Creating a 48GB NVIDIA RTX 4090 GPU. On YouTube.

You cannot prompt your way into a solder joint. You cannot one shot a thermal test. A model can write the plan. It cannot hold the hot air station.

Hard is the moat. It always was.

The end game

At the highest level the scarce thing is no longer only chips. It is power, and a place to connect it. Microsoft's CEO said in 2025 that he has chips in inventory he cannot plug in: "I don't have warm shells to plug into" (DCD, 2025). In the US a new power project waits about five years to connect to the grid (LBNL, 2025). The IEA says about one in five planned data centres could be delayed by grid limits (IEA, 2025).

China is building for this. In 2022 it started a national plan, East Data West Computing, to put data centers in its energy rich west (Sixth Tone, 2022). In 2025 it added more than 430 GW of wind and solar. BloombergNEF expects China to add more than six times as much generating capacity as the US over the next five years (Al Jazeera, 2026).

The US treats this as a race. Its 2025 AI Action Plan names China, and orders faster permits for data centers and the power plants behind them (White House, 2025).

Elon Musk is looking past Earth for the same reason. He plans to merge xAI into SpaceX. He said: "The lowest-cost place to put AI will be in space, and that will be true within two years, maybe three at the latest" (Fortune, 2026). In orbit the sun shines almost all the time. There are no neighbors and no grid queue. The power is not free. You pay for the launch, the panels and the radiators. A Georgetown researcher in the same article says it is "just technologically not feasible at the moment."

Power that does not wait in a queue is the last moat.

References