arkey researchRSS
All posts / Journal
note2026-09-28training

Why I Post-Train My Models

Investors told me to own the stack. For an AI product the model is the layer that decides what happens, and if I don't train it, someone else trains it for someone else's product. So I own every layer: the product, one app that trains, one engine that runs and trains the model, one GPU. It is more work. Here is why I do it anyway.

trainingloradaycaretool-callingBy Julian AbeledaProject daycareModels nemotron-3-nano-4bCreated 2026-09-28Edited 2026-09-28

Technical words are explained in plain language at the bottom, under Words used here.

Own the stack

The best advice I got from investors was three words. Own the stack.

Every layer you don't own is a dependency. A dependency sets your price. It sets your roadmap. It decides how you fail.

You can see where the money goes. When Andreessen Horowitz mapped generative AI, infrastructure vendors were "capturing the majority of dollars flowing through the stack" (a16z, 2023). Some of the developers they talked to went further: "the product is the model" (a16z, 2023).

So I asked the dumb question. What is the stack of an AI product, and who owns each layer?

Who owns what

Here is a product that buys tokens, next to mine.

Two columns of six layers. If you buy tokens: the product is yours; what the model is trained to do belongs to the model maker; running the model to the API provider; training it happens only on their terms; the hardware is their data center; price, limits and retirement are theirs. Mine: GameTerm, DayCare, tinygrad-arkey, the same engine, one GPU on my desk, and me.
Buy tokens and you own one layer. Own the stack and you own all six.

When you buy tokens, you own one layer. The product. Everything under it belongs to someone else.

That is how almost everyone does it. Enterprises spent $12.5 billion on foundation model APIs in 2025, and open-source models, the kind you can run and train yourself, hold "only 11% of today's market" (Menlo Ventures, 2025). Companies say they "want direct access to the latest model", so "more enterprises are hosting either directly with model providers or via Databricks" (a16z, 2025).

The layer that matters most is the second one. The model is the part that decides what happens. The model maker trains it for everyone: for chat, for benchmarks, for the average customer. Not for your tools. Not for your rules. When they change it, your product changes with it. When they retire it, you follow. OpenAI's rule is short: "At the time of the shut down, the model or endpoint will no longer be accessible" (OpenAI, deprecations). Anthropic promises "at least 60 days' notice" (Anthropic, model deprecations). Some providers will fine-tune a model for you, but only the models they list, on their platform (OpenAI, model optimization).

That is the dependency. If I don't train my models for GameTerm, I am relying on people who make models for somebody else.

What it looks like in practice

GameTerm is my agent harness. A model sits in a terminal with 35 tools and does things for you.

A small model calls those tools fine. The trouble is the turn after. The calculator rejects a call, and the model gives up on the calculator and reaches for a shell. A file is blocked, and the model tries another way in. Sometimes it just says nothing.

No model maker is going to fix that for me.

Two pieces

The popular way to post-train a model uses two engines. One writes the model's answers fast. The other one learns from them. After every step, the new numbers get copied from one to the other. Hugging Face's library describes it plainly: "The server only generates. After each optimizer step the trainer streams the updated weights into it" (TRL docs). The two engines do the math slightly differently. With the same numbers loaded, "they can produce significantly different token probabilities" (Yao et al., 2025). So the one that learns is learning from answers the other one never quite gave. I wrote about how many files that route needs in Simple AI Training for Smart People.

I cut the training stack down to two pieces.

Left, the popular way: engine 1 writes the answers, engine 2 learns from them, and new numbers are copied back after every step. Right, mine: one engine, tinygrad-arkey, writes the answers and learns from them, and one app, DayCare, holds the lessons, the marks and the record.
Two engines copy numbers back and forth. One engine has nothing to copy.

  • One engine. tinygrad-arkey, my fork of tinygrad, runs the model and trains it. Same engine, same math. Every step it checks that what it learned from is exactly what it wrote. Nothing to copy, nothing to drift apart.
  • One app. DayCare builds the lessons, marks the answers, runs the loop, and keeps the record.

The product stays the product. The trained model goes back into GameTerm the same way a stock model does.

How it learns

Training a model after it ships is called post-training. Here is the whole idea, step by step.

  1. Take real moments. A real request in GameTerm, the model's real tool call, and GameTerm's real answer, down to the byte.
  2. Let the model try. From each moment, the model writes its next turn several times.
  3. Mark each try. Right, wrong, or no answer. No answer is marked worst.
  4. Push toward the better tries. Each try is compared with the others on the same moment. Better than average becomes more likely. Worse becomes less likely. This method is called RLOO, REINFORCE Leave-One-Out: each try is judged against the average of the others, leaving itself out.
  5. Change very little. The model's own numbers never move. A small patch beside them learns instead. The patch is called LoRA, Low-Rank Adaptation. A leash, the KL divergence (Kullback–Leibler divergence), stops the patch from pulling the model too far from where it started.
  6. Stop if it goes wrong. Automatic checks watch every step. If the model drifts, rambles, or gets worse, training stops by itself.
  7. Grade on an exam it never saw. The test questions are kept apart from the lessons, and what counts as a pass is written down before training starts.

Five boxes in a loop: a real moment, try it 8 times, mark each try, push toward the tries that beat the others, a small patch learns while the model stays frozen, then repeat. Checks watch every step; the exam is kept apart and the pass mark is written first.
Try, mark, push toward the better tries, repeat.

Because a program, not a person, marks right and wrong, the whole approach has a name: RLVR, reinforcement learning with verifiable rewards.

That is it. No second engine. No pile of formats.

What it bought

The first model I did this for is NVIDIA's Nemotron 3 Nano 4B, "a small language model trained from scratch by NVIDIA". One run, 86 minutes, one GPU.

When the calculator rejected its call, it used to get the answer right 76% of the time. After training, 92%. Answers left blank fell from 97 to 16 out of 2,620. Wrong answers fell too, so it was not just guessing to fill the silence.

One example. "A farm collects 3,095 eggs and packs them into cartons of 36. How many are left over?" The calculator rejects the model's first call. Before training, it got this right 7 times out of 20, and most of the misses gave up on the calculator. After training, 19 out of 20. It fixes the call and answers 35.

When a file is blocked, the rule is the one Codex and Claude Code follow: ask for permission once, then say plainly that you can't. Codex asks for "a short question asking for approval" (Codex CLI). The trained model got better at that too, from 27% to 39%. That is less than I predicted, and I will say so.

It also held up on 78 blocked situations worded nothing like its lessons: 38% to 49%.

Then the last check failed. On 132 everyday questions, every plain answer held. But on math, it got two questions wrong that the original got right, and one right that the original got wrong. My rule, written before the run, is no net loss. So by my own rule, this patch failed.

Then I looked closer, asking each question 20 times instead of once. Two of the three were coin flips for both models. One was real. Asked for a total when one person spent "2/5 times more" than $400, the trained model now reads that as "plus 2/5" far more often: 7 misses in 20, against 1 for the original. It still uses the calculator. It just feeds it the wrong sum.

Across all 68 math questions, asked 8 times each, the trained model was level with the original. The check that failed had asked each question once, and that was the noise. So I shipped it anyway, as a written exception, not a pass. The rule stays. The exception is on the record, with its costs next to it.

The bigger lesson was about my alarms, not my model. From the fourth run on, every failure came from a check sitting inside its own noise. Now every check has to prove, on past runs, that it stays quiet on a healthy run and fires on a broken one, before any new run starts.

It took five runs to get here. Each of the first four failed one check, and each failure changed the plan for the next.

Do I have to?

This is the honest objection. If the big models work, why spend the time?

It is a fair question.

  • It costs time. Five runs. A fork to maintain. Exams to build. Someone buying tokens spends none of that.
  • The big models keep getting better. A model maker might fix the turn after a tool call next quarter, for free.
  • It is narrow. This patch covers a calculator and blocked files. Nothing else.

Here is my answer. The first run is the expensive one. The engine, the app, the exams, the checks, those get built once. After that, teaching GameTerm's model a new behavior is one more run on one GPU. It is not a request to a model maker. It is not a bigger bill.

Renting is cheaper on day one. Owning is cheaper on day one thousand. People who do the math find "exactly one point where the fixed cost of an idle-tolerant GPU beats the variable cost of a metered API" (bex.co, 2026). And nobody can change it under me.

The rule

If a behavior matters to my product, it should be trained into my model, on my stack, and graded by my exam.

If I have to ask someone else to get it, I don't own it.

Words used here

  • Post-training: training a model after its maker has released it, to teach it one specific job.
  • Inference: running a model to get answers. Buying tokens means paying someone else for inference.
  • LoRA (Low-Rank Adaptation): the patch. A small set of extra numbers trained beside a model whose own numbers stay frozen (Hu et al., 2021).
  • REINFORCE: the name of a basic trial-and-error learning rule from 1992: make the choices that scored well more likely (Williams, 1992).
  • RLOO (REINFORCE Leave-One-Out): trial and error. The model answers the same question several times, and each answer is scored against the average of the others (Ahmadian et al., 2024).
  • RLVR (reinforcement learning with verifiable rewards): training where a program checks each answer as right or wrong, instead of a person or another model judging it (Lambert et al., 2024).
  • KL divergence (Kullback–Leibler divergence): the leash. A number for how far the trained model has moved from the original. The further it moves, the harder training pulls it back (Ouyang et al., 2022).
  • Tokens: the pieces of text a model reads and writes. API providers charge by the token.
  • API (application programming interface): the door a program uses to talk to another company's service. Model APIs take your text and send back the model's answer.
  • GPU (graphics processing unit): the chip that runs a model. Mine is one NVIDIA RTX 5090.

Sources

What investors told me is my account. The published sources below were opened on 28 September 2026.

Investors and the market: Andreessen Horowitz (Bornstein, Appenzeller, Casado), "Who Owns the Generative AI Platform?", 19 January 2023, a16z.com. Andreessen Horowitz, "How 100 Enterprise CIOs Are Building and Buying Gen AI in 2025", 10 June 2025, a16z.com. Menlo Ventures, "2025: The State of Generative AI in the Enterprise", menlovc.com. bex.co, "The Self-Hosted GPU Breakeven Point", 8 July 2026, bex.co.

Providers: OpenAI, deprecations and model optimization. Anthropic, model deprecations. OpenAI Codex CLI, approval-policy prompt, commit 1b1835f7, on_request_rule_request_permission.md. Claude Code, sandboxing and permission modes.

Tools and models: Hugging Face TRL, vLLM integration. tinygrad and my fork, tinygrad-arkey. NVIDIA, Nemotron 3 Nano 4B model card.

My records: DayCare, 28 September 2026: "Run 5: PREDECLARED (run 4's recipe, seed 20260930, composition-adjusted entropy trigger, blocked rule (c))", commit 7ea1f93; "posttool_blocked: rule (c), after a refusal only one authority request, then a reply", commit 7946ebf. The run's plan, log and results are in DayCare's research/rloo-posttool-calculator-r5.md.

Papers: Williams, "Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning", Machine Learning 8, 1992, doi:10.1007/BF00992696. Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models", 2021, arXiv:2106.09685. Ouyang et al., "Training Language Models to Follow Instructions with Human Feedback", 2022, arXiv:2203.02155. Ahmadian et al., "Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs", 2024, arXiv:2402.14740. Lambert et al., "Tulu 3: Pushing Frontiers in Open Language Model Post-Training", 2024, arXiv:2411.15124. Yao, Liu, Zhang, Dong, Shang and Gao, "Your Efficient RL Framework Secretly Brings You Off-Policy RL Training", 5 August 2025, fengyao.notion.site.