Simple AI Training for Smart People, While the Not So Smart Admire Complexity
"An idiot admires complexity, a genius admires simplicity." Terry A. Davis said that. He wrote a whole operating system by himself, so he had earned the right.
Smarter people have said the same thing more politely. Edsger Dijkstra wrote it in 1984: "Simplicity is a great virtue but it requires hard work to achieve it and education to appreciate it. And to make matters worse: complexity sells better." (Dijkstra, EWD 896, 1984) Steve Jobs said it in 1998: "Simple can be harder than complex: You have to work hard to get your thinking clean to make it simple." (BusinessWeek, 1998)
AI training is the best example of this I know. Complexity sells there, and everyone admires it.
Count the files
How many files does it take to teach a model one thing?
I counted. On the popular route it is eight different file formats. Not eight files. Eight kinds of file, each with its own rules, each written by a different tool, each telling part of the story and none of it the whole.
This week I did the same job with one.
How it works
Two ideas get mixed up here, so I will keep them apart. One is the shape of the change. The other is how the model learns it.
The shape is LoRA. A model is mostly big tables of numbers. Retraining all of them is expensive, and it is an easy way to break what the model already knew. LoRA leaves every big table frozen. Beside a table it adds two thin ones, and only the thin ones learn. Multiply the two thin tables together and you get a correction the same size as the big table, added on top of it (Hu et al., 2021). On GPT-3, that cut the numbers being trained "by 10,000 times" (Hu et al., 2021). How thin the tables are is called the rank. In my trial-and-error runs the patch sat on the model's last block, at rank 4. That small patch is the file the rest of this post is about.
How the patch learns is the second idea, and there are two ways. The first is by example. You show the model the right answer and it copies the pattern. That is called supervised fine-tuning, or SFT.
The second is by trial and error. You ask the model a question four times and let it answer four ways. A checker marks each answer right or wrong. The model is pushed toward the answers that did better than the others. That method is called RLOO (Ahmadian et al., 2024). When the checker is exact, a plain right or wrong, the whole approach is called RLVR, reinforcement learning with verifiable rewards (Lambert et al., 2024).
So LoRA is not the alternative to RLOO. LoRA is the patch. Learning by example or by trial and error is how the patch gets its numbers. Both can train the same patch.
Here is the same trial-and-error job done both ways, one step at a time.
| Step | The popular way | DayCare |
|---|---|---|
| Model | safetensors and two JSON files | the GGUF it runs from |
| Questions | JSON Lines, copied into Arrow | tasks.json, still JSON today |
| Checker | Python, handed to TRL (docs) | Python, its fingerprint in the record |
| Patch | LoRA in safetensors and JSON (PEFT) | LoRA in adapter.xml, as in SFT |
| Progress | trainer_state.json, pickled optimizer.pt (Trainer) | run.xml |
| Before an update | not in the recipe I cite | odds checked twice, within 0.005 nats |
| Saving | a checkpoint folder | must read back exactly |
| Serving | merge, save, convert, shrink | adapter.xml to a GGUF patch |
On the left, trial and error means a second library, a second config and a second set of files beside the ones learning by example already made. On the right it is one more mode of the same trainer, writing the same two files.
Notice the name of that RLOO paper: "Back to Basics". Its finding is that "far simpler" methods beat the complicated ones everyone had been using (Ahmadian et al., 2024). Even the method is a case for simple.
This week DayCare, my training stack, did the second kind on NVIDIA's Nemotron 3 Nano 4B. The run was small, and I will not claim it made the model better. That is not the news.
The news is the files. The trial-and-error run read the same model file, wrote the same kind of patch and kept the same kind of record as the learning-by-example runs before it. One process. One format a person can read.
If it fits in one, why does everyone else use eight?
The "popular" way
Here is the popular way to do the same two things and then run the result on a laptop. Watch the formats pile up.
Your training data is a JSON Lines file. The Hugging Face datasets library loads it and saves a second copy as Apache Arrow, which is "the file format used by Hugging Face Datasets under the hood" (Hugging Face datasets). The model itself arrives as safetensors: an 8 byte length, then "a JSON UTF-8 string", then the raw numbers (safetensors). Beside it sit a settings file and a vocabulary file, both JSON.
Learning by example runs through a library called PEFT. When it saves the patch, it "saves three files": the numbers in safetensors, the settings in JSON, and a README (PEFT checkpoint format). Trial and error runs through another library, TRL, whose RLOO trainer takes the checker as Python code (TRL RLOO). Every save point along the way writes a JSON progress file next to pickled Python objects (transformers Trainer). The training log goes to yet another format, TensorBoard's, or to an online service (transformers callbacks).
Then you want to run it. You fold the patch back into the model and save the whole model again, after which "you can't use any of the PEFT-specific methods anymore" (PEFT checkpoint format). You convert it to GGUF, the format llama.cpp runs, and then shrink it into a second, smaller GGUF (llama.cpp). GGUF keeps its own description of the model in its own binary table (GGUF spec).
Count them. JSON Lines, Arrow, safetensors, JSON in four different layouts, pickle, TensorBoard, GGUF, and Markdown. Eight.
And what is the industry's answer to eight formats? More layers on top. Axolotl, torchtune and LLaMA-Factory each add one more settings file, in YAML (Axolotl; torchtune; LLaMA-Factory). torchtune writes its patch in the same safetensors and JSON pair "for compatibility with the PEFT ecosystem". Unsloth's export quietly downloads and runs llama.cpp's converter for you (Unsloth save.py). The formats are still there. The tools just hide them from you.
Each format solved a real problem on the day it was made. That is how complexity grows: one reasonable step at a time. What nobody did was go back and ask whether a model needs all eight at once. I asked people who train models for a living, and nobody could tell me why. A stack nobody can explain is not sophisticated. It is just old.
The simple way
DayCare starts from the file the model already runs from, the GGUF. It does not convert anything first. The model's own numbers stay frozen.
Learning by example writes one file for the patch, adapter.xml. It holds the patch's numbers and says what they fit: which model, how big the patch is, which layers it belongs to. A second record beside it, also XML, holds the settings and how the training went.
Trial and error writes the same adapter.xml. It can even start from a patch that learning by example made. And a run does not count unless its saved patch reads back exactly as it was written.
To run the result, one step turns adapter.xml into a GGUF patch, and llama.cpp loads it beside the model it already had. No folding it back in. No second shrink. The model file never changed, so there is nothing to convert back.
The training data goes in as XML too, though DayCare still accepts JSON there.
Every record is flat XML. Flat means every value is one line with a name, a type and the value itself:
<field name="source_id" type="str">train-60</field>
<field name="stock_calls" type="int">33</field>
<field name="sft_err" type="bool">true</field>
The code that reads and writes it is 61 lines of plain Python. Whatever goes in as a number comes back out as a number.
That is the whole trick. It does not look impressive. That is the point.
Why simple wins
It is not better because the model scores higher. A file format does not teach anything. It is better because there is one system, and I can read it.
Getting to one format meant throwing things away and checking that nothing was lost. But once you are there, the work gets lighter every day.
On the popular route, a question about a run means holding a map in your head. The data is in one file, the patch size in another, the training curve in a third, the numbers in a fourth, and later in a fifth that describes them all over again. Each one tells part of the story in its own layout. None of them tells the whole story. That map is the complexity people admire, and you carry it for them.
On DayCare there is one shape to remember. Any record opens in a text editor and says what it is. There is no manual in another file to go and find.
I feel that in three places every day.
Reading. Learning by example or by trial and error, the record looks the same. I only had to learn one.
Searching. A flat record answers questions with grep, the plain text search every computer has. On 24 September a run showed Nemotron stuck in a loop on its calculator. One search over the run's record found why: when the calculator refused a call, the harness answered with the rules for writing files. The bug was in the harness, not the model, and it had been there for weeks.
Trying things. Every experiment reads and writes the same shape. The trial-and-error trainer did not need a format of its own. It reused the patch, the record and the export that learning by example already had.
One source, every format
The first thing people say is: but everything speaks JSON. Fine. So does DayCare, whenever it needs to.
A flat record reads back into plain values: text, numbers, true or false, lists. Turning that into JSON is one call, json.dumps, and nothing is lost on the way, because the XML already says what type every value is. DayCare does it on every run: the inputs GameTerm needs are written out as JSON for it, and thrown away after.
The model works the same way. adapter.xml becomes a GGUF patch that llama.cpp runs. A whole DayCare model package becomes a GGUF file through one export step, which tries every format the package asks for. The Hugging Face map is still a placeholder, and I will say so.
That is the difference. On the popular route, eight formats each hold a piece of the truth, and every move between them is another conversion step. On DayCare, one format holds all of it, and every other format is made from it. Compatibility is not a feature I bolted on. It falls out of the design.
So even if the whole world switched to JSON tomorrow, nothing here would change. JSON would just be one more export.
The rule
One format, flat and typed, for everything I will need to read later. The model's numbers are data inside it, never the only copy of what they mean.
If I cannot answer a question about a run with a text editor and a search, the record is wrong, not the question.
Words used here
Fine-tune: change some of a trained model's numbers so it behaves differently.
LoRA: a small patch of numbers trained beside a model that is otherwise left alone.
RLOO: trial and error with a fair baseline. The model answers the same question several times, and each answer is scored against the average of the others.
RLVR: reinforcement learning where a checker, not a person, decides if the answer is right.
safetensors, GGUF: two common file formats for a model's numbers. GGUF is the one llama.cpp runs.
Flat XML: an XML file where every value is one element with a name and a type, and nothing is nested except lists.
Grep: the command-line search that finds lines of text in files.
Sources
DayCare, 15 July 2026: "Persist the Spryt self as flat XML, not JSON"; "M1+M2: flat adapter XML artifact; consolidate emits it (drops safetensors)"; "M4: emit a GGUF from the .xpkg".
DayCare, 20 to 22 September 2026: "Add standardized native tool SFT runner"; "Add finite-action RLOO verifier-driven training smoke for Nemotron"; "Add bounded Countdown RLOO pilot and lossless live adapter evaluation"; "Make experiment records XML-only". 24 September 2026: the GameTerm rejection-text fix (TC-060).
Quotes: Edsger W. Dijkstra, "On the nature of Computing Science", EWD 896, 10 August 1984, transcription. Steve Jobs, "There's Sanity Returning", BusinessWeek, 25 May 1998, Bloomberg. Terry A. Davis: widely attributed, no primary recording found.
Papers: Hu et al., "LoRA: Low-Rank Adaptation of Large Language Models", 2021, arXiv:2106.09685. Ahmadian et al., "Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs", 2024, arXiv:2402.14740. Kool, van Hoof and Welling, "Buy 4 REINFORCE Samples, Get a Baseline for Free!", 2019, OpenReview. Lambert et al., "Tulu 3", 2024, arXiv:2411.15124.