arkey researchRSS
All posts / Journal
note2026-09-28tool calling

PSA: Coding Slop with AI is a choice. Here's how not to do it.

tool-callingagentic-workflowsoftware-designgitsize-budgetsBy Julian AbeledaProject arkeyCreated 2026-09-28Edited 2026-09-28

Coding agents are very good at making code appear.

That is also the danger.

Give an agent fifty tasks and it can solve every task in front of it while making the system behind those tasks worse. One task adds a helper. The next adds a wrapper around the helper. A third puts the same setting in a second file because that is the easiest place to make the screen work. Every answer looks reasonable by itself. Six months later, nobody knows which file is in charge.

That is code slop. It is not just ugly code. It is code that solved today's prompt by making tomorrow's prompt harder.

I use three controls to stop it:

  • coding principles are the map: where should this change live?
  • sz.py is the wall: how much did the repository grow?
  • commit hooks are the gate: did the change prove what it claims before it entered the history?

Two left-to-right lanes compare coding with and without controls. Without controls, a task becomes an isolated patch and leaves the next agent a more confusing system. With controls, the task passes through coding principles, a size budget, and commit hooks before becoming a small reviewable commit.
Coding slop is a choice: a map, a wall, and a gate turn agent speed into reviewable progress.

The point is not to slow the model down. The point is to stop its mistakes from compounding at the same speed as its successes.

Where the three pieces came from

I did not invent these three ideas together.

I got sz.py from George Hotz. Tinygrad keeps a line counter at the root of the repository. The first Python version was committed by Hotz in May 2023 as “move line counter to python.” The current file still makes the size of the project visible. I took that idea and turned it into budgets for the parts of my own repositories.

I got the commit discipline from Google. Google's engineering guide says a change should do one self-contained thing, include its tests, and leave the system working. It also treats the description as a permanent record that says what changed and why (small changes; change descriptions). Gerrit supplies a commit-msg hook that runs when a commit is created. My hooks enforce different local rules, but the lesson is the same: make the normal path carry the discipline.

I got the coding principles from work. I spent years in audit and systems implementation. In an audit, you find the authoritative record, trace what moves through the system, reconcile the total, and investigate what does not tie. In software, the names change but the problem does not. If two files both define the same rule, there is no authority. If a change spreads everywhere, the system has no clean boundary. If the result cannot be tied back to the task, it is not finished.

I assembled those lessons for a new problem: a coding agent can now make changes much faster than a person can understand them.

The map: coding principles

A complex system is hard to change when nobody can answer one question:

Where does this decision belong?

My principles give the agent a map:

  • centralize what defines the system;
  • modularize how the system carries work out;
  • abstract only when the interface becomes simpler;
  • keep unrelated concerns orthogonal, so changing one does not disturb the other.

Take a real kind of problem from GameTerm: should a newly discovered application be approved for the agent to use?

Without a map, an agent can make the toggle look approved in Settings and add a separate default in the tool executor. The screen says yes. The action says no. Both pieces of code “work,” but they disagree because each became its own authority.

With the principles, there is one application-permission policy. Settings reads it. The tool executor asks it. A fresh install and an existing user can have different rules, but those rules live in the same place.

The principle did not write the code. It told the agent where the truth had to live. That is what a map is for.

This connects directly to Simple AI Training for Smart People, While the Not So Smart Admire Complexity. That paper argues for one readable source instead of parallel files that drift apart. The repository needs the same thing: one source for each important decision.

The wall: sz.py

An agent has almost no natural resistance to adding code. Adding a new file is cheap. Adding another helper is cheap. Keeping both forever is expensive, but that bill arrives after the task is over.

sz.py makes the bill arrive now.

It counts the nonblank lines the repository owns, shows which part owns them, and compares them with limits. Mine can answer:

  • how large is the shipping product?
  • which subsystem grew?
  • did code land in a directory nobody accounts for?
  • is one file becoming a system by itself?

Imagine an agent solves a small permission bug by adding 350 lines: a new policy object, a compatibility wrapper, a cache, and tests for all three. The feature works. Without a size account, “tests pass” ends the conversation.

With sz.py, the change has to explain why a small behavior needed 350 permanent lines. Maybe it really did. More often, the count reveals that an existing authority should have been repaired instead of surrounded.

The wall is not “short code is good.” Deleting tests to lower the count is bad. Compressing readable code into a puzzle is bad. The wall says something narrower:

Growth is a decision, not a side effect.

That is the same role the roofline plays in Chips and Speed: a control total does not explain the whole result, but it gives you a number the story must reconcile against.

The gate: commit hooks

Agents are good at saying what they intended to do:

I updated the setting, added a test, and everything passes.

A repository needs evidence, not a sentence.

My hooks turn a commit into a gate:

  • the commit message must name the part of the system that owns the change;
  • the fast checks and size account run before the commit is accepted;
  • the full checks run before the branch leaves the machine.

Suppose the agent changed the Settings screen but forgot the executor test. The fast gate fails. Suppose the change passes its focused test but breaks the packaged app. The full gate fails. Suppose the subject says only “fix bug.” The message gate rejects a history that will be useless next month.

The important part is that the agent goes through the hook. It does not tell me what the hook would probably have said. It makes the commit and lets the repository answer.

This is the same preference behind BoltBeam: a claim becomes a result only after it passes a stated gate, and the verdict stays in a record another person can inspect.

One task, two futures

Here is the difference in practice.

The task is: “New applications should start approved unless the user explicitly turns them off.”

Without the setup, the agent searches for the visible toggle, changes its default, adds a helper when a test is awkward, and reports success. The interface looks correct. The real action path may still use the old default. The next agent sees two plausible answers and adds a third patch between them.

With the setup:

  1. The map sends the agent to the existing permission authority.
  2. The wall makes any new layer of policy visible.
  3. The gate proves both the fresh-install behavior and the explicit user override before the commit lands.

The result is not merely fewer lines. It is one answer to the question, one change in the history, and one proof that can be run again.

That is why I use the setup. Agent speed is only useful when the system remains understandable after the agent is gone.

Copy the setup

The generic version lives in structure-template. It includes the coding-principles document, versioned Git hooks, the shared commit-message checker, CI, and a configurable sz.py.

The template is intentionally not ready to use unchanged. A repository has to name its real owners, measure its current size, choose its limits, and decide which checks are fast enough for every commit versus broad enough for every push. The shape transfers. The numbers do not.

That is the lesson from A Universal Kernel: keep the mechanism general and keep the facts that vary as data.

What would measure it

I have not run a controlled comparison proving that this workflow makes an agent produce better software.

The clean test would give the same set of repository tasks to matched agent sessions. One copy would have only the tasks and tests. The other would add the map, wall, and gate. Then I would compare:

  • how often the first review accepts the change;
  • how often a task creates a second source of truth;
  • how much shipping code each accepted task adds;
  • how many regressions appear after the handoff;
  • how long it takes to reach an accepted result.

That would separate a useful control system from ceremony. Until then, this is the setup I use, where it came from, and the failure each piece is meant to prevent.