Todd Watts

Agentic Coding Is a Process Problem, Not a Model Problem

Todd Watts13 min read

Picking a model is the smallest decision in agentic coding. A fractional CTO's guide to alignment, vertical slices, token budgets and guardrails, with a checklist for founders and engineering leads.

Cover illustration: the headline "It's the loop, not the model." beside a bold lime loop connecting three nodes labelled Budget, Gate and Checkpoint around a small chip labelled Agent.
It's the loop, not the model: budget, gate and checkpoint around the agent.Illustration · Todd Watts

The model is the easy part

The first question most teams ask about coding agents is which model to standardize on. It is a fair question, and it is rarely the one that decides whether agentic coding works.

I build AI-heavy finance, tax and B2B apps. A lot of my learning time goes to listening to the people who build agent harnesses talk about what actually holds up. They keep landing in the same place. The teams that get real value from coding agents have answered three unglamorous questions well. How do we agree on what to build? How much context are we willing to pay for? What is the agent allowed to touch?

So this is not a model comparison. It is the operating model I would put in place before I cared which model sits underneath it.

A small model at the center of three larger rings labelled process, budget and guardrails, showing that the model is the smallest part of an agentic coding system.

The model is the smallest decision. Process, budget and guardrails decide whether agent work ships.

1. Alignment is the one step you cannot delegate

The failure that comes up again and again is not bad code. It is a well-written implementation of the wrong thing.

Matt Pocock's full walkthrough of his AI coding workflow (AI Engineer, April 2026) attacks this first. Before any code, he runs a deliberately tiny "grill me" skill: the agent interviews him about the feature until the two of them share the same design concept. He calls misalignment one of the main problems in AI work. Only then does the conversation become a PRD. He admits he tends not to read the generated PRD, because once the grilling has aligned him with the model, reading it mostly checks the model's ability to summarize. In his words, planning "has to be human in the loop." Implementation is the part that can run away from the keyboard.

Nate B Jones describes the other side of this in The AI Failure Mode Nobody Warned You About (January 2026). His example is an agent told to clean up old docs on a laptop. It does exactly what it was told. That's the problem. His point: intent isn't in the text the way context is. And a bigger context window can sometimes make things worse, not better. What he suggests is practical. Have the model ask targeted clarifying questions. Write the intent down as its own document. Put business logic in code, not only in a prompt. And run cheap checks that escalate when uncertainty is high.

For a founder, this becomes very concrete. The planning session is human work. It should feel like a requirements meeting, not a prompt.

Synthetic example. Say a fictional bookkeeping startup tells an agent to "add recurring invoices." Ten minutes of grilling turns up the questions that actually matter. What happens when a customer's card fails? Is tax recalculated every cycle? Who gets to pause a schedule? None of those answers were in the original request, and each one would otherwise have been a guess baked into code.

2. Vertical slices and TDD are the feedback loop

Once the requirements are clear, the next mistake is handing an agent one giant task. Pocock breaks the PRD into small issues. He calls them tracer bullets, or vertical slices, and credits the idea to The Pragmatic Programmer. Every slice goes through every layer (front end, API, database). That way the agent sees the whole flow working early, instead of finishing one layer before starting the next. Issues block each other like cards on a kanban board, so an agent can grab the next unblocked one without asking.

Within a slice, the loop is plain test-driven development. One failing test. Make it pass. Refactor. "TDD I found is absolutely essential for getting the most out of agents," Pocock says, and he teaches red, green, refactor to the agent as a skill. Tests, types and lint tell you a slice is done. The agent's own report does not.

Two more ideas from the same talk are worth stealing:

  • Humans own the interfaces. Pocock leans on John Ousterhout's deep modules from A Philosophy of Software Design: a small, simple interface with a lot of functionality behind it. His advice is to "design the interface for these modules, but then delegate the implementation." That is also a clean line between a senior engineer's job and an agent's.
  • Autonomy comes last. Only after planning does he let implementation run unattended, using Sandcastle, a TypeScript library he built that creates a git worktree, sandboxes it in a Docker container and runs a prompt inside it. The result is a branch you merge later.

Nik Pash, head of AI at Cline, offers a useful counterweight in Hard Won Lessons from Building Effective AI Coding Agents (AI Engineer, December 2025). His headline is "capability beats scaffolding": frontier models bulldoze the clever tool tricks we built for weaker ones. I agree, and it sharpens my point. Scaffolding is not process. What Pash still treats as the art is the verifier, and his example is a teakettle whistle. It is "pure outcome verification." The kettle does not care whether you used gas, induction or a campfire. A good test works the same way.

A human-led grill me interview becomes a PRD and a set of ordered issues, which split into four vertical slices. Each slice passes through the UI, API and database layers and runs its own red, green, refactor test loop. The human designs interfaces and reviews at merge; the agent implements each slice in its own worktree and sandbox.

Tracer-bullet slices cut through every layer, and tests, not the agent, say when a slice is done.

3. Context is a budget, and you are paying for it

Every token an agent reads costs money, and past a point it also costs quality.

The best number I've found on this is in a video from the Playwright team itself, Playwright CLI vs MCP (February 2026). They ran one browser task through a coding agent twice, once each way: open the Playwright docs, search for locators, check that the page exists for each language, and screenshot each one. The MCP run took about 114,000 tokens, because the full accessibility snapshot and the screenshot landed in the model's context. The CLI run took about 26,800 tokens, because the CLI saved its output to files and the coding agent decided what it actually needed to read. Their recap is balanced: for coding and testing inside a coding agent, use the CLI; if you are authoring a general agentic loop, MCP "is still the way to go."

Quality is the second cost. Pocock uses the "smart zone" and "dumb zone" framing, which he credits to Dex Horthy of HumanLayer. Each added token is like adding a team to a football league: the number of matches, or attention relationships, grows quadratically. His current marker for leaving the smart zone is "around 40% or around... 100K," whatever the size of the window. Past that, clear the context rather than push on. That is why his workflow produces durable files, the PRD and the issues, that a fresh context can pick up.

For a team lead, this translates into four habits:

  • Treat token spend like cloud spend. Measure it per task, not per month.
  • Prefer tools that write to disk over tools that pour everything into the conversation.
  • Hand off through files, not through ever-growing chat histories.
  • Size tasks to the smart zone. If a slice cannot be done in a clean context, the slice is too big.

Bar chart showing about 114,000 tokens for the MCP approach versus about 26,800 tokens for the CLI approach on the same browser task, as reported by the Playwright team.

Same browser task, roughly a quarter of the context. Figures as reported by the Playwright team in "Playwright CLI vs MCP"; not independently reproduced.

4. Governance: gates, not trust

This is the part most teams skip. A coding agent on a developer's laptop usually runs with that developer's access.

In A Conversation with Jiquan Ngiam About Agent + MCP Security on Daniel Miessler's Unsupervised Learning (February 2026), Ngiam puts it plainly: coding agents "have the same permissions as the engineer." They can read your files, make changes to production and potentially read your SSH directory for secrets. His approach uses hooks at each point in the agent's life cycle: when a prompt is submitted, before a tool runs and after it runs. The examples are concrete. Flag a secret pasted into a prompt before it reaches the model. Put MCP servers behind one gateway, because on their own they tend to have inconsistent authentication, too many permissions and too many tools. Give each rule one of three settings: allow, block or ask. Both he and Miessler admit to being uneasy about running agents with permission checks skipped. Ngiam's answer is a network sandbox whose only way in or out is the gateway.

This matters even more in finance and tax software. An agent that can read a production .env file or trigger a deploy is a risk to customer data, however good the model behind it is.

The other half of governance is cost-aware judgment. IndyDevDan's 10 Levels of Jev For Agentic Engineers (September 2026) walks through putting a small, fast decision model in front of expensive agents. As presented, it starts with "a smart, cheap, fast if statement," such as checking an API input for prompt injection. It moves through multiple choice and weighted scores to confidence gating, for cases where "a wrong answer costs more than asking a human." Then come "one cheap decision in front of a bunch of expensive things" for routing, and guards on the agent's own tool calls. In the demo, a request to push to origin main is blocked every time it is attempted. I would not build your architecture around a product you saw in a demo video. The pattern is what matters: cheap, typed, logged decisions in front of expensive, open-ended ones.

Mario Zechner, who created the open-source Pi coding agent, adds the vendor angle in Building pi in a World of Slop (AI Engineer, April 2026). He used to work construction: "if my hammer breaks every day, I'm getting really mad," and the same goes for development tools. His deeper complaint about a closed harness was that "my context wasn't my context." The system prompt and tool definitions changed on every release. For a CTO, the lesson is to own the parts of the stack that encode your policy: permissions, hooks, skills and context files. Do not leave them as defaults in someone else's product.

An agent action flows through three gates: a prompt check that scans for secrets, a pre-tool check that allows, asks or blocks, and a post-tool check that logs and verifies. A small, fast judgment model feeds every gate and escalates to an expensive model or a human only when needed. An illustrative log records each decision, including a blocked force push to main.

Hooks at each point in the agent's life cycle, with cheap, typed, logged decisions in front of expensive ones. The log lines are illustrative.

5. A CTO checklist before you pick a model

Before the checklist, here is how those pieces fit together. The short video below walks through the loop around a coding agent in a real, time-compressed Pi session, where the permission gate asks, blocks a push and waits for approval.

IN MOTION
Budget, Gate, Checkpoint: The Loop Around a Coding Agent

A coding agent works inside a loop: plan, request a tool, pass a permission gate, spend from a token budget, run the tests, and stop at a human checkpoint. This explainer walks that loop step by step, then shows a real, time-compressed Pi session in a throwaway repo where the gate asks, blocks a push, and waits for approval. Companion video for the post "Agentic Coding Is a Process Problem, Not a Model Problem."

Open video in a new tab
Read the video transcript

Everyone argues about which model to use. This is about the part you actually control: the loop the model runs inside, and the walls around it. An agent session is a loop. Plan. Request a tool. Pass the gate. Spend from the budget. Run the tests. Then, at the end of a slice, check in with a human. And around again. First, plan. Before touching code, the agent says what it will change, and how it will know it's done. A plan you can read is a plan you can reject, cheaply, before tokens go into code. Second, the request. The model can't act on its own. It can only ask: read this file, run this command, write this change. Every action shows up as a structured request. That's the seam where control lives. Third, the gate. Each request is checked against a written policy with three answers. Allow, for routine things like reading files or running the tests. Deny, for things this session should never do, like pushing code, installing packages, or reaching the network. And ask, for anything that changes code. That one pauses and waits for a person. The details matter. A chained command is only as safe as its least safe part, so the gate judges every piece. Fourth, the budget. Every turn spends tokens, and the context window fills up as the session runs. So the session gets caps. Treat them as a signal, not a penalty. When the meter climbs faster than the work moves, the slice was too big, or the agent is going in circles. The fix is to stop and re-scope, not to buy more room. Fifth, tests. The test is written first, so done has a definition the agent can't talk its way around. Red, then green. Still red means another lap, out of a smaller budget. Sixth, the checkpoint. When the tests pass, the agent stops and shows a human what changed. Approve, and the slice is done. Revise, and the next lap starts with better instructions. Small slices keep this honest. A short diff is easy to really review. Here's a real session, in a throwaway repo, sped up. The task: add a length limit to a tiny slug function. Even the first look around chains in an unlisted command, so the gate asks. The test comes first, and it fails. Red. The fix changes code, so the gate asks again. Green. Then the agent tries to commit and push in one command. Push is on the deny list, so the whole chain is blocked. It doesn't argue. It reports the block, calls the checkpoint, and waits for a human. A little over half the token budget, and nothing got pushed. None of this depends on the model. A stronger model wastes fewer laps. But the policy, the caps, and the checkpoint are yours to design. Write them down before the next session starts.

If you are a founder or engineering lead rolling out coding agents, this is the list I would work through first. Nothing on it cares which model wins next month.

Process

  • No feature starts from a one-line prompt. A human sits through the requirements interview first.
  • The requirements are written to a file, so a fresh agent context can pick them up.
  • Work comes in vertical slices, and you know what blocks what.
  • Every slice runs red, green, refactor. Nothing is "done" until types, tests and lint pass.
  • Senior engineers design the interfaces; agents implement behind them.

Cost

  • Token usage is measured per task and reviewed like any other infrastructure cost.
  • Tools that write artifacts to disk are preferred over tools that flood the context.
  • Tasks fit comfortably in a fresh context, with handoff through files.
  • Cheap models handle narrow, typed decisions; expensive models handle the hard work.

Governance

  • By default, agents don't get a developer's full permissions.
  • Prompts get scanned for secrets before anything goes to a model provider.
  • MCP servers and tools sit behind one gateway or allowlist with consistent authentication.
  • Force pushes, production deploys and data deletion are blocked or need a human to approve them.
  • Your permissions, hooks, skills and context files are in your repo, not just in some vendor's settings screen.
  • If a vendor changes pricing, limits or behavior, you have somewhere else to go.

A three-column CTO checklist covering process, cost and governance, with items such as a human-in-the-loop requirements interview, vertical slices with tests, token usage measured per task, prompts scanned for secrets and destructive actions blocked or approved. The requirements interview is highlighted as the one step that cannot be delegated.

The same checklist as a one-page summary. The text list above is the accessible version.

Get that list right and changing models is a configuration change. Skip it and a better model mostly helps you build the wrong thing faster.

If you want help putting this operating model in place for your team, that is a big part of my fractional CTO work. For a related decision, see build, buy or integrate.

toddwatts.dev: Agentic coding is a process problem, not a model problem.

Sources

Quotes are taken from each video's published captions and lightly trimmed. Dates are YouTube publication dates. The bookkeeping example is fictional and does not describe any client engagement. Figures are as reported by the linked creators and have not been independently reproduced. Diagrams and cover are original work by Todd Watts.

Todd Watts

Software engineer and fractional CTO through Shell Command, LLC. Technical direction, architecture and hands-on development.

← All posts