Picking a model is the smallest decision in agentic coding. A fractional CTO's guide to alignment, vertical slices, token budgets and guardrails, with a checklist for founders and engineering leads.

The model is the easy part
The first question most teams ask about coding agents is which model to standardize on. It is a fair question, and it is rarely the one that decides whether agentic coding works.
I build AI-heavy finance, tax and B2B apps. A lot of my learning time goes to listening to the people who build agent harnesses talk about what actually holds up. They keep landing in the same place. The teams that get real value from coding agents have answered three unglamorous questions well. How do we agree on what to build? How much context are we willing to pay for? What is the agent allowed to touch?
So this is not a model comparison. It is the operating model I would put in place before I cared which model sits underneath it.

The model is the smallest decision. Process, budget and guardrails decide whether agent work ships.
1. Alignment is the one step you cannot delegate
The failure that comes up again and again is not bad code. It is a well-written implementation of the wrong thing.
Matt Pocock's full walkthrough of his AI coding workflow (AI Engineer, April 2026) attacks this first. Before any code, he runs a deliberately tiny "grill me" skill: the agent interviews him about the feature until the two of them share the same design concept. He calls misalignment one of the main problems in AI work. Only then does the conversation become a PRD. He admits he tends not to read the generated PRD, because once the grilling has aligned him with the model, reading it mostly checks the model's ability to summarize. In his words, planning "has to be human in the loop." Implementation is the part that can run away from the keyboard.
Nate B Jones describes the other side of this in The AI Failure Mode Nobody Warned You About (January 2026). His example is an agent told to clean up old docs on a laptop. It does exactly what it was told. That's the problem. His point: intent isn't in the text the way context is. And a bigger context window can sometimes make things worse, not better. What he suggests is practical. Have the model ask targeted clarifying questions. Write the intent down as its own document. Put business logic in code, not only in a prompt. And run cheap checks that escalate when uncertainty is high.
For a founder, this becomes very concrete. The planning session is human work. It should feel like a requirements meeting, not a prompt.
Synthetic example. Say a fictional bookkeeping startup tells an agent to "add recurring invoices." Ten minutes of grilling turns up the questions that actually matter. What happens when a customer's card fails? Is tax recalculated every cycle? Who gets to pause a schedule? None of those answers were in the original request, and each one would otherwise have been a guess baked into code.
2. Vertical slices and TDD are the feedback loop
Once the requirements are clear, the next mistake is handing an agent one giant task. Pocock breaks the PRD into small issues. He calls them tracer bullets, or vertical slices, and credits the idea to The Pragmatic Programmer. Every slice goes through every layer (front end, API, database). That way the agent sees the whole flow working early, instead of finishing one layer before starting the next. Issues block each other like cards on a kanban board, so an agent can grab the next unblocked one without asking.
Within a slice, the loop is plain test-driven development. One failing test. Make it pass. Refactor. "TDD I found is absolutely essential for getting the most out of agents," Pocock says, and he teaches red, green, refactor to the agent as a skill. Tests, types and lint tell you a slice is done. The agent's own report does not.
Two more ideas from the same talk are worth stealing:
- Humans own the interfaces. Pocock leans on John Ousterhout's deep modules from A Philosophy of Software Design: a small, simple interface with a lot of functionality behind it. His advice is to "design the interface for these modules, but then delegate the implementation." That is also a clean line between a senior engineer's job and an agent's.
- Autonomy comes last. Only after planning does he let implementation run unattended, using Sandcastle, a TypeScript library he built that creates a git worktree, sandboxes it in a Docker container and runs a prompt inside it. The result is a branch you merge later.
Nik Pash, head of AI at Cline, offers a useful counterweight in Hard Won Lessons from Building Effective AI Coding Agents (AI Engineer, December 2025). His headline is "capability beats scaffolding": frontier models bulldoze the clever tool tricks we built for weaker ones. I agree, and it sharpens my point. Scaffolding is not process. What Pash still treats as the art is the verifier, and his example is a teakettle whistle. It is "pure outcome verification." The kettle does not care whether you used gas, induction or a campfire. A good test works the same way.

Tracer-bullet slices cut through every layer, and tests, not the agent, say when a slice is done.
3. Context is a budget, and you are paying for it
Every token an agent reads costs money, and past a point it also costs quality.
The best number I've found on this is in a video from the Playwright team itself, Playwright CLI vs MCP (February 2026). They ran one browser task through a coding agent twice, once each way: open the Playwright docs, search for locators, check that the page exists for each language, and screenshot each one. The MCP run took about 114,000 tokens, because the full accessibility snapshot and the screenshot landed in the model's context. The CLI run took about 26,800 tokens, because the CLI saved its output to files and the coding agent decided what it actually needed to read. Their recap is balanced: for coding and testing inside a coding agent, use the CLI; if you are authoring a general agentic loop, MCP "is still the way to go."
Quality is the second cost. Pocock uses the "smart zone" and "dumb zone" framing, which he credits to Dex Horthy of HumanLayer. Each added token is like adding a team to a football league: the number of matches, or attention relationships, grows quadratically. His current marker for leaving the smart zone is "around 40% or around... 100K," whatever the size of the window. Past that, clear the context rather than push on. That is why his workflow produces durable files, the PRD and the issues, that a fresh context can pick up.
For a team lead, this translates into four habits:
- Treat token spend like cloud spend. Measure it per task, not per month.
- Prefer tools that write to disk over tools that pour everything into the conversation.
- Hand off through files, not through ever-growing chat histories.
- Size tasks to the smart zone. If a slice cannot be done in a clean context, the slice is too big.

Same browser task, roughly a quarter of the context. Figures as reported by the Playwright team in "Playwright CLI vs MCP"; not independently reproduced.
4. Governance: gates, not trust
This is the part most teams skip. A coding agent on a developer's laptop usually runs with that developer's access.
In A Conversation with Jiquan Ngiam About Agent + MCP Security on Daniel Miessler's Unsupervised Learning (February 2026), Ngiam puts it plainly: coding agents "have the same permissions as the engineer." They can read your files, make changes to production and potentially read your SSH directory for secrets. His approach uses hooks at each point in the agent's life cycle: when a prompt is submitted, before a tool runs and after it runs. The examples are concrete. Flag a secret pasted into a prompt before it reaches the model. Put MCP servers behind one gateway, because on their own they tend to have inconsistent authentication, too many permissions and too many tools. Give each rule one of three settings: allow, block or ask. Both he and Miessler admit to being uneasy about running agents with permission checks skipped. Ngiam's answer is a network sandbox whose only way in or out is the gateway.
This matters even more in finance and tax software. An agent that can read a production .env file or trigger a deploy is a risk to customer data, however good the model behind it is.
The other half of governance is cost-aware judgment. IndyDevDan's 10 Levels of Jev For Agentic Engineers (September 2026) walks through putting a small, fast decision model in front of expensive agents. As presented, it starts with "a smart, cheap, fast if statement," such as checking an API input for prompt injection. It moves through multiple choice and weighted scores to confidence gating, for cases where "a wrong answer costs more than asking a human." Then come "one cheap decision in front of a bunch of expensive things" for routing, and guards on the agent's own tool calls. In the demo, a request to push to origin main is blocked every time it is attempted. I would not build your architecture around a product you saw in a demo video. The pattern is what matters: cheap, typed, logged decisions in front of expensive, open-ended ones.
Mario Zechner, who created the open-source Pi coding agent, adds the vendor angle in Building pi in a World of Slop (AI Engineer, April 2026). He used to work construction: "if my hammer breaks every day, I'm getting really mad," and the same goes for development tools. His deeper complaint about a closed harness was that "my context wasn't my context." The system prompt and tool definitions changed on every release. For a CTO, the lesson is to own the parts of the stack that encode your policy: permissions, hooks, skills and context files. Do not leave them as defaults in someone else's product.

Hooks at each point in the agent's life cycle, with cheap, typed, logged decisions in front of expensive ones. The log lines are illustrative.
5. A CTO checklist before you pick a model
Before the checklist, here is how those pieces fit together. The short video below walks through the loop around a coding agent in a real, time-compressed Pi session, where the permission gate asks, blocks a push and waits for approval.
A coding agent works inside a loop: plan, request a tool, pass a permission gate, spend from a token budget, run the tests, and stop at a human checkpoint. This explainer walks that loop step by step, then shows a real, time-compressed Pi session in a throwaway repo where the gate asks, blocks a push, and waits for approval. Companion video for the post "Agentic Coding Is a Process Problem, Not a Model Problem."
Open video in a new tabRead the video transcript
Everyone argues about which model to use. This is about the part you actually control: the loop the model runs inside, and the walls around it. An agent session is a loop. Plan. Request a tool. Pass the gate. Spend from the budget. Run the tests. Then, at the end of a slice, check in with a human. And around again. First, plan. Before touching code, the agent says what it will change, and how it will know it's done. A plan you can read is a plan you can reject, cheaply, before tokens go into code. Second, the request. The model can't act on its own. It can only ask: read this file, run this command, write this change. Every action shows up as a structured request. That's the seam where control lives. Third, the gate. Each request is checked against a written policy with three answers. Allow, for routine things like reading files or running the tests. Deny, for things this session should never do, like pushing code, installing packages, or reaching the network. And ask, for anything that changes code. That one pauses and waits for a person. The details matter. A chained command is only as safe as its least safe part, so the gate judges every piece. Fourth, the budget. Every turn spends tokens, and the context window fills up as the session runs. So the session gets caps. Treat them as a signal, not a penalty. When the meter climbs faster than the work moves, the slice was too big, or the agent is going in circles. The fix is to stop and re-scope, not to buy more room. Fifth, tests. The test is written first, so done has a definition the agent can't talk its way around. Red, then green. Still red means another lap, out of a smaller budget. Sixth, the checkpoint. When the tests pass, the agent stops and shows a human what changed. Approve, and the slice is done. Revise, and the next lap starts with better instructions. Small slices keep this honest. A short diff is easy to really review. Here's a real session, in a throwaway repo, sped up. The task: add a length limit to a tiny slug function. Even the first look around chains in an unlisted command, so the gate asks. The test comes first, and it fails. Red. The fix changes code, so the gate asks again. Green. Then the agent tries to commit and push in one command. Push is on the deny list, so the whole chain is blocked. It doesn't argue. It reports the block, calls the checkpoint, and waits for a human. A little over half the token budget, and nothing got pushed. None of this depends on the model. A stronger model wastes fewer laps. But the policy, the caps, and the checkpoint are yours to design. Write them down before the next session starts.
If you are a founder or engineering lead rolling out coding agents, this is the list I would work through first. Nothing on it cares which model wins next month.
Process
- No feature starts from a one-line prompt. A human sits through the requirements interview first.
- The requirements are written to a file, so a fresh agent context can pick them up.
- Work comes in vertical slices, and you know what blocks what.
- Every slice runs red, green, refactor. Nothing is "done" until types, tests and lint pass.
- Senior engineers design the interfaces; agents implement behind them.
Cost
- Token usage is measured per task and reviewed like any other infrastructure cost.
- Tools that write artifacts to disk are preferred over tools that flood the context.
- Tasks fit comfortably in a fresh context, with handoff through files.
- Cheap models handle narrow, typed decisions; expensive models handle the hard work.
Governance
- By default, agents don't get a developer's full permissions.
- Prompts get scanned for secrets before anything goes to a model provider.
- MCP servers and tools sit behind one gateway or allowlist with consistent authentication.
- Force pushes, production deploys and data deletion are blocked or need a human to approve them.
- Your permissions, hooks, skills and context files are in your repo, not just in some vendor's settings screen.
- If a vendor changes pricing, limits or behavior, you have somewhere else to go.

The same checklist as a one-page summary. The text list above is the accessible version.
Get that list right and changing models is a configuration change. Skip it and a better model mostly helps you build the wrong thing faster.
If you want help putting this operating model in place for your team, that is a big part of my fractional CTO work. For a related decision, see build, buy or integrate.
toddwatts.dev: Agentic coding is a process problem, not a model problem.
Sources
- Matt Pocock, "Full Walkthrough: Workflow for AI Coding", AI Engineer, April 24, 2026.
- Nate B Jones, "The AI Failure Mode Nobody Warned You About (And how to prevent it from happening)", AI News & Strategy Daily, January 2, 2026.
- Nik Pash (Cline), "Hard Won Lessons from Building Effective AI Coding Agents", AI Engineer, December 12, 2025.
- Playwright team, "Playwright CLI vs MCP: a new tool for your coding agent", Playwright, February 6, 2026.
- Jiquan Ngiam with Daniel Miessler, "A Conversation with Jiquan Ngiam About Agent + MCP Security", Unsupervised Learning, February 5, 2026.
- IndyDevDan, "10 Levels of Jev For Agentic Engineers", IndyDevDan, September 28, 2026.
- Mario Zechner, "Building pi in a World of Slop", AI Engineer, April 16, 2026.
Quotes are taken from each video's published captions and lightly trimmed. Dates are YouTube publication dates. The bookkeeping example is fictional and does not describe any client engagement. Figures are as reported by the linked creators and have not been independently reproduced. Diagrams and cover are original work by Todd Watts.