The short version
- For five weeks I ran my coding agent in a split: the main model, Opus and later Fable, reads, decides and reviews, and a worker agent writes the code from a written brief. I call it the architect pattern and it is now a public skill: architect-skill on GitHub , MIT licensed.
- The main model stopped typing. Across 49 architect sessions it edited a source file 6 times in 1,467 prompts. In 275 plain sessions over the same weeks it edited 244 files in 3,081 prompts. That is 0.4 against 7.9 per 100 prompts.
- Total tokens per prompt did not fall. They rose, badly, for two weeks, then came back to where plain sessions sit. What changed is who spent them: the Opus share of worker tokens went from 100% to about 4%, and since 14 September a prompt in an architect session costs 364k weighted tokens on the main model, 205k on Sonnet workers and 80k on Opus workers.
- That split is what a subscription cares about. Claude Max meters its all-models weekly limit at each model’s rate, so a Sonnet token draws it at 40% of an Opus token, and Fable has a separate, tighter weekly limit of its own that no worker touches.
- I also tried routing the workers to DeepSeek through a headless Claude Code. Priced at list, the same traffic was $10 instead of $202 on Sonnet or $505 on Opus. On a subscription that saving does not exist, because the subscription was already paid. I dropped it. If you pay per token, read that section anyway, because the numbers are the point.
This is the third post in a series. The first measured what one setting did to a context window, 17% fewer tokens , and the second explains what a context window costs and why. This one is about the other lever: not how big the context is, but which model is holding it.
What the pattern is
A coding agent session does two kinds of work. It thinks: reads the codebase, works out what to change, decides names and signatures and edge cases, checks the diff. And it types: writes the code, runs the test file, fixes the lint. The thinking is where a frontier model earns its price. The typing is not. Most of it is determined by the decision that came before it, and a smaller model executes a decided plan about as well as a large one.
So the skill splits the session. The main model is the architect. It does everything in the main loop except edit a source file:
- Interrogate the requirement and explore the codebase.
- Make the design decisions: approach, data flow, names, signatures, edge cases. The nearest existing pattern to mirror. The smallest change that fixes the root cause.
- Decompose into slices that own disjoint sets of files.
- Write one implementation brief per slice.
- Review the diff that comes back, run the gates, commit.
The worker is a subagent with a fresh context window. It gets the brief and nothing else, writes the code, runs the test files the brief names, and returns a short structured report. Then its context is thrown away. It never becomes part of the main session’s history, which matters more than it sounds: everything the main session reads is re-sent on every later call, so a worker that reads twelve files and a test log in its own context keeps all of that off the architect’s bill.
The brief is the whole contract. Mine has seven fields and the skill is mostly about filling them well:
## Implementation brief: <slice>
Goal: one sentence, observable outcome
Context: what exploration found: excerpts, the pattern to mirror, line anchors, gotchas
Files: paths to create/modify, exclusive ownership for this slice
Design: the decided approach: data flow, names, signatures, edge cases
Constraints: project rules that apply: style, layers, security
Out of scope: what not to touch, including files owned by other slices
Verify: exact commands and expected outcome
A brief contains decisions, not options. If I am still weighing two approaches I am not ready to delegate.
There is one exception. Delegation has a fixed overhead: writing the brief, worker startup, reading the report, reviewing the diff. A typo fix, a config value, a version bump, deleting a dead key, a one-line fix and its regression test: for those the architect edits directly, because the brief would cost more than the edit. The test is cost, not line count. The moment the edit needs a design judgment or touches a path the repo flags as critical, it goes to a worker.
Why a subscription changes the arithmetic
If you pay per token, the case is simple. Opus 5 lists at $5 per million input tokens and $25 per million output; Sonnet 5 at $2 and $10. Every token of implementation that moves from Opus to Sonnet costs 40% of what it did.
I pay a subscription, and there the case is different and, I think, stronger. Claude Max has two weekly limits, not one: a limit across all models, drawn by every token at that model’s rate, and a separate, tighter limit that only Fable draws. Opus has no limit of its own. That gives the split two effects. Work a Sonnet worker does draws the shared limit at 40% of the rate an Opus worker would, so moving two thirds of a session’s tokens to Sonnet makes the week last longer even if the total token count is unchanged. And when the architect is Fable, every token it does not spend typing is a token left on the limit that runs out first. So even if the architect pattern spent exactly the same number of tokens as a plain session, who spends them would still decide when the week ends.
That “even if” turned out to be doing a lot of work.
What the transcripts say
Claude Code writes every session to a JSONL transcript, and every worker run to a sibling file under the session’s directory. I have all of them since 19 August: 324 interactive sessions with three or more prompts, of which 49 used the pattern, meaning they spawned an implementer agent or sent a brief to one. Tokens are weighted the way the bill is: fresh input at 1, cache writes at 1.25 or 2, cache reads at 0.1. CI runs are excluded.
Read it left to right and it is not a success story for the first two weeks.
In the week of 24 August an architect session cost 1.40M weighted tokens per prompt. Plain sessions the same week cost 0.42M. Three and a half times more, and almost all of the excess was orange: Opus workers. I had given the workers a good model and no rules, and they behaved like a good model with no rules. The median Opus worker run that week was 176 turns. Several ran the full backend test suite, which puts about 100 KB of output into a context that is re-sent on every turn after it. One six-worker session on 14 September, after I had already started tightening things, still burned 245M input tokens over 1,585 turns to produce about six thousand lines of diff. Prompt caching was working at a 91.6% hit rate the whole time. The bill was volume, not pricing.
The worker strip shows where the rules landed. Three of them did most of the work, and I wrote them into the agent definitions rather than into the brief, so they apply whether or not I remember:
- A worker never runs a full test suite. It runs the test files its brief names. The architect runs the gates.
maxTurnsis capped in the agent definition, 140 for Sonnet and 160 for Opus. A slice that needs 250 turns was briefed wrong.- A slice stays under about twelve files, and the brief under about 8 KB. The brief is re-sent every turn too.
By the week of 31 August the median worker run was 57 turns and 0.8M weighted tokens. From 14 September, when I moved the default route to Sonnet, it was 40 turns and 0.6M, and no worker has run a full suite since.
The per-prompt total followed: 711k, 425k, 684k, 483k across the last four weeks, against plain sessions at 505k, 425k, 318k, 431k. Architect sessions now cost about what plain sessions cost. Slightly more in two of the four weeks, slightly less in one. What they buy for that is the next two charts.
The main model stopped typing
This is the number I set out to change and it changed. In plain sessions the main model made 244 Edit or Write calls on source files across 3,081 prompts. In architect sessions it made 6 across 1,467. The 6 are all small-edits exceptions I can account for.
Meanwhile the workers were productive. The 85 worker runs made 2,510 source-file edits. Counting main loop and workers together, an architect session produced 2,516 code edits from 1,467 prompts; plain sessions produced 943 from 3,081. Per code edit, an architect session spent 403k weighted tokens in total and 330k on the Opus-family model. A plain session spent 1,410k and 1,324k. I want to be careful with that comparison, because the two populations are not the same work: a session gets classified as architect by delegating code, so architect sessions are coding-heavy by construction, and plain sessions include research, reviews and debugging that never edit anything. Read it as “the pattern gets a lot of code written per token”, not as a four-times efficiency claim.
The cleaner claim is the model split. Since 14 September, a prompt in an architect session costs 364k weighted tokens in the main loop, 205k in Sonnet workers and 80k in Opus workers. The Opus workers’ share of all worker tokens was 100% in August and about 4% in the week of 21 September. On a Max plan that is the shared weekly limit being drawn at Sonnet’s rate instead of Opus’s for almost all of the implementation work, and none of it touching the Fable limit.
Two things did not improve, and I would rather say so than let the charts imply otherwise. The main session’s own context did not shrink: architect sessions compacted 8.7 times per 100 prompts against 4.7 for plain sessions, because the architect reads a lot before it briefs. And the pattern is slower in wall-clock terms on small tasks, because a worker’s startup and report are minutes you do not pay when you just make the edit. The small-edits exception exists because the first version of the skill did not have one, and I watched myself write a 3 KB brief for a two-line change.
The DeepSeek detour
Once the workers were Sonnet, the obvious next question was whether they had to be a Claude model at all. From 13 to 18 September I pointed the worker at DeepSeek V4.1 Flash through a headless Claude Code with ANTHROPIC_BASE_URL set to a compatible endpoint, first through OpenRouter and then natively. Same brief format, same report format, same architect. 2,732 calls, 431M input tokens of which 88% were cache reads, 1.7M output tokens.
At list prices on 22 September, that traffic costs about $10 on DeepSeek off-peak, $21 at peak, $202 on the Sonnet API and $505 on the Opus API. If I paid per token, this would be the headline of the post and I would have kept it.
I do not pay per token, and that is the whole argument. A Max subscription is $100 a month whether I use 30% of the weekly limit or 100%. Moving worker traffic to DeepSeek did not lower that bill by a cent. It added a second bill of $10 to $21 a week on top of it, to save capacity on a plan whose capacity I had already paid for. I had confused “cheaper than the API” with “cheaper than what I pay”, and they are not the same thing when what you pay is flat. The right comparison for a subscriber is DeepSeek against a second subscription, and there a second Claude plan or a ChatGPT plan with Codex costs about the same as a month of moderate DeepSeek use and comes with its own weekly limit, its own model and no metering to watch.
Quality was the other half. Five backend slices came back clean and were committed as-is. One frontend slice removed an exported symbol and left every consumer of it untouched, so the tree did not compile and a later slice had to clean it up. That is exactly the failure a model that executes a specification has and a model that infers a component’s shape from its siblings does not, and the skill now says so: shared-API-surface changes go to Sonnet, and the worker rule reads “never remove or rename an exported symbol without migrating every consumer in the same run”. DeepSeek Flash is a good executor. The architect pattern needs a worker that notices when the brief undercounted the consumers, and that is the one thing a brief cannot supply.
So the routing table today has two rows. Rote work, backend work against a detailed brief with tests to satisfy, frontend work that follows an existing design grammar, and any change to a shared API surface go to sonnet-implementer. Work whose difficulty is in the reasoning goes to opus-implementer: an async race, a multi-site refactor whose correctness depends on sites the brief cannot enumerate, a backfill migration no single test proves. Route by shape, never by risk. Risk changes how hard the diff gets reviewed and whether a human signs it. It never changes who writes it.
The review loop, which is where the pattern earns or loses its keep
A worker’s report is a claim. The skill’s review loop is short and every line came from being burned:
- Read the report and the actual
git diff. Cross-check the diffstat the worker quotes against the one you see. - Re-run the key verification yourself when the claim matters.
- Found a problem? Send the correction to the same worker. It still has its context and picks up where it stopped. Do not fix it by hand, do not respawn.
- Two failed corrections means the brief was underspecified, not that the model is too small. Decompose, put the missing excerpt in Context, rerun. Do not climb to a bigger model, because it will implement the same misunderstanding more fluently.
- Run the scoped gates before accepting anything: the touched tests, the linter on the touched files, whatever convention script the repo carries.
In the transcripts I sent about one correction per worker run, 52 over 85 runs, and 7 Opus runs hit their 160-turn cap, every one on a slice I would now split. Those are the two numbers I watch. If corrections per run go up, my briefs are getting lazy. If workers hit the cap, my slices are too big.
What to take from this
If you pay per token: put the expensive model on decisions and review, put a cheap one on implementation, and price your worker traffic at list before you choose which cheap one. The four-bar chart is the argument.
If you pay a subscription: the pattern does not cut your token count. It moves the implementation tokens to the model that draws the shared weekly limit slowest, keeps the Fable limit for the work only Fable should do, and stops your most expensive model spending its turns on typing. Do not add a metered API on the side to “save” capacity you have already bought. If you run out of worker capacity, buy a second subscription.
Either way, the rules that made it work were the boring ones. No full test suites in a worker. Cap turns. Small slices, short briefs. Correct the same worker once, then fix the brief.
The skill, the three agent definitions and the guard hook are at github.com/LivinTribunal/architect-skill under MIT. The companion skill that tells the worker how to write code that matches the repo it lands in is grain-skill , and the post about it is code with the grain .
Method
Everything the post describes is in github.com/LivinTribunal/architect-skill : the skill itself, the brief template, the routing table, the three implementer agent definitions with their turn caps and context rules, and the guard hook. Install it and you have the setup the numbers below came from.
Sessions are Claude Code JSONL transcripts from 19 August to 22 September 2026, entrypoint cli, three or more user prompts. A session is classified as architect if it spawned opus-implementer, opus-implementer-xhigh or sonnet-implementer, ran a ds-slice worker, invoked the architect skill, or sent a message containing an implementation brief to a general-purpose agent, which is how the early weeks worked before the named agents existed. Worker runs are the subagent transcripts under each session’s directory. Usage is deduplicated by request id and weighted fresh input 1, five-minute cache write 1.25, one-hour cache write 2, cache read 0.1. “Code edit” is an Edit or Write call on a .py, .ts, .tsx, .js, .sql or similar source file, excluding markdown and anything under ~/.claude. DeepSeek pricing is native list price on 22 September 2026, off-peak and peak, with cache reads at the cache-hit rate and cache writes priced as fresh input; Claude API prices are the published Opus 5 and Sonnet 5 list prices with cache reads at a tenth of input. The week of 21 September is two days of data.

