The short version
- I took one real issue from the FlowHunt backlog and had Claude Code implement it 13 times on Opus 5.5, under seven different instruction setups. Every run started from the same commit, in its own clone with its own database, and went through our real GitHub CI. Two blind Opus reviews graded each result.
- The repository’s 33 KB CLAUDE.md did not buy measurable quality. Runs with it got 2.0 blocker and major findings per review. A run with no instruction file at all got the same. A 5.5 KB file got 1.0.
- A separate planning stage made things worse. Both setups that ran a planning prompt and then an implementation prompt cost more than their plain-prompt twins, and they produced four of the five blockers in all 26 reviews.
- The winner was an 8 KB CLAUDE.md, three skills, Sonnet workers, and a plain prompt. Its two runs averaged $6.68 at API list price. The full setup with the pipeline prompts averaged $11.49, about 1.7 times as much, and had more findings.
- What survived from all the structure I had built is the brief: a short, decided description of one slice of work that the main model writes for a cheaper worker. Plans for the main model itself did not pay off. The model plans fine on its own.
This follows the architect pattern post, which measured who should type the code. This one asks how much the main model needs to be told before it starts.
Why I ran it
Our monorepo had collected a lot of agent instructions. The root CLAUDE.md was 33 KB: build commands, test cadence, code style, architecture, critical paths, security rules, permission and credit rules, the two agent surfaces, issue filing, PR conventions. Ten skills sat under .claude/skills/. On top of that, our CI pipeline (we call it harnext) ran an issue through a tagger, a triage step, a planning step that wrote a plan as an issue comment, and an implementation step that read the plan back and wrote the code, each with its own long system prompt.
Every line of that was added for a reason, usually after an agent got something wrong. The question nobody had asked was whether it still helped. A CLAUDE.md is loaded into every session and re-sent on every turn. If the model would have done the right thing anyway, those tokens are a tax, and as what a context window costs explains, a tax on every turn adds up faster than it looks.
The test
The issue was a timed email sequence for new sign-ups: an admin defines steps like “day 0: welcome” and “day 5: your trial ends soon”, and a background task mails each account on the right day, stops once the person buys a paid plan, and skips accounts created before launch. It needs a migration, a repository, a service, a Celery beat task, REST endpoints, MCP tools for agents, and a console UI in 11 languages. That is a realistic week-one feature, and it has several edge cases that are easy to miss.
Each run was a headless Claude Code session on Opus 5.5, sandboxed to its own clone and database, with a script that pushed its commit to a draft PR and fed the CI result back into the session, up to three rounds. Cost is Claude Code’s own total_cost_usd for the main session plus every subagent, priced at API list. I pay a subscription, so this is a measure of work, not of my bill.
The seven setups:
| Setup | Runs | |
|---|---|---|
| A | Full 33 KB CLAUDE.md and ten skills, pipeline prompts (plan, then implement) | 2 |
| B | Full CLAUDE.md and skills, plain prompt | 3 |
| C | No CLAUDE.md, no skills, plain prompt | 2 |
| D | Full CLAUDE.md plus the architect skill, plain prompt | 2 |
| E | 8 KB CLAUDE.md, architect, grain-ponytail and unslop, pipeline prompts | 1 |
| F | 5.5 KB CLAUDE.md, no skills, plain prompt | 1 |
| G | 8 KB CLAUDE.md, architect, grain-ponytail and unslop, plain prompt | 2 |
A to D came first. E to G were built from what they showed, and the third B run went alongside them as a control. The 8 KB file is the 5.5 KB one plus a short list of the skills and when to load them, and the model-routing rules that make the architect skill delegate to Sonnet.
Every result then went to two Opus reviewers that saw only a letter, the diff, the repository and the issue. They listed every change they would ask for before approving, graded blocker, major, minor or nit, and checked each acceptance criterion against a file and line.
Cost
The winning setup, G, cost $5.96 and $7.39. The full setup with pipeline prompts, A, cost $11.12 and $11.85. The gap is mostly turns and what each turn carries. A’s main sessions ran 172 and 144 turns on Opus with 24 to 28 million input tokens. G’s ran 79 and 95 turns with 13 to 16 million, and handed about 60 more turns to Sonnet workers, which bill at a fraction of the rate. G also finished its work in about 21 minutes against A’s 45.
Two things on this chart surprised me. No instructions at all is not the cheap option: its two runs cost $7.32 and $15.58, and the dear one also failed CI and wrote the largest diff of the test. And the 5.5 KB file without skills, F, was not cheaper than the full setup. It wrote more code, 1,601 product lines, and paid for it. The saving in G does not come from the short file alone. It comes from the short file plus the two skills that change how the work is done.
Quality
The cheapest setup also had the fewest serious findings: no blockers, and 1.5 and 0.5 majors per review on its two runs. The full CLAUDE.md, B, sat at 2.0 majors on all three of its runs, the same as no CLAUDE.md at all.
This is the result I care most about. The file has long sections on security, permissions and credits, written because an agent once skipped them. With or without the file, the reviewers found the same number of serious problems, and the same kinds came back across setups: a sign-up who pays and then cancels, an email lost or sent twice when a worker fails partway, an editor that overwrites content it cannot show. Those are reasoning gaps about state over time. No rule list in CLAUDE.md covered them, because you cannot write a rule for an edge case you have not thought of yet.
What did cover them was a change to the architect skill between rounds. Before the main model writes any brief, it now writes a table: each acceptance criterion, what happens when the state it depends on changes over time, and the decided behaviour. Each row then goes into the brief with a named test. G’s runs wrote 774 and 818 test lines, against 510 to 670 for B.
The planning stage
A and E are the two setups that ran our CI pipeline’s prompts: a planning prompt that writes a plan, then an implementation prompt that reads the plan and executes it. Compare each with its twin that got a plain prompt instead. A cost $11.49 a run against B’s $9.61 and had blockers where B had none. E cost $10.13 against G’s $6.68 and had a blocker in each of its reviews where G had none.
I think the reason is that a plan is a document the model writes for itself and then has to read back as an outside instruction. The second session treats the plan as a specification, follows it where it is wrong, and spends turns reconciling it with what it finds in the code. A model this capable plans better inside the session that is about to do the work, where the plan and the code are in the same context.
The part of planning that did survive is the brief. When the main model hands a slice to a Sonnet worker, the worker starts with an empty context and builds exactly what it is told. That handoff needs a written, decided description: the goal, the files, the pattern to mirror, the edge cases, the test to write. A plan for yourself is overhead. A brief for someone else is the only way the work gets across.
What grain-ponytail did
The skill from code with the grain tells the model to mirror the nearest existing example and ship the least code that works. Every run that had it loaded wrote between 1,068 and 1,353 product lines. The two largest diffs, 1,601 and 1,752, came from runs with no skills. One no-skill run stayed at 1,280, so the skill does not decide the size by itself, but it removed the tail.
It is also the clearest case of what kind of instruction still pays. “Mirror the nearest example” changes what the model does on every file it writes. “Use snake_case for functions” restates something it already does.
What the reviews still caught
The winning run was not clean. After the test I rebased G2’s work onto main and put it through our full PR check, which reviews the diff, proves each finding with a failing test, and drives the change in a real browser. It found that the step editor dropped every block except the first HTML one when an admin edited a step created over MCP, and that a failed claim release could mail some users twice. Both were fixed, with tests, before the PR went up for human review.
The blind reviewers also flagged more over-engineering in the architect runs than in the full-instruction ones: 2.5 items per review for G against 0 to 1.0 for the runs of A and B. Fewer serious findings, more nits. I would take that trade, but it is a trade.
What I changed
- The 33 KB CLAUDE.md became a 12 KB AGENTS.md that Codex reads too, and CLAUDE.md now only imports it. It keeps the commands, the test cadence, the architecture and the rules a model cannot infer from the code, such as which paths are critical and that the two agent surfaces must stay in step.
- The planning stage is gone from the pipeline. Triage hands the issue straight to an implementation stage whose prompt shrank to six steps: read the issue, branch, load the architect and grain-ponytail skills, test, commit, open a draft PR.
- The architect skill gained the edge-case table, a Tests field in the brief, and a final pass that checks the whole diff against the issue rather than against the briefs. That version is public at github.com/LivinTribunal/architect-skill .
The repository change is in review at the time of writing.
What to take from this
If your CLAUDE.md grew one incident at a time, it is probably paying for rules the model already follows. Cut it to what the model cannot find out by reading the code, and measure before and after on a real task rather than trusting that more guidance is safer.
Keep the instructions that change how the work is done: who types the code, how a slice is briefed, which existing pattern to copy. Drop the plan-then-execute pipeline if your main model is a current frontier one. Write briefs for workers, not plans for yourself.
Method
13 headless runs of Claude Code 2.1.289 on 5 October 2026, main model Opus 5.5, workers Sonnet 5.5 where the architect skill delegated. All runs started from one commit of the FlowHunt monorepo, each in an APFS clone with its own Postgres database, sandboxed from the other runs and from the experiment’s own files. A broker pushed each run’s commit to a draft PR against a frozen base branch and resumed the session with the CI result, up to three rounds. Cost is the maximum total_cost_usd per session, summed over the main session and its subagents. Product lines are non-blank added lines outside tests, translations and migrations.
Each result got two blind reviews by Opus with one rubric: correctness, acceptance criteria, grain, private fields and hidden state, over-engineering, security and permissions, tests, migrations. Candidates were shuffled under letters, and the reviewers could read the repository but not run anything. Findings per review are averaged over a setup’s runs. With one to three runs per setup, treat a single setup’s mean as noisy. The conclusions above are the ones that held across both rounds.

