The short version
- AI coding agents are good at adding code and bad at noticing that the code already exists. Left alone they give a repo a second helper for every first one, and every reader has to check both to trust either.
- The fix is two habits written down so an agent will follow them. Grain: find the nearest existing example of the same kind of thing and mirror its shape, place and names. Ponytail: climb a ladder from “does this need to exist” to “the minimum that works” and stop at the first rung that holds.
- Both are now a public skill, MIT licensed: grain-skill on GitHub . It has the authoring rules, a review pass with fixed tags, and the shape of the closing line that makes the agent name what it mirrored.
- In our main repo since August: 386 issues and pull requests carry a plan comment naming the exemplar to mirror, across 1,333 commits into a tree of roughly 18,000 files. 35 places in the code carry a
ponytail:comment naming a deliberate simplification and its upgrade path. - The uncomfortable finding: the agent invoked the skill as a skill zero times. The rules worked because I also put them in the project’s CLAUDE.md, in the worker agent definitions and in the review checklist. Put the rule where the agent already reads.
I wrote about this once before and I did not like the result, because it described the principle and not the mechanism. This is the version with the mechanism.
The failure this is for
Three examples from our backend, all real, all from agent-written pull requests that passed their tests.
A new REST route that pulled the application service directly from a factory. Every sibling route in the repo goes route to controller to service to repository, each hop injected. This one collapsed the chain to route to service. It worked. It also meant the one place where request validation and response shaping live for that feature did not exist, and the next agent to touch the feature had two shapes to choose from.
An oversized service split into _payloads.py, _protocol.py and _tokens.py, each a module of bare async defs, with the service calling protocol.discover(...) by import. The file-length check went green. The thing dependency injection exists to prevent came straight back: a collaborator chosen at import time, unsubstitutable without patching the module, absent from the constructor that documents what the service depends on.
A second normalizer. There was already one. The agent did not find it, wrote another eight lines that did nearly the same thing, and now two functions with different edge-case behaviour both claim to canonicalize the same field.
None of these is a bug in the sense a test catches. Each is an agent choosing the code it could generate over the code that existed. Do that across a few hundred pull requests and the repo stops having a shape.
Grain
New code should be indistinguishable from the code around it. The skill turns that sentence into five steps an agent takes before writing.
Find the exemplar first. For an endpoint, service, repository, query hook, store slice or component, locate the closest existing one and copy its shape, file location and naming. Do not invent structure when a pattern exists. In our repo the branding slice is the named exemplar for the backend because it is small, clean and spans every layer, and the skill’s table says, for each kind of artifact, which file to open and what to mirror.
Respect the layers. Presentation calls application calls domain, and infrastructure implements domain. A new import across those lines fails the architecture check. Read the layers doc, do not re-derive it.
Local grain against repo grain. Match your siblings. But if the siblings themselves diverge from the canonical chain, follow the canonical chain and note the divergence rather than copying the drift. “It matches the other five routes in this file” is not enough if all six skip a layer the canonical slice uses. You compared too locally.
Reuse the canonical helper. If a helper, type or pattern already lives in the repo, use it. Two normalizers force every reader to diff both.
Names carry meaning. Reuse the vocabulary already in the code. A new name for an existing concept is a bug waiting for a reader.
The test at the end is the one I give reviewers: if you can tell which lines the agent wrote purely from their shape, it fought the grain.
Ponytail
Lazy means efficient, not careless. The best code is the code never written. Understand the problem fully first, every file the change touches, the real end-to-end flow, and then climb this ladder and stop at the first rung that holds.
Two rules ride along with the ladder.
A bug fix is a root cause, not a symptom. A ticket names a symptom. Grep every caller of the function you are about to touch. One guard in the shared function is a smaller diff than a guard in every caller, and patching only the path the ticket names leaves the sibling callers broken.
Keep the diff minimal. The diff is what review reads. No drive-by refactors, no renames, no reformatting, no “while I’m here” improvements, no touching files the task did not require. Each one buries the real change. If the diff is bigger than the problem, cut it back. Improvements you spotted are a follow-up issue, not extra commits on this branch.
And the limits, because “lazy” is the word most likely to be misread. Never simplify away input validation at a trust boundary, error handling that prevents data loss, security, accessibility basics, or anything the user asked for. Never be lazy about understanding: the ladder shortens the solution, never the reading. A small diff in the wrong place is a second bug. Non-trivial logic leaves one runnable check behind, the smallest thing that fails if the logic breaks.
When the agent takes a rung on purpose and knows the ceiling, it says so in the code:
# ponytail: global lock, per-account locks if throughput matters
There are 35 of those in our repo now. Each is a decision a future reader would otherwise have to reconstruct, and a marker for where to look when the ceiling is hit.
Making an agent actually do it
A principle in a document is not a behaviour. Three mechanisms made these two into behaviours, and I would rank them in this order.
The brief carries the exemplar. In the architect pattern the main model writes a brief for a worker, and the brief’s Design field names the nearest existing pattern to mirror and the minimal root-cause change. The worker does not have to find the grain. It is handed it. Across our repo that is what the 491 plan comments on 386 issues and pull requests are: a planning step that names what to mirror before anyone writes.
The agent closes with what it mirrored. Every implementer’s report ends with one line, mirrored: <path>, or names the new pattern and why no sibling fit. This is cheap and it is the single best review aid I have added. A mirrored: line pointing at the wrong kind of file tells you the diff is wrong before you open it. A missing one tells you the agent did not look.
Review enforces the same bar, in a fixed order. Our review skill has two passes with fixed tags, and the tags are what make findings comparable across reviewers and across agents.
Structure first. Does this diverge from the nearest exemplar, duplicate an existing helper, cross a layer, sit in the wrong place, reach into another module’s internals, use inconsistent names, take ownership of something it should inject, or break the canonical call chain? Then cuts. What could be deleted, replaced by the standard library, replaced by a native feature, dropped as speculative, or shrunk?
The order is a rule, not a preference. Consistency beats cuts. A cut that leaves the code diverging from its siblings is not a win, because the reader now has to learn two shapes to save eight lines. Run structure first and the cuts pass only touches code that already matches.
What happened in the repo, including the part that did not work
Since 1 August our main repo has taken 1,333 commits, most of them agent-written under review. 386 of the issues and pull requests in that period carry a plan comment that names the exemplar to mirror, 491 such comments in total. The canonical chain is the default shape of a new endpoint, and the two legacy routes that skip the controller are listed in the doc as a divergence being retired, not a second style.
The 35 ponytail: comments are a smaller number than I expected and I think that is right. Most simplifications are not worth a comment. The ones that are, are the ones with a ceiling.
Now the part I would rather not report. I shipped grain and ponytail as a Claude Code skill: a description in the skill index, a body that loads on demand, and a line in CLAUDE.md saying to invoke it before writing code. Across five weeks of transcripts the agent invoked it by name zero times. The mirrored: closing line appeared in seven responses.
The rules still landed. They landed because the same text was also in three places the agent reads without choosing to: the project’s CLAUDE.md, which is in context from the first turn; the worker agent definitions, which say “mirror the nearest existing example’s shape, location and naming” and “reach for the laziest solution that actually works” as rules, not as a skill to consider; and the review checklist, whose tags fail a pull request that ignores them. The skill body became the canonical reference those three point at. As a trigger, it did nothing.
So the lesson I would pass on to anyone packaging conventions for an agent is this. A skill is a good place for the long version of a rule. It is a bad place for the rule itself, because the agent has to decide to load it and it mostly will not. Put the one-paragraph version in the file that is always in context, put the imperative version in the definition of the agent that writes the code, and put the check in the review that gates the merge. Then point all three at the skill for the detail.
The public version at github.com/LivinTribunal/grain-skill is laid out that way: a short SKILL.md with the rules an agent can hold in its head, the full doc with the ladder and the anti-patterns, the review tags, and a README that says which paragraph goes in your CLAUDE.md and which goes in your implementer’s definition. Replace the worked example with your own canonical exemplar and it is done.

