Grouping related commits into a single story is the core of turning a git history into content. Four commits about CSV export are one thing that happened, and one post about that thing is better than four posts about its parts.
The instinct, then, is that more grouping is better. It is not, and the point where it stops being better is sharper than you would expect.
Where it breaks
Running our grouper against a real account produced a story containing forty commits across six unrelated subsystems: billing, email, interviews, job tracking, a console UI and an agent runtime.
That is not a broad story. It is not a story at all. There is no sentence that describes it except "this person did a lot of work", which is not interesting and not specific and would produce a post nobody would read.
Compare it to the version after splitting:
Added plans, usage metering, limits and billing abstraction
Files: packages/core/src/saas/billing.ts, plans.ts, usage.ts Scope: +947 -12 lines
One commit. Thirteen files. A specific thing that a person can picture. Far better material than the forty commit blob, despite containing a fraction of the work.
The one idea rule applies to grouping
Good posts carry one idea. A reader should be able to summarise the post in a sentence.
That constraint has to be enforced at the grouping stage, not the writing stage. If you hand a model forty unrelated commits and ask for a focused post, it will produce something vague, because vagueness is the only thing true of all forty.
Story size is therefore not a formatting concern. It is upstream of everything the writer does.
The backstop
Our merge rules are permissive, because the failure of being too strict is worse: it splits the CSV export story, which is the case the feature exists for. So there is a separate check for when permissiveness goes wrong.
export const MAX_STORY_COMMITS = 12;
Past twelve, the cluster gets split. And the way it splits matters more than the threshold.
First, by the author's own scopes. If the commits carry conventional commit scopes, those are an explicit statement of which subsystem each belongs to. Splitting along them uses the author's judgment rather than inventing our own. The forty commit blob had six distinct scopes, and splitting along them produced exactly the six stories a person would have drawn by hand.
Then, by top level directory. When scopes are absent, path structure is the next best proxy.
Finally, cap and stop. If the block is genuinely undifferentiated, keep the strongest twelve. The rest stay visible in the day's activity breakdown, just not inside a story claiming to be about them.
That last distinction is deliberate. Truncating a story is a claim about what the story covers. Truncating the record would be hiding work. Only the first one happens.
The threshold is not the interesting number
Twelve is roughly where a post stops having a spine. It is not derived from anything and we would not defend it to two significant figures.
What matters is that reaching the backstop is treated as a signal, not a routine operation. If a cluster hits twelve, the edge rules merged something they should not have. The split is a correction, and it errs toward splitting because an over split story is mildly repetitive while an over merged one is unusable.
The other direction
Worth being honest about the failure this created.
After the fix, that same account produced stories containing a single commit each, because nearly every commit there carries a unique scope. Six subsystems, six scopes, six one commit stories.
Is that wrong? Arguably not. They genuinely were six separate features shipped in one day, and a one commit story with thirteen files and a specific title is good material.
But it loses something. A story with four commits carries corroboration: the feature, the follow up fix, the test. That shape is more credible than a single commit, because it looks like how work actually happens.
We have not solved that. The current position is that over splitting produces mildly thin stories while over merging produces unusable ones, so the asymmetry is worth accepting until there is more real output to tune against. Tuning grouping thresholds against synthetic data is how you get the forty commit blob in the first place.