Here is a real day of commits:
feat: add CSV export
feat: add export button
fix: handle empty export
test: export service
That is one story. Someone shipped CSV export. It took four commits because that is how shipping works, not because four separate things happened.
A naive tool writes four posts. A slightly better tool writes one post with four bullet points, which is a commit log with extra steps. What you actually want is one post about shipping CSV export, where the four commits are evidence rather than content.
Getting that grouping right turns out to be the hard part.
The signals available
Without reading the code, you have four things per commit:
- The message, which may or may not follow a convention
- The file paths that changed
- The number of lines added and removed
- The time it happened
From those you can infer relatedness several ways. Two commits probably belong together if they share a conventional commit scope, if they touched the same file, if they worked in the same directory, or if their messages share a distinctive word.
We use union-find over those signals. Each commit starts in its own group, and any relatedness signal merges two groups. The four export commits above merge because they share the word "export" and mostly share a directory.
Where it went wrong
The first version merged on any single signal. That felt safe. An over-merged story reads as slightly broad, we reasoned, while an under-merged one produces four posts about one feature.
Then we ran it against a real account with 89 repositories and it produced a single story containing forty commits spanning six unrelated subsystems.
The cause was transitive chaining. Union-find merges groups, not pairs. Commit A shares a directory with commit B. Commit B shares a word with commit C. A and C now sit in the same group despite having nothing to do with each other. Repeat forty times and everything collapses into one blob.
Two specific inputs caused it.
Directories were too coarse. The original code treated both the immediate directory and its parent as shared context. Every path under src/lib therefore yielded src/lib as a signal, and in a monorepo that is true of nearly every commit.
Lockfiles counted as shared files. Almost every commit touches package-lock.json eventually. Treating that as evidence of related work merges everything that ever installed a dependency.
The fix that did not work
The obvious response is to tighten the merge rule. Require a shared directory and a shared keyword rather than either one.
That broke the canonical case. In the export example, add export button lives in src/components while the rest lives in src/lib/export. It shares a keyword but not a directory, so it split off into its own story. The rule that stopped the forty-commit blob also broke the four-commit story that motivated the whole feature.
So the fix went upstream instead. The merge rule stayed permissive. What changed was what counts as a signal:
- Only the immediate directory, never the parent
- A single coarse top-level segment like
srcorappis not a signal at all - Lockfiles and root config are excluded from file comparison
Same rule, honest inputs.
Two things on top
A differing explicit scope now vetoes a circumstantial merge. If one commit says feat(jobs) and another says feat(email), the author has stated these are different work. That outranks a coincidence of wording, which is what was merging job intelligence with career email intelligence on the word "intelligence".
And a cluster that grows past twelve commits gets split, first along the author's own scopes, then by top-level directory. Reaching that backstop means the edge rules merged something they should not have, so it errs toward splitting.
The lesson
The interesting part was not the algorithm. Union-find is a textbook structure and it worked exactly as documented.
The interesting part was that the inputs were wrong in a way that only showed up at scale. Four commits in a test fixture never chain. Forty commits across a real monorepo chain immediately. No synthetic test would have caught it, because writing a synthetic test for transitive chaining requires already knowing that transitive chaining is the problem.
That is an argument for running new heuristics against real data early, before you have built confidence in them.