Developers write commit messages constantly and think about them rarely. Which is a shame, because a git history is the most honest record of what someone actually did, and it is sitting there already written.
The question is how much you can extract without reading the code.
The conventional commit is a gift
If a message looks like this:
feat(auth): add password reset
you get three things for free. The type says this is a new capability rather than a fix. The scope says which part of the system. The description says what, in the author's own words.
That parses cleanly into a readable sentence:
Added password reset (auth)
We keep a table of about fifty verbs mapped to past tense, plus fallbacks per commit type. add becomes Added. set becomes Set up. support becomes Added support for. When the first word is not a known verb, the type supplies one: fix: implies Fixed.
It is not sophisticated. It is a lookup table and a regex. But it turns a commit log into something a person can read, which is most of the job.
The scope is the most underused field
Most people who use conventional commits treat the scope as decoration. It is the single most useful signal in the whole message.
A scope is the author explicitly stating which part of the system this belongs to. When you are trying to group related commits into a coherent story, that is worth more than any inference you can make from file paths.
It works in both directions. Two commits sharing a scope almost certainly belong together. Two commits with different scopes almost certainly do not, even when their wording overlaps. We found a real case where feat(jobs): job intelligence and feat(email): career email intelligence were being merged because they share the word "intelligence". The differing scopes are the author saying they are different work, and that should outrank a coincidence of vocabulary.
When there is no convention
Plenty of history looks like this:
fixed the thing
more work on export
wip
asdf
You can still get somewhere. Keyword inference handles a lot: fix, bug, crash suggest a fix; add, new, implement suggest a feature; refactor, cleanup, tidy suggest a refactor.
Order matters more than you would expect. Our first version classified add coverage for the parser as a feature, because it checked for add before checking for coverage. And fix typo in comment came out as a bug fix rather than a style change, for the same reason. Specific markers have to be tested before generic verbs.
Everything else in the commit
The message is not the only signal.
File paths say more than the message often does. A change under src/components is probably user facing. A change under .github/workflows is probably not. A change touching only package-lock.json is probably not worth writing about.
We deliberately kept full paths rather than basenames. src/lib/auth/session.ts tells you considerably more than session.ts, and the directory is one of the strongest signals for grouping related commits.
Line counts are weak evidence of substance, and we cap their influence deliberately. A 900 line diff that is entirely a regenerated lockfile means nothing. A 12 line diff can be the most important change of the week.
The author matters because bots exist. Dependabot and Renovate produce real commits that are real maintenance, and nobody wants a post about them.
The trap in chore(deps)
Here is a real bug worth knowing about.
Dependabot's default commit format is:
chore(deps): bump next from 15.3.2 to 15.3.3
Our classifier mapped the type to a category. chore became a chore, which carries a modest penalty. But the thing that makes this a dependency bump is the scope, not the type.
The result was that the single most common automated commit format in the entire JavaScript ecosystem was escaping the noise filter and competing for attention with real work.
The fix is one rule: when the scope is deps, dependencies, deps-dev or dev-deps, the scope wins over the type.
What to do about noise
The temptation is to delete low signal commits before doing anything else. That is wrong, for a reason that took us a while to see.
We score rather than filter. A dependency bump gets a penalty and stays visible in the activity breakdown. A merge commit gets a large penalty and stays. A typo fix gets a penalty and stays.
The reason is that low scoring commits are still evidence. A story about shipping CSV export is more credible when it includes the follow up fix and the test, even though neither of those is interesting alone. If you filter them out before grouping, you strip stories of the corroboration that makes them believable.
We learned this the hard way: an early version applied the noise threshold before grouping, and produced stories containing only the headline commit with all the supporting work discarded. The threshold now applies to the finished story, not to membership.