Give a language model a handful of commits and ask it to write a LinkedIn post. There is a good chance you get something like this:
Shipped a 40% performance improvement to our export pipeline this week.
Nobody said anything about 40%. The model produced it because posts like that contain numbers, and it is completing a pattern. It is not lying in any meaningful sense. It is doing exactly what it was trained to do.
For a product whose entire claim is that posts describe work that actually happened, this is the failure that matters most. One invented statistic published under someone's name costs more trust than fifty good posts build.
Prompting is necessary and not sufficient
The first line of defence is telling the model not to. Our system prompt includes an explicit list:
Never invent: metrics, percentages, timings or benchmark numbers; users, customers, testimonials or user feedback; revenue, growth, signups or any business result.
This helps. It does not solve the problem, because a prompt is a request and a request can be declined. Anything that matters has to be checked, not asked for.
The check that works
The strongest deterministic signal available is numeric.
Every story carries an evidence array: verbatim facts pulled from the commit record. Commit messages, changed file paths, line counts, languages. Nothing inferred.
So: extract every number from the generated post, extract every number from the evidence, and flag any number in the post that does not appear in the evidence.
40% faster fails, because no 40 appears anywhere in the commits. 2,500 lines fails. 2ms per row fails. Meanwhile 4 commits passes if there were four commits, and +180 lines passes if that is the real diff size.
It is a crude check and it is extremely effective, because fabricated metrics are almost always numbers that have no source.
The details that make it usable
A naive version of this is unusable, because ordinary prose contains numbers.
Single digits are ignored. "There were 3 things worth doing here" is not a metric claim. Flagging it produces noise that trains people to ignore the checker.
Unless they carry a metric unit. "Cut it to 2ms per row" is a claim even though 2 is a single digit. The unit is what makes it an assertion rather than a count.
Years are exempt. "We picked this design back in 2024" is not a fabricated statistic.
Thousands separators are normalised. The evidence might say 1240 while the post says 1,240. Those are the same number and treating them as different would produce a false positive on a true statement.
What numbers cannot catch
Some fabrications have no digits at all:
Users love the new export feature.
No number, entirely invented. A commit scan can never evidence user sentiment.
So alongside the numeric check there is a small set of patterns for assertions that commit activity structurally cannot support: user feedback, customer testimony, business metrics, popularity, audience size, launch status, superiority claims. These are blocked regardless of phrasing, because there is no version of "our users are delighted" that a git history justifies.
The escape hatch
This creates a real problem. The best content genuinely does contain numbers. "First paying customer, $4.99" is a far better post than "we got some revenue". If the checker blocks every number not in the commits, the strongest material becomes unreachable.
The answer is to let the user supply it. A milestone is a fact a person types: a label, an optional value, a date. It attaches to the story as evidence, explicitly marked as the writer's own assertion, and the validator then accepts numbers drawn from it.
The distinction that matters is that the system never infers a milestone. It exists because someone entered it. The tool is not deciding that you got your first customer, you are telling it that you did.
No AI checking AI
One deliberate constraint: the validator uses no language model at all. It is pure, deterministic, and unit tested.
Using a model to check another model's output feels appealing and fails in a specific way: the checker shares the failure modes of the thing it checks. If a model is inclined to produce plausible-sounding numbers, it is also inclined to find them plausible.
A regex that compares digits has no opinion. That is the entire point.