There is an appealing pattern for making AI output trustworthy: generate with one model, check with another. Add a judge. Score the output. It looks like defence in depth.
It mostly is not, and the reason is worth being precise about.
Shared failure modes
A language model asked to verify whether a post contains fabricated claims is doing the same kind of task as the model that wrote it. Both are producing plausible text conditioned on context.
If a model is inclined to generate "a 40% performance improvement" because posts like that contain numbers, it is also inclined to find that number plausible when asked to check it. The checker does not have independent access to the truth. It has the same training distribution and the same instinct for what sounds right.
Independence is the entire value of a check. A second opinion from something that thinks the same way is not a second opinion.
What deterministic checking looks like
Noomachy DevTracker's validator uses no model at all. It is a pure function over the generated post and the story's evidence.
The strongest check is numeric. Every story carries an evidence array of verbatim facts from the commit record. Extract every number from the post, extract every number from the evidence, flag anything in the first set that is not in the second.
export function unsupportedNumbers(body: string, story: StoryCandidate): string[] {
const allowed = evidenceNumbers(story);
const found = new Set<string>();
for (const match of body.matchAll(METRIC_NUMBER)) {
// ...
if (!allowed.has(value)) found.add(unit ? `${raw}${unit}` : raw);
}
return [...found];
}
A regex comparing digits has no opinion about whether 40% sounds reasonable. It only knows whether 40 appears in the source material. That indifference is the feature.
What else it checks
Numbers are the sharpest signal but not the only one:
Claims with no possible source. Some assertions can never be supported by a commit scan regardless of phrasing: user feedback, customer testimony, revenue, popularity, audience size, launch status, superiority. Those are pattern matched and blocked outright, because there is no version of "our users love this" that a git history justifies.
Leaked secrets. The post is run through the same sanitiser used on commit messages. If it produces a redaction marker, the raw text contained a credential, an email address, an IP, or an absolute path. The sanitiser covers GitHub tokens, fine grained PATs, AI provider keys, AWS keys, Google API keys, Slack tokens, JWTs and connection strings.
Platform limits. Character counts, per thread part, hashtags included.
Editorial rules. Banned marketing vocabulary, engagement bait, em dashes, emoji budget.
Placeholders. An unfilled {{project}} reaching a published post.
Blocking versus warning
Not everything should stop a post.
Blocking violations prevent persistence entirely: fabricated numbers, unsupported claims, leaked secrets, over length, malformed URLs, unfilled placeholders. These are wrong in a way the user cannot consent to.
Warnings are surfaced and not enforced: banned phrases, emoji overuse, hashtag counts, em dashes. These are style, and the user gets to disagree with us.
That split matters. A validator that blocks everything gets disabled. A validator that blocks nothing is decorative. The line is roughly whether the violation makes the post false or merely makes it worse.
The model still declares its claims
There is one place the model participates in its own checking, and it is structural rather than judgmental.
Every post schema requires a claims array: every factual assertion the post makes, one per entry. The model has to enumerate what it said.
This is not the model marking its own homework. It is the model producing a structured artifact that a deterministic checker can then use, and it makes the fabrication check tractable in cases where parsing prose would not be.
Retry, then degrade
When validation fails, the pipeline retries once with the violations appended to the prompt:
Your previous attempt was rejected for these reasons: [...]. Rewrite it so none of those apply.
If it fails again, generation falls through to a deterministic generator that assembles posts purely from the story's own evidence lines. That path cannot fabricate by construction, and there is a test asserting exactly that: every line of its output must be traceable to the source material.
Posts written that way are labelled in the UI. The user can see that a post came from the fallback rather than silently receiving a weaker one and wondering why.