Most AI content pipelines have a prompt and an output. Source material, if any, is interpolated into the prompt as text and then forgotten.
That works until you need to answer a question the pipeline cannot: is this claim true?
Making evidence an object
In Noomachy DevTracker, every story carries an explicit array:
interface StoryCandidate {
title: string;
summary: string;
evidence: string[];
relatedCommits: string[];
// ...
}
evidence holds verbatim, source traceable facts. Nothing inferred:
Commit: Added CSV export
Commit: Added export button
Files changed (3): src/lib/export/csv.ts, src/components/ExportButton.tsx
Languages: typescript
Scope: 3 commits in penny, +180 -12 lines
At least one change touched user-facing code
It goes into the prompt, as it would anyway. The difference is that it also survives as data, and three things become possible that were not before.
1. Validation
The validator receives both the generated post and the evidence, and can compare them.
Every number in the post must appear in the evidence. 40% faster fails, because no 40 exists in the source. This is the single most effective check in the product and it is only possible because the evidence outlived the prompt.
2. Controlled extension
Some true things are not in the git history. "First paying customer, $4.99" is real and unprovable from commits.
Because evidence is an array, a user-entered milestone can be appended to it:
Milestone stated by the writer: First paying customer: $4.99 (2026-09-15)
The validator now accepts $4.99, because it is in the evidence. And the label is explicit about its provenance: this is the writer's assertion, not something the system inferred.
That distinction is the entire design. There is no code path where Noomachy decides you got a customer.
3. Traceability
A post stores the commit SHAs it came from. If someone asks where a claim came from, there is an answer.
That matters more than it sounds. The failure mode for generated content is not usually a dramatic lie, it is a small claim nobody can source and everyone assumes someone checked.
The ordering trap
One detail that caused a real bug.
Evidence is extended with milestones before generation, in one place, so the prompt and the validator see identical evidence.
The first version attached milestones only to the prompt. Generation legitimately used a milestone, then validation rejected the number because its copy of the evidence did not include it. The pipeline was arguing with itself.
Any system with a generator and a checker needs them reading the same source. Two paths that assemble context independently will diverge, and the divergence surfaces as a confusing intermittent failure rather than an obvious one.
What it costs
Evidence has to be generated deterministically, which means a real extraction layer before any model call: parsing commits, grouping them, scoring them, formatting facts.
That is most of the engineering in this product. The AI part is comparatively small.
Which is the general shape of building something trustworthy on a language model. The model is the last step and the least of the work. Everything before it exists to constrain what the model is allowed to say, and everything after it exists to check that it did.