Engineering3 min read

Every AI feature needs a path that works without the AI

Providers go down, models refuse, output fails validation. A feature that only works when all three go right is not a feature, it is a demo.

If your product's core action depends on a language model, you have four failure modes that a normal API does not:

  1. The provider is down or rate limiting you
  2. The model refuses the request outright
  3. The output does not match your schema
  4. The output parses fine and is wrong

Any one of those, at three in the morning, and your scheduled job produces nothing. The user opens the app and finds an empty screen with no explanation.

The shape

The pipeline reads like this:

story -> AI -> validate -> retry once with the violations
                              |
                              v
                     deterministic fallback

Nothing unvalidated is ever returned. When the AI path cannot produce a clean post, the deterministic path does, and the caller can see which happened.

What the fallback actually is

The important property is that it cannot fabricate. Not that it is unlikely to, that it structurally cannot.

Every line it emits is assembled from the story's own evidence array:

function commitLines(story: StoryCandidate): string[] {
  return story.evidence
    .filter((item) => item.startsWith("Commit: "))
    .map((item) => item.slice("Commit: ".length));
}

The title comes from the story. The bullets come from commit lines that were already humanised and sanitised upstream. The stats line comes from the scope evidence. There is no step at which new content enters.

There is a test that asserts this directly: for every line in the fallback's output, that line must appear in the source material. Not "looks reasonable". Appears.

Degrading is a product decision

The fallback produces plainer posts. Shorter, more literal, less of a story and more of a summary.

That is the right trade. The alternative to a plain post is no post, and a user who opens the app to find nothing has no idea whether the product is broken, their GitHub is disconnected, or they genuinely did not commit anything.

But it has to be visible. A post written by the fallback carries a -fallback suffix in its generationVersion and renders a small "basic" badge in the interface. Silently serving weaker output and letting people conclude the product is mediocre is worse than the empty screen.

Typed failures

The provider layer maps everything onto a small set of kinds, each with copy safe to show a user:

Kind Retryable Message
not_configured no AI generation is not configured yet.
refusal no The model declined to write about this activity.
invalid_output no Generation produced an unusable result. Your activity is safe.
rate_limit yes Generation is busy right now. We will retry.
overloaded yes As above.
network yes Could not reach the service. We will retry.

"Your activity is safe" is doing deliberate work. The scan is persisted before generation runs, precisely so that a generation failure does not lose the day's data. Saying so removes the worst interpretation the user might otherwise reach for.

The refusal case

One subtlety specific to current models: a declined request is not an error.

It returns HTTP 200 with a stop_reason of refusal and content that is empty or partial. Code that reads response.content[0] without checking stop_reason first throws on a refusal, and the stack trace will point at the content access rather than the cause.

Every path checks the stop reason before touching content, and a refusal retries once on a configured fallback model before giving up.

Where the fallback comes from

The satisfying part is that the deterministic path was not extra work.

The original version of this product was entirely deterministic: regex parsing of commit messages, a past tense verb table, an emoji categoriser, a language detector. When we added AI, those functions were not deleted. They were extracted, hardened, and given test coverage they never had.

They now do three jobs. They pre process commits before anything reaches the model, which cuts token cost and gives it cleaner input. They validate output afterwards, since the same sanitiser catches leaked secrets in generated text. And they are the fallback.

If you are adding AI to something that already works, resist deleting the thing that already works. It is your floor.

aiengineeringreliability

Keep reading