Noomachy DevTracker's scheduler runs every fifteen minutes. For each workspace it asks whether the user's local time has just passed their configured hour, and if so it generates the day's content.
That means the same workspace and day pair gets evaluated up to 96 times per day. Almost all of those are no ops. But "almost all" is doing dangerous work in that sentence, because the failure mode is a user opening the app to find four copies of the same post.
That is the fastest way an automated product loses trust. Not a crash, not an error page. Just visible evidence that it does not know what it already did.
The key
Idempotency needs a key that identifies the unit of work. Here it is the workspace plus the local date, held as a document whose ID is the date:
workspaces/{workspaceId}/generations/{localDate}
Using the date as the document ID rather than a field is deliberate. It makes duplication structurally impossible rather than something you check for. There cannot be two documents for 2026-09-18 in the same workspace, because a collection cannot contain two documents with the same ID.
Acquiring it
The naive version reads the document, sees nothing, and writes one. Two schedulers racing on the same tick both read nothing, both write, both generate.
So acquisition is a transaction:
return adminDb().runTransaction(async (tx) => {
const snap = await tx.get(ref);
if (snap.exists) {
const status = snap.get("status");
if (status === "completed" || status === "skipped") {
return { acquired: false, reason: "already_completed" };
}
// ...
}
tx.set(ref, { status: "running", lockExpiresAt: /* ... */ });
return { acquired: true };
});
Only one transaction commits. The other sees the write and backs off.
The part people forget
A lock that is only ever acquired is a lock that eventually wedges.
If a run crashes between acquiring the lock and completing, the document sits at running forever. Every subsequent tick sees an in progress run and declines. That workspace silently stops generating content and nobody finds out until the user complains.
So the lock carries an expiry:
lockExpiresAt: Timestamp.fromMillis(now + LOCK_TTL_MS)
A later run that finds a running lock past its expiry takes it over. Twenty minutes, in our case, which is comfortably longer than a scan can legitimately take and short enough that a crash costs at most one cycle.
The state machine
Four states, and each one means something specific:
| Status | Next tick does |
|---|---|
running, unexpired |
declines |
running, expired |
takes over |
completed |
declines |
skipped |
declines |
failed |
retries |
skipped and completed both decline, but they are different facts. skipped means we looked and there was nothing to do: no commits, no connection, no repositories. completed means we generated. Keeping them apart matters when you are debugging why a user got nothing, because "we found nothing" and "we did not look" require different fixes.
failed retries, because a transient error should not cost someone their day.
Deliberate regeneration
A finished day is never regenerated automatically. That is the whole point.
But users do want to regenerate, usually because they did not like the output. So there is exactly one path that clears the record:
if (body.force) await releaseGenerationLock(workspaceId, localDate);
That runs only when a person explicitly asks. The lock is not bypassed, it is released and then reacquired normally.
Why this also bounds the sweep
There is a second benefit that is easy to miss.
Because re-evaluation is safe, the sweep does not have to process every due workspace in one tick. It handles at most twenty and lets the next tick pick up the rest. No queue, no backlog tracking, no risk of a long tick hitting the request timeout.
That only works because the lock makes a repeated attempt harmless. Idempotency bought us a simpler scheduler as a side effect of buying correctness.