Postmortem · Bug 03 of 03

Webhook retry poisoning: when idempotency becomes the bug

Bug 03 from our own hardening suite — the recovery path itself was broken.

caught by: supabase/tests/hardening/stripe-webhook-reliability.test.tsfixed in: 000026_stripe_event_ordering.sql

The one-sentence version: a webhook that failed partway through processing left its event marked as already seen — so when Stripe retried, exactly as it is designed to do, our own dedup layer rejected the retry and the event was never applied. The protection mechanism had become the failure mode.

Why we deduplicate at all

Stripe delivers at-least-once: the same event can arrive two, three, ten times. Without deduplication, that means double activations and double state transitions. The standard fix is correct and we had implemented it correctly: a unique constraint on the event id in a stripe_events table, and the handler inserts the event record first — if the insert hits the duplicate constraint, the event was already processed, so return 200 and move on. (You must return 200 either way, or Stripe retries forever.)

The bug

The event record was committed before the work it describes. If processing failed after that commit — a constraint violation on the subscription update, a timeout, anything — the event record survived, because it was its own transaction. Now watch what happens on retry:

the-failure-sequence.txttext
attempt 1:
  insert event record    → committed ✓
  apply state change     → fails ✗
  handler returns error

attempt 2 (Stripe's automatic retry):
  insert event record    → duplicate constraint ✗
  handler: "already seen this event" → return 200

Result: the event is acknowledged and permanently skipped.
The side effects it describes never happened.
Idempotency without a correct failure path isn't idempotency. It's a permanent, self-inflicted message drop.

This is the nastiest class of bug we found, because every component is individually reasonable: the dedup constraint is best practice, returning 200 is required, the failure was transient. Stack them together and the system loses events precisely when something goes wrong — the moment reliability matters most.

How we caught it

The hardening suite's stripe-webhook-reliability test injects a failure partway through processing and then replays the event, asserting the failed event can be re-claimed by the retry and driven to completion. Ours couldn't be — the retry hit the poisoned event record. Recovery was structurally impossible. That single test encoded the question most webhook implementations never get asked: what happens on the second attempt after the first one failed halfway?

The fix: make the event record a state machine

The event row is no longer just a "seen it" marker — it carries the attempt's lifecycle: inserted on first delivery, processed_at marks completion, and processing_started_at acts as a concurrency lock (added by migration 000026_stripe_event_ordering.sql). The retry path now distinguishes the two cases the old handler conflated:

  • Duplicate of a PROCESSED event → acknowledge with 200. Deduplication still works.
  • Duplicate of an UNFINISHED event (previous attempt failed) → do not acknowledge. Atomically re-claim the event and re-process it.
  • Claim fails because another worker holds the lock → acknowledge; that worker owns the retry.
  • Re-claim succeeded but processing failed again → release the lock, return 500, and the next retry can reclaim. Recovery is never permanently blocked.
retry-claim.sqlsql
-- Retry claim: atomic and concurrency-safe
update stripe_events
set processing_started_at = now(), error = null
where stripe_event_id = :event_id
  and processed_at is null            -- never completed
  and processing_started_at is null   -- no worker holds the lock
returning id;

-- 0 rows  → someone else owns this retry → acknowledge.
-- 1 row   → we own it → process, then set processed_at.
-- failure → release the lock → the next retry reclaims.

A failed attempt now leaves the event in a reclaimable state instead of a poisoned one. The hardening suite's webhook-reliability test proves exactly this: duplicate inserts are blocked by the constraint, stale ordering updates are rejected, and a failed event can be re-claimed by its retry and driven to completion. It runs in CI on every push.

The honest summary

This isn't a claim that our app is secure or that these were the only bugs that could exist. It's the record of three real bugs we found in our own codebase by attacking it, the migrations that fixed them, and the executable tests that now guard them — tests that ship with the product, so you can run the same attacks against your own copy before your customers run them against you.

Where does your app stand?

These three bugs lived in an app with RLS enabled, idempotent webhooks, and careful code. A 10-question self-assessment tells you which areas of your architecture deserve the same scrutiny.