Idempotency, Deduplication, and Safe Replays
Design idempotent API operations and deduplication records so clients can retry after timeouts without creating duplicate payments, jobs, or updates.
Idempotency keys let clients retry after a timeout without repeating a payment, job, or other business effect. The guide explains how to scope keys, compare request fingerprints, reserve work atomically, handle in-progress requests, and retain results for the right retry window. It also covers provider-side references and reconciliation for cases where an external effect may have succeeded before the response was lost.
Idempotency, Deduplication, and Safe Replays
Introduction
A payment request can reach the server and charge the card even if the response never reaches the client. If the client retries after a timeout, a new operation may create a second charge. An idempotency key gives both attempts the same operation identity:
curl -X POST https://api.example.com/payments \
-H 'Idempotency-Key: order-8421-payment' \
-H 'Content-Type: application/json' \
-d '{"orderId":"8421","amount":4200}'
If the response is uncertain, the client sends the same key and request again; the service returns the existing result instead of charging twice. This guide covers how to scope and reserve keys, handle concurrent or in-progress requests, and reconcile effects that cross into an external provider.
Key lifecycle and storage
Treat an idempotency key as the identity of one logical operation, not as a globally unique secret. Scope the lookup to the authenticated caller and operation, then store a canonical request fingerprint alongside the key so a retry with changed input can be rejected.
Reserve the key atomically before starting work. A new reservation moves from in_progress to complete; a matching replay can return the stored result or an in-progress status, while a mismatched fingerprint is a conflict. Keep records for at least the documented retry window, with longer retention for costly or irreversible effects. If the side effect runs in another system, carry a stable provider reference and reconcile uncertain outcomes.
flowchart TD
A[First request reserves key] --> B[Durable key and result store]
B --> C[Operation completes and result is saved]
C --> D[Client retries with same scoped key]
D --> E[Key lookup finds completed record]
E --> F[Authorize caller and replay saved result]
The stored result makes a retry a lookup rather than a second business operation. If the original work is still running, the lookup can return its in-progress state; it should not start the effect again.
Capacity planning example
Estimate capacity from the retry window and measured request shape. For one illustrative workload, assume 10 million new idempotent operations per day, a seven-day retention period, and an average durable record of 1.5 KB including the response reference and metadata. This yields about 70 million live keys (10 million × 7 days) and 105 GB of raw record data (70 million × 1.5 KB, using decimal units). If indexes, storage overhead, and replicas together require 2.5 times the raw space, budget roughly 263 GB. Measure the actual record and deployment overhead before provisioning.
At 10 million operations per day, average arrival rate is about 116 operations per second (10 million ÷ 86,400). If a load test shows a peak-to-average ratio of 8, plan for about 930 lookups per second at peak (116 × 8); a further 2× capacity margin would set the initial lookup target near 1,860 per second. This is a starting estimate: include retries, status polling, and hot-key behavior in the load test. The seven-day TTL is appropriate only if it covers the documented client retry window and any longer reconciliation period the business requires.
Implementation sketch
The reservation must be atomic. The business effect and idempotency record may span separate systems, so a transaction alone may not cover both. Use the downstream provider’s idempotency facility where available, persist operation state, and reconcile ambiguous outcomes rather than assuming failure.
async function createPayment(
key: string,
input: PaymentInput,
): Promise<PaymentResult> {
const fingerprint = hashCanonicalJson(input);
// Resolve by authenticated caller, operation, and key, not key alone.
const existing = await store.reserveOrRead(key, fingerprint);
if (existing.kind === "conflict")
throw new Error("key reused with different request");
if (existing.kind === "complete") return existing.response;
if (existing.kind === "in_progress")
return { status: "processing", id: existing.operationId };
const result = await provider.charge(input, { idempotencyKey: key });
await store.complete(key, result);
return result;
}
When to use and when not to
Use idempotency for payments, order creation, provisioning, job submission, and any operation clients may retry after an ambiguous timeout. Use event IDs for duplicate message or webhook deliveries. It is less useful for naturally repeatable reads, though read endpoints should still avoid surprising side effects. Do not accept keys without limits: validate syntax and length, scope them to a principal, and define expiration and collision behavior.
| Strategy | Benefit | Trade-off |
|---|---|---|
| Client supplied key | Retries preserve the caller’s operation identity | Clients must reuse the same key for a retry |
| Server generated operation ID | Easy status lookup after acceptance | Client may not receive the ID if the response is lost |
| Payload hash deduplication | Can detect exact repeated messages | Similar payloads may represent separate legitimate actions |
| Downstream provider key | Extends protection across a boundary | Provider retention and semantics may differ |
Trade-Off Table
| Decision | Safer default | Alternative | Trade-off |
|---|---|---|---|
| Key scope | Authenticated caller + operation + key | Global key namespace | Scoped keys reduce collisions and prevent one caller from retrieving another caller’s result; global lookup is simpler but risks cross-operation collisions. |
| Durable key store | Database record with a unique constraint and atomic reservation | Cache-only record | A durable store survives restarts and coordinates workers; a cache is faster to expire but eviction can allow a duplicate effect. |
| Request fingerprint | Hash canonical input with a stored fingerprint version | Hash raw serialized bytes | Canonical, versioned hashing handles field order and future canonicalization changes; raw bytes are simpler but equivalent requests can hash differently. |
| Retention window | Cover the documented retry window, extending it for costly effects | Short fixed TTL | Longer retention uses storage and may complicate cleanup; early expiry can turn a late retry into a second operation. |
| Retry behavior | Return the stored result or an in-progress status for a matching key | Start the operation again | Reusing the existing operation avoids duplicate work; callers need a documented way to check pending outcomes. |
| Replay safety | Reauthorize the caller and reuse downstream provider references | Trust possession of the key | Authorization and stable downstream references limit data exposure and duplicate external effects; key-only trust can leak stored responses or repeat side effects. |
Production failures and mitigations
Two servers race to reserve the same key; use a database uniqueness constraint and return the established operation. A crash occurs after the payment provider charges but before local completion is stored; retry the provider call with the same provider key or reconcile by stable external reference. A client reuses one key for two separate purchases; reject mismatched fingerprints and document that a new logical action needs a new key. Expiring records too early can make a late retry duplicate an effect; choose retention based on retry windows and irreversible impact.
Observability checklist
- Count new, replayed, in-progress, conflict, and expired-key requests.
- Record operation IDs and key hashes, never raw sensitive request bodies.
- Alert on a rise in in-progress records beyond expected processing time.
- Track reconciliation outcomes for ambiguous external effects.
- Preserve correlation IDs through retries and asynchronous work.
Security notes and pitfalls
An idempotency key is not authorization. Authenticate and authorize every replay just as you would a new request; otherwise a leaked key could expose another user’s stored response. Avoid predictable keys and store only hashes if that fits the lookup design. Protect responses containing payment or personal data. A key prevents duplicate execution only when every relevant side effect is covered; partial protection can still create inconsistent records.
Quick Recap Checklist
- Reuse one idempotency key for retries of the same logical action.
- Compare a canonical request fingerprint to catch key reuse with different input.
- Reserve keys atomically and persist the operation outcome.
- Handle in-progress work and downstream side effects explicitly.
Interview Questions
Further Reading
- Retries, Timeouts, Backoff, and Circuit Breakers — How to bound the retry policy that uses idempotency.
- API Keys, Sessions, and Service Credentials — How caller identity and credentials shape key scope and replay authorization.
- Sync, Async Messaging, and Webhooks — Apply deduplication to repeated webhook and message deliveries.
- Stripe: Idempotent requests — A production example of key scope, replay behavior, and retention.
- RFC 9110: HTTP Semantics — The intended idempotency semantics of HTTP methods.
Conclusion
A timeout leaves the client guessing. A stable idempotency record lets the service recognize a retry and return the existing outcome. Scope keys to the caller and operation, reserve them atomically, and authorize every replay.
Category
Related Posts
API Clients, Servers, and Network Boundaries Explained
Understand what API clients and servers each own, how network boundaries fail, and how timeouts, retries, and trust boundaries shape reliable integrations.
Partial Failure, Ordering, and Eventual Consistency
Understand partial API failures, message ordering, and eventual consistency, then design status models and recovery paths clients can reason about.
Retries, Timeouts, Backoff, and Circuit Breakers
Set API deadlines, bounded retries, exponential backoff, and circuit breakers so transient failures recover without multiplying load or hiding outages.