Idempotency, Deduplication, and Safe Replays

Design idempotent API operations and deduplication records so clients can retry after timeouts without creating duplicate payments, jobs, or updates.

published: reading time: 10 min read author: GeekWorkBench
Quick Summary

Idempotency keys let clients retry after a timeout without repeating a payment, job, or other business effect. The guide explains how to scope keys, compare request fingerprints, reserve work atomically, handle in-progress requests, and retain results for the right retry window. It also covers provider-side references and reconciliation for cases where an external effect may have succeeded before the response was lost.

Idempotency, Deduplication, and Safe Replays

Introduction

A payment request can reach the server and charge the card even if the response never reaches the client. If the client retries after a timeout, a new operation may create a second charge. An idempotency key gives both attempts the same operation identity:

curl -X POST https://api.example.com/payments \
  -H 'Idempotency-Key: order-8421-payment' \
  -H 'Content-Type: application/json' \
  -d '{"orderId":"8421","amount":4200}'

If the response is uncertain, the client sends the same key and request again; the service returns the existing result instead of charging twice. This guide covers how to scope and reserve keys, handle concurrent or in-progress requests, and reconcile effects that cross into an external provider.

Key lifecycle and storage

Treat an idempotency key as the identity of one logical operation, not as a globally unique secret. Scope the lookup to the authenticated caller and operation, then store a canonical request fingerprint alongside the key so a retry with changed input can be rejected.

Reserve the key atomically before starting work. A new reservation moves from in_progress to complete; a matching replay can return the stored result or an in-progress status, while a mismatched fingerprint is a conflict. Keep records for at least the documented retry window, with longer retention for costly or irreversible effects. If the side effect runs in another system, carry a stable provider reference and reconcile uncertain outcomes.

flowchart TD
    A[First request reserves key] --> B[Durable key and result store]
    B --> C[Operation completes and result is saved]
    C --> D[Client retries with same scoped key]
    D --> E[Key lookup finds completed record]
    E --> F[Authorize caller and replay saved result]

The stored result makes a retry a lookup rather than a second business operation. If the original work is still running, the lookup can return its in-progress state; it should not start the effect again.

Capacity planning example

Estimate capacity from the retry window and measured request shape. For one illustrative workload, assume 10 million new idempotent operations per day, a seven-day retention period, and an average durable record of 1.5 KB including the response reference and metadata. This yields about 70 million live keys (10 million × 7 days) and 105 GB of raw record data (70 million × 1.5 KB, using decimal units). If indexes, storage overhead, and replicas together require 2.5 times the raw space, budget roughly 263 GB. Measure the actual record and deployment overhead before provisioning.

At 10 million operations per day, average arrival rate is about 116 operations per second (10 million ÷ 86,400). If a load test shows a peak-to-average ratio of 8, plan for about 930 lookups per second at peak (116 × 8); a further 2× capacity margin would set the initial lookup target near 1,860 per second. This is a starting estimate: include retries, status polling, and hot-key behavior in the load test. The seven-day TTL is appropriate only if it covers the documented client retry window and any longer reconciliation period the business requires.

Implementation sketch

The reservation must be atomic. The business effect and idempotency record may span separate systems, so a transaction alone may not cover both. Use the downstream provider’s idempotency facility where available, persist operation state, and reconcile ambiguous outcomes rather than assuming failure.

async function createPayment(
  key: string,
  input: PaymentInput,
): Promise<PaymentResult> {
  const fingerprint = hashCanonicalJson(input);
  // Resolve by authenticated caller, operation, and key, not key alone.
  const existing = await store.reserveOrRead(key, fingerprint);
  if (existing.kind === "conflict")
    throw new Error("key reused with different request");
  if (existing.kind === "complete") return existing.response;
  if (existing.kind === "in_progress")
    return { status: "processing", id: existing.operationId };
  const result = await provider.charge(input, { idempotencyKey: key });
  await store.complete(key, result);
  return result;
}

When to use and when not to

Use idempotency for payments, order creation, provisioning, job submission, and any operation clients may retry after an ambiguous timeout. Use event IDs for duplicate message or webhook deliveries. It is less useful for naturally repeatable reads, though read endpoints should still avoid surprising side effects. Do not accept keys without limits: validate syntax and length, scope them to a principal, and define expiration and collision behavior.

Strategy Benefit Trade-off
Client supplied key Retries preserve the caller’s operation identity Clients must reuse the same key for a retry
Server generated operation ID Easy status lookup after acceptance Client may not receive the ID if the response is lost
Payload hash deduplication Can detect exact repeated messages Similar payloads may represent separate legitimate actions
Downstream provider key Extends protection across a boundary Provider retention and semantics may differ

Trade-Off Table

Decision Safer default Alternative Trade-off
Key scope Authenticated caller + operation + key Global key namespace Scoped keys reduce collisions and prevent one caller from retrieving another caller’s result; global lookup is simpler but risks cross-operation collisions.
Durable key store Database record with a unique constraint and atomic reservation Cache-only record A durable store survives restarts and coordinates workers; a cache is faster to expire but eviction can allow a duplicate effect.
Request fingerprint Hash canonical input with a stored fingerprint version Hash raw serialized bytes Canonical, versioned hashing handles field order and future canonicalization changes; raw bytes are simpler but equivalent requests can hash differently.
Retention window Cover the documented retry window, extending it for costly effects Short fixed TTL Longer retention uses storage and may complicate cleanup; early expiry can turn a late retry into a second operation.
Retry behavior Return the stored result or an in-progress status for a matching key Start the operation again Reusing the existing operation avoids duplicate work; callers need a documented way to check pending outcomes.
Replay safety Reauthorize the caller and reuse downstream provider references Trust possession of the key Authorization and stable downstream references limit data exposure and duplicate external effects; key-only trust can leak stored responses or repeat side effects.

Production failures and mitigations

Two servers race to reserve the same key; use a database uniqueness constraint and return the established operation. A crash occurs after the payment provider charges but before local completion is stored; retry the provider call with the same provider key or reconcile by stable external reference. A client reuses one key for two separate purchases; reject mismatched fingerprints and document that a new logical action needs a new key. Expiring records too early can make a late retry duplicate an effect; choose retention based on retry windows and irreversible impact.

Observability checklist

  • Count new, replayed, in-progress, conflict, and expired-key requests.
  • Record operation IDs and key hashes, never raw sensitive request bodies.
  • Alert on a rise in in-progress records beyond expected processing time.
  • Track reconciliation outcomes for ambiguous external effects.
  • Preserve correlation IDs through retries and asynchronous work.

Security notes and pitfalls

An idempotency key is not authorization. Authenticate and authorize every replay just as you would a new request; otherwise a leaked key could expose another user’s stored response. Avoid predictable keys and store only hashes if that fits the lookup design. Protect responses containing payment or personal data. A key prevents duplicate execution only when every relevant side effect is covered; partial protection can still create inconsistent records.

Quick Recap Checklist

  • Reuse one idempotency key for retries of the same logical action.
  • Compare a canonical request fingerprint to catch key reuse with different input.
  • Reserve keys atomically and persist the operation outcome.
  • Handle in-progress work and downstream side effects explicitly.

Interview Questions

1. What does idempotent mean for an API operation?
Repeating the same logical operation has the same intended effect as performing it once. The response may differ in timing or metadata, but the system does not create duplicate business effects.
2. Why store a request fingerprint with the key?
It detects accidental key reuse with different input. Returning the first result for a different payload can make a client believe its new request succeeded when it did not.
3. Does an idempotency key create exactly-once delivery?
No. Requests can still be delivered multiple times. The key lets the receiver collapse repeats into one logical effect, and every side-effect boundary must participate for the guarantee to hold.
4. What should a server do when a matching key is already in progress?
Do not start the operation again. Return a documented in-progress response or status resource tied to the existing operation, so the caller can check its outcome without duplicating the side effect.
5. Why scope an idempotency key to a caller or operation?
Scoping prevents unrelated callers or endpoints from colliding on the same key and limits access to a stored result. Authenticate and authorize the replay before returning that result.
6. How should a team choose how long to retain idempotency records?
Cover the documented client retry window and account for how costly or irreversible a duplicate effect would be. Expiring a record too early can make a late retry look like a new operation.
7. Why should the server reauthorize a replay before returning a stored response?
Access may have changed since the first request, and a leaked key is not proof of identity or permission. Apply the current caller and resource policy before exposing the prior result.
8. How can idempotency cover a payment provider call outside the local database?
Pass a stable provider-side idempotency key or external reference when supported, and persist the local operation state. If the outcome is ambiguous, reconcile against that reference instead of issuing a new charge blindly.
9. Why might hashing the payload alone incorrectly deduplicate two legitimate operations?
Two separate logical actions can have identical payloads, such as buying the same item twice. A client-supplied operation key distinguishes an intentional new action from a retry of the earlier one.

Further Reading

Conclusion

A timeout leaves the client guessing. A stable idempotency record lets the service recognize a retry and return the existing outcome. Scope keys to the caller and operation, reserve them atomically, and authorize every replay.

Category

Related Posts

API Clients, Servers, and Network Boundaries Explained

Understand what API clients and servers each own, how network boundaries fail, and how timeouts, retries, and trust boundaries shape reliable integrations.

#api-design #networking #distributed-systems

Partial Failure, Ordering, and Eventual Consistency

Understand partial API failures, message ordering, and eventual consistency, then design status models and recovery paths clients can reason about.

#api-design #distributed-systems #consistency

Retries, Timeouts, Backoff, and Circuit Breakers

Set API deadlines, bounded retries, exponential backoff, and circuit breakers so transient failures recover without multiplying load or hiding outages.

#api-design #resilience #retries