Partial Failure, Ordering, and Eventual Consistency
Understand partial API failures, message ordering, and eventual consistency, then design status models and recovery paths clients can reason about.
Distributed API workflows can partially succeed, leaving callers unsure whether to retry, wait, or recover. This guide explains how to expose operation states, handle eventual consistency and per-entity event ordering, and choose between bounded retries and saga compensation. It also shows how version checks and reconciliation help detect stale or missing updates, so teams can build clients and operations that respond safely to delays and failures.
Partial Failure, Ordering, and Eventual Consistency
Introduction
A distributed operation can succeed in one place and fail in another. An order may be saved while inventory reservation times out. A service may publish an event, then lose the connection before returning its HTTP response. These are partial failures: the system has no single moment at which every participant knows the same outcome.
Eventual consistency means replicas or services may temporarily show different values but converge if updates stop and processing succeeds. Patterns such as event sourcing make those changes explicit. It is not a promise that everything will fix itself. APIs need to expose pending states, ordering rules, and recovery actions so callers do not mistake “not visible yet” for “failed.”
Model the outcome explicitly
Return acceptance separately from completion when work continues asynchronously. A response can give the caller a stable operation ID and status URL:
{
"operationId": "op_7f32",
"status": "processing",
"statusUrl": "/operations/op_7f32"
}
Define what each state means. accepted means the service recorded the request; processing means work is underway; succeeded and failed are terminal outcomes; compensating means the system is undoing or offsetting earlier steps. A client can poll the status resource with bounded intervals or receive a notification. Include a version or update time so clients can recognize stale status, and avoid reporting succeeded until the operation’s promised effects are complete.
When to use eventual consistency
It fits workflows where independent services own their data, where read models can lag briefly, or where asynchronous processing protects availability. It is a poor fit for an invariant that must be checked atomically, such as preventing a bank account from spending the same balance twice. Keep strong consistency for the narrow operation that protects the invariant, then distribute resulting events afterward.
| Design choice | Advantage | Trade-off |
|---|---|---|
| Synchronous transaction | Immediate invariant across one database | Couples components and may limit scale |
| Asynchronous workflow | Services recover independently | Pending states and compensation logic |
| Per-entity ordering | Preserves meaningful sequence | Hot entities can become throughput bottlenecks |
| Versioned events | Detects stale or missing updates | Consumers must handle gaps and schema evolution |
Compensation is not rollback
In a workflow spanning several services, a later failure cannot usually undo earlier commits with one database rollback. A saga handles this by issuing compensating actions, such as releasing a reservation after payment fails. Compensation is a new business operation: it can fail, take time, or produce a result that differs from the original state. For example, a refund may reverse a charge, but it does not erase the charge from the ledger.
Choose between retrying the failed step and compensating based on the business rule. If the dependency may recover and the action remains valid, retry with a bounded policy. If the workflow can no longer continue, run an explicit compensation. Make each step and compensation safe to repeat, record its outcome, and expose intermediate states such as compensating or compensation_failed through the operation status API. Provide a manual repair path for cases where compensation itself cannot complete.
Key Takeaways
- Model compensation as a durable, repeatable operation, not an automatic rollback.
- Keep the original and compensating actions visible in the operation history.
- Decide whether to retry, compensate, or require operator repair for each failure point.
Implementation snippet: version check
A consumer can ignore stale events and flag a sequence gap for repair. The update and stored version should be atomic.
async function apply(event: AccountEvent): Promise<void> {
await db.transaction(async (tx) => {
const current = await tx.getVersion(event.accountId);
if (event.version <= current) return; // duplicate or stale
if (event.version !== current + 1)
throw new Error("sequence gap; retry or reconcile");
await tx.apply(event);
await tx.setVersion(event.accountId, event.version);
});
}
Production failures and mitigations
The event is committed but the API response is lost, so the client creates a duplicate workflow. Use an idempotency key. A consumer crashes after updating its database but before acknowledging the message; make the handler idempotent. The projection lags for minutes while clients poll aggressively; expose processing status and use bounded polling or notifications. A poison event blocks a strict-order partition; quarantine it, alert, and define whether later events may proceed. Reconciliation jobs should compare authoritative state with projections and repair gaps.
Observability checklist
- Measure end-to-end workflow duration and time spent in each state.
- Track consumer lag, retry counts, dead letters, and sequence gaps.
- Compare projection freshness with an explicit lag objective.
- Log operation IDs, entity versions, and event IDs across services.
- Alert on stuck nonterminal operations and reconciliation backlog.
Security and Compliance Notes
Authorize status lookups against the same tenant boundary as the original request. Do not expose internal event payloads through a public status endpoint. Validate event producer identity and schema. Keep enough operation and event history to support audit and reconciliation, while applying retention limits and redacting sensitive fields from logs and status responses.
Common Pitfalls / Anti-Patterns
A frequent mistake is labeling a command accepted as completed, or returning stale data without identifying it as a projection. Avoid “exactly once” claims: use idempotent effects, sequence checks, and reconciliation instead. Do not retry a non-idempotent side effect blindly after a timeout; the first attempt may already have succeeded.
Quick Recap Checklist
- State which records are authoritative and which views may lag.
- Expose operation status when completion can take time or remain uncertain.
- Define ordering at the entity or partition level and use versions to detect stale events.
- Make consumers replay-safe and provide reconciliation for inconsistent state.
Interview Questions
Further Reading
- Event sourcing explains how to represent domain changes as an append-only event history.
- API idempotency and safe replays covers duplicate requests and replay-safe side effects.
- Retries, timeouts, backoff, and circuit breakers covers bounded retry behavior when dependencies fail.
- Microsoft Azure Architecture Center: Saga pattern — workflow coordination and compensating transactions across services.
- Microsoft Azure Architecture Center: Event Sourcing — event history, projections, and eventual consistency.
Conclusion
Partial failure is normal once an operation crosses service boundaries. Make uncertainty visible through operation states, define ordering at the entity level, and give operators reconciliation tools. Clients can then handle delay without turning every temporary mismatch into a duplicate command.
Category
Related Posts
Ordering Guarantees in Distributed Messaging
Understand how message brokers provide ordering guarantees, from FIFO queues to causal ordering across partitions, and the trade-offs in distributed systems.
API Clients, Servers, and Network Boundaries Explained
Understand what API clients and servers each own, how network boundaries fail, and how timeouts, retries, and trust boundaries shape reliable integrations.
Idempotency, Deduplication, and Safe Replays
Design idempotent API operations and deduplication records so clients can retry after timeouts without creating duplicate payments, jobs, or updates.