Partial Failure, Ordering, and Eventual Consistency

Understand partial API failures, message ordering, and eventual consistency, then design status models and recovery paths clients can reason about.

published: reading time: 7 min read author: GeekWorkBench
Quick Summary

Distributed API workflows can partially succeed, leaving callers unsure whether to retry, wait, or recover. This guide explains how to expose operation states, handle eventual consistency and per-entity event ordering, and choose between bounded retries and saga compensation. It also shows how version checks and reconciliation help detect stale or missing updates, so teams can build clients and operations that respond safely to delays and failures.

Partial Failure, Ordering, and Eventual Consistency

Introduction

A distributed operation can succeed in one place and fail in another. An order may be saved while inventory reservation times out. A service may publish an event, then lose the connection before returning its HTTP response. These are partial failures: the system has no single moment at which every participant knows the same outcome.

Eventual consistency means replicas or services may temporarily show different values but converge if updates stop and processing succeeds. Patterns such as event sourcing make those changes explicit. It is not a promise that everything will fix itself. APIs need to expose pending states, ordering rules, and recovery actions so callers do not mistake “not visible yet” for “failed.”

Model the outcome explicitly

Return acceptance separately from completion when work continues asynchronously. A response can give the caller a stable operation ID and status URL:

{
  "operationId": "op_7f32",
  "status": "processing",
  "statusUrl": "/operations/op_7f32"
}

Define what each state means. accepted means the service recorded the request; processing means work is underway; succeeded and failed are terminal outcomes; compensating means the system is undoing or offsetting earlier steps. A client can poll the status resource with bounded intervals or receive a notification. Include a version or update time so clients can recognize stale status, and avoid reporting succeeded until the operation’s promised effects are complete.

When to use eventual consistency

It fits workflows where independent services own their data, where read models can lag briefly, or where asynchronous processing protects availability. It is a poor fit for an invariant that must be checked atomically, such as preventing a bank account from spending the same balance twice. Keep strong consistency for the narrow operation that protects the invariant, then distribute resulting events afterward.

Design choice Advantage Trade-off
Synchronous transaction Immediate invariant across one database Couples components and may limit scale
Asynchronous workflow Services recover independently Pending states and compensation logic
Per-entity ordering Preserves meaningful sequence Hot entities can become throughput bottlenecks
Versioned events Detects stale or missing updates Consumers must handle gaps and schema evolution

Compensation is not rollback

In a workflow spanning several services, a later failure cannot usually undo earlier commits with one database rollback. A saga handles this by issuing compensating actions, such as releasing a reservation after payment fails. Compensation is a new business operation: it can fail, take time, or produce a result that differs from the original state. For example, a refund may reverse a charge, but it does not erase the charge from the ledger.

Choose between retrying the failed step and compensating based on the business rule. If the dependency may recover and the action remains valid, retry with a bounded policy. If the workflow can no longer continue, run an explicit compensation. Make each step and compensation safe to repeat, record its outcome, and expose intermediate states such as compensating or compensation_failed through the operation status API. Provide a manual repair path for cases where compensation itself cannot complete.

Key Takeaways

  • Model compensation as a durable, repeatable operation, not an automatic rollback.
  • Keep the original and compensating actions visible in the operation history.
  • Decide whether to retry, compensate, or require operator repair for each failure point.

Implementation snippet: version check

A consumer can ignore stale events and flag a sequence gap for repair. The update and stored version should be atomic.

async function apply(event: AccountEvent): Promise<void> {
  await db.transaction(async (tx) => {
    const current = await tx.getVersion(event.accountId);
    if (event.version <= current) return; // duplicate or stale
    if (event.version !== current + 1)
      throw new Error("sequence gap; retry or reconcile");
    await tx.apply(event);
    await tx.setVersion(event.accountId, event.version);
  });
}

Production failures and mitigations

The event is committed but the API response is lost, so the client creates a duplicate workflow. Use an idempotency key. A consumer crashes after updating its database but before acknowledging the message; make the handler idempotent. The projection lags for minutes while clients poll aggressively; expose processing status and use bounded polling or notifications. A poison event blocks a strict-order partition; quarantine it, alert, and define whether later events may proceed. Reconciliation jobs should compare authoritative state with projections and repair gaps.

Observability checklist

  • Measure end-to-end workflow duration and time spent in each state.
  • Track consumer lag, retry counts, dead letters, and sequence gaps.
  • Compare projection freshness with an explicit lag objective.
  • Log operation IDs, entity versions, and event IDs across services.
  • Alert on stuck nonterminal operations and reconciliation backlog.

Security and Compliance Notes

Authorize status lookups against the same tenant boundary as the original request. Do not expose internal event payloads through a public status endpoint. Validate event producer identity and schema. Keep enough operation and event history to support audit and reconciliation, while applying retention limits and redacting sensitive fields from logs and status responses.

Common Pitfalls / Anti-Patterns

A frequent mistake is labeling a command accepted as completed, or returning stale data without identifying it as a projection. Avoid “exactly once” claims: use idempotent effects, sequence checks, and reconciliation instead. Do not retry a non-idempotent side effect blindly after a timeout; the first attempt may already have succeeded.

Quick Recap Checklist

  • State which records are authoritative and which views may lag.
  • Expose operation status when completion can take time or remain uncertain.
  • Define ordering at the entity or partition level and use versions to detect stale events.
  • Make consumers replay-safe and provide reconciliation for inconsistent state.

Interview Questions

1. What is a partial failure?
It is an outcome where some components completed work while another component or the caller did not observe completion. Distributed systems cannot assume a timeout means that nothing happened.
2. How can a consumer preserve event order?
Partition or serialize events by the entity whose order matters, attach a sequence number, and reject or buffer events that arrive with a gap. Global ordering is usually expensive and rarely needed.
3. When is eventual consistency inappropriate?
When a decision must enforce an immediate invariant, such as preventing duplicate spending or overselling a scarce item. Keep that invariant within a strongly consistent boundary, then publish its result asynchronously.
4. How does a saga differ from a database transaction?
A database transaction can roll back uncommitted changes within its boundary. A saga coordinates committed steps across services and uses retries or compensating actions when a later step fails.
5. Why is compensation not the same as rollback?
Compensation is a new business action that may fail or have its own visible effects. A refund can offset a charge, for example, but the ledger still records both entries.
6. When should a workflow retry instead of compensate?
Retry when the dependency may recover and the operation is still valid, using a bounded policy. Compensate when the workflow cannot continue and the business rule calls for reversing earlier effects.
7. What should an API expose while compensation is running?
Return an operation status that distinguishes the original failure from compensation in progress or compensation failure. Include a stable operation ID so clients and operators can check the same workflow.

Further Reading

Conclusion

Partial failure is normal once an operation crosses service boundaries. Make uncertainty visible through operation states, define ordering at the entity level, and give operators reconciliation tools. Clients can then handle delay without turning every temporary mismatch into a duplicate command.

Category

Related Posts

Ordering Guarantees in Distributed Messaging

Understand how message brokers provide ordering guarantees, from FIFO queues to causal ordering across partitions, and the trade-offs in distributed systems.

#distributed-systems #messaging #kafka

API Clients, Servers, and Network Boundaries Explained

Understand what API clients and servers each own, how network boundaries fail, and how timeouts, retries, and trust boundaries shape reliable integrations.

#api-design #networking #distributed-systems

Idempotency, Deduplication, and Safe Replays

Design idempotent API operations and deduplication records so clients can retry after timeouts without creating duplicate payments, jobs, or updates.

#api-design #idempotency #retries