Distributed Tracing: Trace Context and OpenTelemetry

Master distributed tracing for microservices. Learn trace context propagation, OpenTelemetry instrumentation, and how to debug request flows across services.

published: reading time: 46 min read author: GeekWorkBench updated: June 17, 2026
Quick Summary

Distributed tracing connects spans across services so a request can be followed through HTTP calls and message consumers. This guide explains OpenTelemetry instrumentation, W3C context propagation, sampling, trace storage, and the signals that reveal problems in the tracing pipeline. It also shows how to investigate missing spans, correlate traces with logs and metrics, and limit sensitive telemetry. Use the examples to trace a boundary, choose what to retain, and keep incident evidence useful without treating traces as an audit log.

Distributed Tracing: Trace Context, OpenTelemetry, and Correlation

Introduction

Suppose an order service calls a payment service, but the payment span appears as a separate trace. The missing piece is usually trace context at the HTTP boundary. With OpenTelemetry JavaScript, supported HTTP instrumentation normally injects and extracts this context for you; use the API directly when the transport or library needs manual propagation.

This TypeScript sketch assumes a configured OpenTelemetry SDK and context manager. It shows the API shape for a plain header carrier; adapt carrier creation to your HTTP framework and installed propagator.

import { context, propagation, trace } from "@opentelemetry/api";

const tracer = trace.getTracer("orders");
type TraceCarrier = Record<string, string>;

// Sender: inject the active span context into the outgoing headers.
async function callPayment(url: string): Promise<Response> {
  return tracer.startActiveSpan("orders.call-payment", async (span) => {
    try {
      const headers: TraceCarrier = {};
      propagation.inject(context.active(), headers);
      return await fetch(url, { headers });
    } finally {
      span.end();
    }
  });
}

// Receiver: extract headers, start a child span, then run work in its context.
async function handlePayment(headers: TraceCarrier): Promise<void> {
  const parentContext = propagation.extract(context.active(), headers);
  const span = tracer.startSpan("payments.handle", undefined, parentContext);
  const spanContext = trace.setSpan(parentContext, span);

  try {
    await context.with(spanContext, async () => {
      await authorizePayment();
    });
  } finally {
    span.end();
  }
}

The shared trace ID connects the two services, while the payment span’s parent ID points to the caller span. The rest of this guide covers span design, sampling, storage, and diagnosing broken propagation.

Core Concepts

Traces and Spans

A trace represents an entire request journey. It contains one or more spans, where each span represents a single operation within that trace.

sequenceDiagram
    participant C as Client
    participant A as API Gateway
    participant O as Order Service
    participant P as Payment Service
    participant N as Notification Service

    C->>A: GET /orders/123
    A->>O: GetOrder(123)
    O->>P: ProcessPayment(order)
    P-->>O: Payment confirmed
    O->>N: SendConfirmation(order)
    N-->>O: Notification sent
    O-->>A: Order details
    A-->>C: Response

Each span captures:

  • Operation name
  • Start and end time
  • Parent span ID (linking)
  • Attributes (key-value metadata)
  • Events (timestamped points within the span)

Trace Context

Trace context propagates across service boundaries through HTTP headers. When service A calls service B, it passes trace context in headers. Service B creates a child span using that context, ensuring the spans stay connected in a single trace.

The W3C Trace Context specification standardizes these headers:

traceparent: 00-0af7651916cd43dd8448eb211c80319c-b7ad6b7169203331-01
tracestate: congo=t61rcWkgMzE

The traceparent header contains:

  • Version (00)
  • Trace ID (32 hex characters)
  • Parent ID (16 hex characters)
  • Flags (01 = sampled)

OpenTelemetry Architecture

OpenTelemetry (OTel) is the open standard for observability. It gives you APIs, SDKs, and instrumentation for collecting traces, metrics, and logs.

graph TB
    subgraph "Application Code"
        A[Your Service]
        B[OTel SDK]
        C[Language-specific auto-instrumentation]
    end

    subgraph "Exporters"
        D[OTLP Exporter]
        E[Jaeger Exporter]
        F[Zipkin Exporter]
    end

    subgraph "Collecting Infrastructure"
        G[OTel Collector]
        H[Jaeger]
        I[Zipkin]
    end

    A --> B
    B --> C
    B --> D
    D --> G
    G --> H
    G --> I

OTel collector

The OTel collector receives, processes, and exports telemetry data. Think of it as middleware between your application and your observability backend.

# otel-collector-config.yaml
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  batch:
    timeout: 5s
    send_batch_size: 1024
  memory_limiter:
    check_interval: 1s
    limit_percentage: 90

exporters:
  otlp:
    endpoint: jaeger-collector:4317
    tls:
      insecure: false
      cert_file: /certs/cert.pem
      key_file: /certs/key.pem

  prometheus:
    endpoint: "0.0.0.0:8889"

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [otlp]
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [prometheus]

Manual Instrumentation

Auto-instrumentation covers many frameworks automatically, but you need manual instrumentation for business-specific operations and custom spans.

Starting Traces

When you need to trace operations that auto-instrumentation cannot reach, you create spans manually. You get a tracer instance and call startSpan. The tracer carries your service name and version, and these become default attributes on every span it creates. Every manual span needs a clear operation name that describes the work being tracked.

Attributes attach business context to each span. They turn into queryable fields in your tracing backend, so you can filter traces by customer ID, order amount, or any other dimension that matters to your domain. Keep attribute values small and typed. Strings and numbers work best.

Always wrap startSpan with a try/catch/finally block. The finally clause guarantees span.end() runs even when an error occurs. Without it, orphaned spans leak memory and produce misleading timing data. On error paths, call recordException and set the span status to ERROR so your tracing backend flags the trace for investigation.

import { trace, SpanStatusCode } from "@opentelemetry/api";

const tracer = trace.getTracer("order-service", "1.0.0");

async function createOrder(orderData: OrderData): Promise<Order> {
  const span = tracer.startSpan("OrderService.createOrder", {
    attributes: {
      "order.customer_id": orderData.customerId,
      "order.item_count": orderData.items.length,
      "order.total": orderData.total,
    },
  });

  try {
    const order = await db.orders.create(orderData);
    span.setStatus({ code: SpanStatusCode.OK });
    return order;
  } catch (error) {
    span.recordException(error as Error);
    span.setStatus({
      code: SpanStatusCode.ERROR,
      message: (error as Error).message,
    });
    throw error;
  } finally {
    span.end();
  }
}

Creating Child Spans

OpenTelemetry JS creates a child span from the active span in the current context. Use startActiveSpan to make a span active while its callback runs; nested spans then inherit its trace and parent IDs automatically. This also keeps parentage intact across awaited operations when a context manager is configured.

async function processPayment(payment: Payment): Promise<PaymentResult> {
  return tracer.startActiveSpan("PaymentService.process", async (span) => {
    try {
      const result = await tracer.startActiveSpan(
        "PaymentService.verifyCard",
        async (verificationSpan) => {
          try {
            await api.verifyCard(payment.card);
          } catch (error) {
            verificationSpan.recordException(error as Error);
            verificationSpan.setStatus({ code: SpanStatusCode.ERROR });
            throw error;
          } finally {
            verificationSpan.end();
          }
        },
      );

      const charged = await chargeCard(payment);
      span.setAttribute("payment.transaction_id", charged.transactionId);
      span.setStatus({ code: SpanStatusCode.OK });
      return charged;
    } catch (error) {
      span.recordException(error as Error);
      span.setStatus({
        code: SpanStatusCode.ERROR,
        message: (error as Error).message,
      });
      throw error;
    } finally {
      span.end();
    }
  });
}

Context Propagation

Proper context propagation connects spans across service boundaries. Without it, spans become orphaned and useless for debugging.

HTTP Propagation

Middleware that extracts incoming trace context and propagates it to downstream calls:

import { trace, context, propagation } from "@opentelemetry/api";

function httpMiddleware(req, res, next) {
  // Extract context from incoming headers
  const extractedContext = propagation.extract(context.active(), req.headers);

  // Run the rest of the request handler within that context
  context.with(extractedContext, () => {
    // All spans created here are linked to the incoming trace
    next();
  });
}

// When making outgoing requests, inject context into headers
async function callDownstreamService(url: string, data: any): Promise<any> {
  const headers = {};
  propagation.inject(context.active(), headers);

  return fetch(url, {
    method: "POST",
    headers: {
      ...headers,
      "Content-Type": "application/json",
    },
    body: JSON.stringify(data),
  });
}

Messaging Propagation

Propagate context through message queues so spans stay connected even with asynchronous processing:

import { trace, context, propagation } from "@opentelemetry/api";

// Producer: inject context into message
async function sendOrderCreatedEvent(order: Order): Promise<void> {
  const headers: Record<string, string> = {};
  propagation.inject(context.active(), headers);

  await kafka.send({
    topic: "order.created",
    messages: [
      {
        key: order.id,
        value: JSON.stringify(order),
        headers: headers,
      },
    ],
  });
}

// Consumer: extract context from message and create linked span
async function handleOrderCreated(message: KafkaMessage): Promise<void> {
  const extractedContext = propagation.extract(
    context.active(),
    message.headers,
  );

  await context.with(extractedContext, async () => {
    const span = tracer.startSpan("OrderConsumer.handleOrderCreated");
    try {
      const order = JSON.parse(message.value.toString());
      await processOrder(order);
      span.setStatus({ code: SpanStatusCode.OK });
    } catch (error) {
      span.recordException(error as Error);
      span.setStatus({ code: SpanStatusCode.ERROR });
      throw error;
    } finally {
      span.end();
    }
  });
}

Diagnosing Broken Propagation

When a trace stops at a service boundary, compare the context on both sides before changing sampling. Check the incoming traceparent, the outgoing request or message headers, and the span IDs recorded by each service. A new trace ID usually means extraction failed or the downstream client started a root span; a matching trace ID with a missing parent span points to dropped or unsampled data instead.

Symptom Check first Likely next step
A downstream service starts a new trace ID Compare inbound and outbound traceparent values; inspect middleware and client instrumentation Confirm extraction runs before the handler and injection runs inside the active context
HTTP spans connect, but queue consumer spans do not Inspect producer message headers and consumer extraction, including header casing and encoding Add propagation to the queue adapter and test a round trip with a known trace ID
Child spans are missing after an await or callback Inspect the runtime context manager and any custom async boundary Use the SDK’s context-aware APIs or explicitly bind the active context
Trace appears only in some services Compare each service’s sampler configuration and parent-based behavior Align sampler settings so downstream services honor the upstream sampling decision

For asynchronous work that outlives the request, do not keep a request span open just to preserve a tree. End the request span and use a span link when the later work is causally related but is not a synchronous child operation.

Adding Business Context

Rich span attributes turn traces from timing diagrams into debugging tools.

Semantic Attributes

Use standard attribute names for common data:

// HTTP attributes
span.setAttribute("http.method", "POST");
span.setAttribute("http.url", "https://api.example.com/orders");
span.setAttribute("http.status_code", 201);
span.setAttribute("http.response_content_length", 1024);

// Database attributes
span.setAttribute("db.system", "postgresql");
span.setAttribute("db.name", "orders_db");
span.setAttribute("db.statement", "SELECT * FROM orders WHERE id = $1");
span.setAttribute("db.operation", "SELECT");

// Messaging attributes
span.setAttribute("messaging.system", "kafka");
span.setAttribute("messaging.destination", "order.created");
span.setAttribute("messaging.operation", "publish");

Custom Business Attributes

Add domain-specific context:

span.setAttribute("order.id", order.id);
span.setAttribute("order.status", order.status);
span.setAttribute("order.customer_tier", customer.tier);
span.setAttribute("order.is_first_purchase", customer.orderCount === 0);

These attributes let you filter traces by business properties: find all traces for premium customers, or analyze timing for first-time purchasers.

Correlation with Logs and Metrics

Traces work best when linked to your logs and metrics.

Trace-Log Correlation

Include trace ID in logs:

import { trace, span } from "@opentelemetry/api";

function logInfo(message: string, data?: Record<string, unknown>): void {
  const span = trace.getActiveSpan();
  const traceId = span?.spanContext().traceId;

  const logEntry = {
    timestamp: new Date().toISOString(),
    level: "INFO",
    message,
    traceId,
    ...data,
  };

  console.log(JSON.stringify(logEntry));
}

// Now every log entry includes the trace ID
logInfo("Order created successfully", { orderId: "ord_123" });
// {"timestamp":"2026-03-22T14:30:00Z","level":"INFO","message":"Order created successfully","traceId":"abc123...","orderId":"ord_123"}

Trace-Metric Correlation

Link metrics to traces through span events:

const meter = metrics.getMeter("payment-service");

const paymentDuration = meter.createHistogram("payment.duration", {
  unit: "ms",
  description: "Payment processing duration",
});

async function processPayment(payment: Payment): Promise<void> {
  const span = tracer.startSpan("PaymentService.process");

  const startTime = Date.now();
  try {
    await doPayment(payment);
    paymentDuration.record(Date.now() - startTime, {
      "payment.method": payment.method,
      "payment.status": "success",
    });
  } catch (error) {
    paymentDuration.record(Date.now() - startTime, {
      "payment.method": payment.method,
      "payment.status": "failure",
    });
    throw error;
  } finally {
    span.end();
  }
}

Sampling Strategies

At high traffic, you cannot capture every trace. Sampling reduces volume while preserving useful data.

Common Sampling Strategies

Head-based sampling decides at trace start whether to record the trace. This Node.js example keeps a deterministic 1% sample; it cannot select errors based on a status that has not happened yet:

import { NodeSDK } from "@opentelemetry/sdk-node";
import { TraceIdRatioBasedSampler } from "@opentelemetry/sdk-trace-node";

const sdk = new NodeSDK({
  sampler: new TraceIdRatioBasedSampler(0.01),
  traceExporter: exporter,
});

sdk.start();

Tail-based sampling buffers spans in the collector, then decides what to keep after evaluating the trace:

# OTel Collector tail-based sampling
processors:
  tail_sampling:
    decision_wait: 10s
    num_traces: 100000
    policies:
      - name: errors
        type: status_code
        status_code: { status_codes: [ERROR] }
      - name: slow-traces
        type: latency
        latency: { threshold_ms: 1000 }
      - name: probabilistic
        type: probabilistic
        probabilistic: { sampling_percentage: 10 }
      - name: keep-all
        type: always_sample

This captures slow traces, errors, and a percentage of everything else.

Tail sampling only works when the collector sees the complete trace. Route spans for a trace ID to the same sampling instance, size its memory for the arrival rate and decision_wait, and alert on late spans, evictions, and dropped data. If traces routinely arrive after the decision window, increase the wait only after checking the added memory cost. Keep head-based sampling decisions consistent across services; a downstream sampler that discards a parent-sampled trace can leave a partial request.

Cloud-Native Tracing Solutions

AWS X-Ray

AWS X-Ray integrates with services like API Gateway, Lambda, ECS, and EKS:

import { AWSXRay } from "aws-xray-sdk";

// Automatic tracing for AWS SDK calls
AWSXRay.captureAWSv3Client(s3Client);
AWSXRay.captureHTTPClient(httpAgent);

// For Lambda, use the wrapper
export const handler = AWSXRay.captureAsyncHandler(async (event, context) => {
  // Your handler code
  return await processOrder(event);
});

X-Ray uses a daemon that buffers traces and sends them to the AWS backend. In ECS, run the X-Ray daemon as a sidecar container.

Google Cloud Trace

GCP Cloud Trace integrates with Cloud Run, GKE, and Compute Engine:

import { TraceAgent } from "@google-cloud/trace-agent";

// Initialize before other imports
TraceAgent.start({
  projectId: process.env.GCP_PROJECT_ID,
  keyFilename: "/path/to/service-account.json",
  logLevel: 1,
});

// OpenTelemetry SDK with GCP exporter
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-grpc";

const traceExporter = new OTLPTraceExporter({
  url: "collector.googleapis.com:443",
  headers: {
    "x-goog-api-key": process.env.GCP_API_KEY,
  },
});

Azure Application Insights

Azure uses the OpenTelemetry SDK with its own exporter:

import { ApplicationInsights } from "@microsoft/applicationinsight-web";

// Auto-instrument HTTP and AJAX calls
const appInsights = new ApplicationInsights({
  config: {
    instrumentationKey: process.env.AZURE_INSTRUMENTATION_KEY,
    enableCorsCorrelation: true,
    autoTrackPageVisit: true,
  },
});

appInsights.loadAppInsights();
appInsights.trackTrace({
  message: "Distributed tracing initialized",
  severityLevel: 1,
});

Multi-Cloud Trace Correlation

When running across cloud providers, maintain trace context using W3C headers. The traceparent header works across all providers:

import { context, propagation } from "@opentelemetry/api";

async function forwardToExternalService(
  url: string,
  headers: Headers,
): Promise<Response> {
  const outgoingHeaders: Record<string, string> = Object.fromEntries(
    headers.entries(),
  );
  propagation.inject(context.active(), outgoingHeaders);

  return fetch(url, { headers: outgoingHeaders });
}

Visualization with Jaeger

Jaeger is a popular distributed tracing backend. It stores traces and provides a UI for exploration.

Key Jaeger Views

Jaeger’s UI is organized around three main views. Each one serves a different purpose in your debugging workflow, and knowing how to use them together turns a trace dump into actionable insight.

Search is where every investigation starts. You filter by service name, operation, time range, and tags — those tag values come directly from the span attributes you set, like order.customer_tier or http.status_code. The search results show matching traces sorted by recency or duration, with a sparkline of span count and total time. If you know an error happened between 14:00 and 14:05 in the payment service, search narrows you down to the handful of traces that actually matter instead of forcing you to scroll through thousands. Pro tip: combine a broad time range with a specific error tag filter for the fastest path to the root cause.

Trace Detail is the flame-graph view that shows every span as a horizontal bar. The bar’s length represents duration, and its position on the timeline shows when it ran relative to the parent span. A wide bar that takes up most of the trace is your first bottleneck candidate. Click any span to see its start time, duration, and attribute set. The critical path — the chain of spans that accounts for the longest sequential execution — is often highlighted or visually obvious: look for bars that are both wide and nested deep. This view is where you discover that what you thought was a database timeout is actually a serialization delay in the upstream auth call.

Span Detail is the inspection panel. When you click a single span in the flame graph, this panel opens with the full set of attributes, events, and logs attached to that span. This is where the recordException calls you wrote in your instrumentation pay off — the error message, stack trace, and any custom attributes you set will be right here. You can also see the parent-child hierarchy with span IDs, which helps confirm the propagation chain is intact. If a trace looks broken (orphan spans, missing children), the span detail panel is where you verify the traceparent header made it through.

Analyzing Trace Flame Graphs

A flame graph shows the parent-child span relationships:

order-service.createOrder (2.3s)
├── auth-service.validateToken (50ms)
├── inventory-service.checkStock (150ms)
│   └── external-partner.getAvailability (120ms)
├── payment-service.process (1.8s)
│   ├── fraud-check.analyze (400ms)
│   │   └── external-api.call (380ms)
│   └── payment-gateway.charge (1.2s)
│       └── external-bank.authorize (1.1s)
├── notification-service.send (100ms)
└── db.orders.insert (30ms)

Long spans are easy to spot. Here, payment-service.process dominates. Drilling in, payment-gateway.charge is the bottleneck. Further still, external-bank.authorize is where time is spent.

Service Mesh Tracing (Istio and Linkerd)

Service meshes add automatic tracing to all service-to-service communication without requiring code changes.

Istio Integration

Istio’s Envoy sidecar proxy automatically instruments all HTTP, gRPC, and TCP traffic:

# istio tracing config
apiVersion: install.istio.io/v1alpha1
kind: IstioOperator
metadata:
  name: tracing-config
spec:
  meshConfig:
    enableTracing: true
    defaultProviders:
      tracing:
        - opentelemetry
    extensionProviders:
      - name: otel
        opentelemetry:
          service: otel-collector.observability
          port: 4317

Envoy extracts trace context from traceparent headers and creates spans for every request. Your application code only needs to propagate context for async operations.

Linkerd Integration

Linkerd uses service profiles to enable tracing on specific routes:

# service-profile.yaml
apiVersion: linkerd.io/v1alpha2
kind: ServiceProfile
metadata:
  name: order-service.default.svc.cluster.local
spec:
  routes:
    - condition:
        method: GET
        path: /api/orders/{id}
      timeout: 5s
      retryBudget:
        retryRatio: 0.2
        minRetriesPerSecond: 10
        maxRetries: 100

Trade-offs: Mesh vs SDK Tracing

Aspect Service Mesh Tracing SDK Manual Tracing
Setup effort Minimal (config only) Code changes required
Network span coverage Automatic for all traffic Only where you add it
Business context Limited to mesh metadata Full custom attributes
Performance impact Sidecar overhead ~1-2ms Minimal with sampling
Portability Tied to mesh implementation Portable across platforms

Common Patterns

Database Query Tracing

Database queries are one of the most common sources of latency in web services. Tracing them reveals which queries are slow, which are called repeatedly in a single request, and how database time adds up in the overall request duration.

The PostgreSQL instrumentation plugin monkey-patches the pg client to wrap every query in a span. Turn on enhancedDatabaseReporting to capture the full query text and bind parameters. This lets you pinpoint exactly which query pattern is causing trouble. The addSqlCommenterCommentToQueries option embeds trace context as SQL comments. This lets you correlate traced queries with database-level performance monitoring.

Similar plugins exist for MySQL, SQLite, Redis, MongoDB, and most other databases. Auto-instrumentation covers the common cases. For custom query builders or ORM operations the built-in plugins do not handle, you can create manual spans.

// Monkey-patch your database client for auto-tracing
import { dbplugin } from "@opentelemetry/instrumentation-pg";

new dbplugin.DatabaseDetector({
  enhancedDatabaseReporting: true,
  addSqlCommenterCommentToQueries: true,
});

HTTP Client Tracing

Outgoing HTTP calls to external services are natural tracing boundaries. Instrumenting your HTTP client captures every external request as a span. You get the target URL, HTTP method, response status code, and duration. This shows you how external dependencies affect your service’s latency.

The fetch instrumentation plugin wraps the browser’s native fetch API or Node.js fetch to create spans for outgoing requests automatically. The propagateCorrelationHeader option enables end-to-end tracing by injecting trace context headers into the request. The downstream service can continue the trace if it also supports W3C Trace Context.

For environments using axios, got, or other HTTP libraries, OpenTelemetry provides dedicated instrumentation packages. They capture request and response metadata, handle context propagation automatically, and require no changes to your existing HTTP call patterns.

import { fetchInstrumentation } from "@opentelemetry/instrumentation-fetch";

new fetchInstrumentation({
  propagateCorrelationHeader: true,
  timingOrigin: (origin) => origin !== window.location.origin,
});

gRPC Tracing

gRPC is common in microservice architectures. Its binary protocol carries metadata channels that work well for trace context propagation. The gRPC instrumentation plugin intercepts both client and server calls. It creates spans for each RPC with attributes for the service name, method, status code, and message size.

On the server side, the plugin extracts incoming trace context from gRPC metadata and creates server spans as children of the caller’s span. On the client side, it injects context into outgoing metadata and creates client spans. This bidirectional coverage means the full RPC call shows up as a single connected path in your trace visualization.

The plugin supports both unary and streaming RPCs. For streaming calls, it creates a span for the overall stream connection with events marking message sends and receives. This is useful for debugging streaming bottlenecks where individual messages arrive slowly even though the connection remains open.

import { grpcInstrumentation } from "@opentelemetry/instrumentation-grpc";

new grpcInstrumentation({
  yaml: true,
});

When to Use Distributed Tracing

Use distributed tracing when:

  • Understanding request flow through complex architectures
  • Finding which service causes cascading failures
  • Root cause analysis when errors propagate across boundaries
  • Performance optimization by identifying bottlenecks
  • Validating service dependencies and communication patterns

When Not to Use Distributed Tracing:

  • Single monolithic applications (local debugging suffices)
  • Low-traffic services where logs provide sufficient context
  • When you only need aggregate metrics (use Prometheus)
  • Very high-throughput paths where tracing overhead matters (use sampling)
  • Systems without clear request boundaries (batch jobs)

Trade-off Analysis

Aspect Distributed Tracing Traditional Logging Metrics Only
Debugging Speed Minutes (full context) Hours (manual correlation) N/A
Storage Cost High (span data) Medium (log volume) Low
Overhead ~1-5% latency Minimal Minimal
Root Cause Full causal chain visible Requires ID correlation Aggregates only
Cardinality High (many traces) Medium Low
Error Context Full request path Per-service only None

SLI/SLO/Error Budget Templates for Tracing

Distributed tracing does not typically have traditional SLIs/SLOs since it is qualitative debugging tooling rather than quantitative reliability measurement. However, you can define SLOs around tracing coverage and health.

Trace Health SLI Template

# tracing-sli-config.yaml
service: tracing-observability
environment: production

slis:
  - name: trace_ingestion_success_rate
    description: "Percentage of started traces successfully exported"
    query: |
      sum(rate(otel_exporter_sent_spans_total[5m]))
      /
      sum(rate(otel_span_started_total[5m]))

  - name: trace_context_propagation_success
    description: "Percentage of requests with valid propagated trace context"
    query: |
      sum(rate(otel_trace_context_propagated_total{status="success"}[5m]))
      /
      sum(rate(otel_trace_context_propagated_total[5m]))

  - name: tail_sampling_efficiency
    description: "Percentage of traces retained by tail sampling"
    query: |
      sum(rate(otel_tail_sampling_traces_sampled_total[5m]))
      /
      sum(rate(otel_tail_sampling_traces_evaluated_total[5m]))

  - name: span_error_rate
    description: "Percentage of spans with error status"
    query: |
      sum(rate(otel_span_status_code_total{code="ERROR"}[5m]))
      /
      sum(rate(otel_span_started_total[5m]))

Trace SLO Template

# tracing-slo-config.yaml
objectives:
  - display_name: "Trace Ingestion Availability"
    sli: trace_ingestion_success_rate
    target: 99.5
    window: 30d
    description: "99.5% of started traces should be exported"

  - display_name: "Context Propagation Success"
    sli: trace_context_propagation_success
    target: 99.9
    window: 30d
    description: "99.9% of requests should have valid trace context"

  - display_name: "Tail Sampling Coverage"
    sli: tail_sampling_efficiency
    target: 95.0
    window: 30d
    description: "95% of sampled traces should match sampling policies"

Error Budget Calculator for Tracing

def calculate_tracing_budgets():
    """
    Calculate error budgets for tracing SLOs (30-day window).
    """
    window_minutes = 30 * 24 * 60

    slos = {
        "99.5% (Trace Ingestion)": window_minutes * 0.005,
        "99.9% (Context Propagation)": window_minutes * 0.001,
        "95.0% (Tail Sampling)": window_minutes * 0.050,
    }

    for slo, budget in slos.items():
        print(f"{slo}: {budget:.1f} minutes allowed degradation")
        print(f"  = {budget / 60:.2f} hours")
        print(f"  = {budget / 60 / 24:.2f} days")

calculate_tracing_budgets()

Multi-Window Burn-Rate Alerting for Tracing

Tracing is diagnostic infrastructure, so treat these as coverage alerts rather than service-availability SLOs. The SLI templates above define what to measure; this example pages when exported span coverage stays below its threshold.

Trace Coverage Alert (1h Window)

# Tracing coverage alert
groups:
  - name: tracing-burn-rate
    rules:
      # Fast burn: Trace ingestion dropping significantly
      - alert: TracingCoverageFastBurn
        expr: |
          (
            sum(rate(otel_exporter_sent_spans_total[1h]))
            /
            sum(rate(otel_span_started_total[1h]))
          )
          < 0.95
        for: 5m
        labels:
          severity: critical
          category: tracing
          window: 1h
        annotations:
          summary: "Trace coverage dropping fast (1h window)"
          description: "Trace ingestion success rate is {{ $value | humanizePercentage }}. Investigate OTel collector health or exporter issues."

Observability Hooks for Distributed Tracing

This inventory names the signals to emit and collect for the tracing pipeline. Use the SLI section to define coverage objectives and the preceding alert example to set paging policy.

Log (What to Emit)

Event Fields Level
Collector started version, endpoint, exporters INFO
Exporter failure exporter_type, error, retry_count WARN
Sampling decision sampling_policy, trace_id, decision DEBUG
Context propagation failure service, direction, error WARN
Span queue full service, queue_size, drop_count ERROR
Batch export success exporter, spans_count, bytes DEBUG

Measure (Metrics to Collect)

Metric Type Description
otel_span_started_total Counter Total spans started
otel_span_ended_total Counter Total spans ended
otel_exporter_sent_spans_total Counter Spans successfully exported
otel_exporter_failed_spans_total Counter Spans that failed to export
otel_trace_context_propagated_total Counter Context propagation attempts
otel_tail_sampling_traces_evaluated_total Counter Traces evaluated by tail sampler
otel_tail_sampling_traces_sampled_total Counter Traces retained by tail sampler
otel_span_queue_depth Gauge Pending spans in export queue
otel_collector_receive_latency_seconds Histogram Time to receive spans
otel_exporter_send_latency_seconds Histogram Time to send to backend

Trace (Correlation Points)

Operation Trace Attribute Purpose
Span started tracing.otel.version Track OTel SDK version
Sampling decision tracing.sampling.decision Monitor sampling efficiency
Export batch tracing.export.batch_size Track export efficiency
Context inject/extraction tracing.context.direction Monitor propagation health

Alert (When to Page)

Alert Condition Severity Purpose
Trace Silence No spans exported for 5 minutes P1 Critical Tracing pipeline down
Export Failure Rate Export failures > 5% for 5 min P1 Critical Data loss imminent
Context Propagation Failure Propagation failures > 1% P2 High Incomplete traces
Span Queue Critical Queue > 90% capacity P2 High Risk of drops
Tail Sampling Bypass Sampled < expected with high errors P3 Medium Sampling misconfigured
Collector Latency Receive latency > 1s p95 P3 Medium Performance issue

Keep dashboard panels and alerts for queue saturation, export failures, context propagation failures, and collector latency. Add service-specific thresholds and runbooks after measuring normal traffic; metric names and availability vary by SDK and collector distribution.

Trace Storage and Retention Considerations

Choosing a trace storage backend and defining retention policies are critical decisions for production tracing systems. The storage layer affects query performance, operational costs, and your ability to debug issues after the fact.

Storage Operations

Storage Backend Options

Self-hosted options give you control over data and infrastructure:

Backend Best For Limitations
Jaeger with Elasticsearch Flexible querying, multi-tenant Operational complexity
Jaeger with Cassandra High write throughput Limited query capabilities
Jaeger with badger Small-scale, simplicity Not distributed
Zipkin with Elasticsearch Basic needs Fewer features than Jaeger

Managed cloud options reduce operational overhead:

Service Advantages Trade-offs
AWS X-Ray Deep AWS integration Vendor lock-in
GCP Cloud Trace Auto-scaling, strong perf GCP dependency
Azure Application Insights Full APM features Azure dependency
Honeycomb Sophisticated queries Cost at high volume
Datadog Comprehensive platform Expensive at scale

Retention Planning

Trace data follows a lifecycle pattern:

# Tiered retention example
retention_tiers:
  hot_storage:
    duration: 7 days
    sampling: 100% for errors, 10% for normal
    compression: none
    storage: fast SSD

  warm_storage:
    duration: 30 days
    sampling: 100% errors, 1% normal
    compression: lz4
    storage: standard block storage

  cold_storage:
    duration: 1 year
    sampling: errors only
    compression: zstd
    storage: object storage (S3, GCS)

Sampling and retention solve different problems: sampling decides which traces enter storage, while retention decides how long stored traces remain available. Set retention by investigative value, data sensitivity, and contractual requirements, then verify that expiry covers replicas and backups where the backend supports it. Revisit the policy when traffic or incident response needs change; a long retention period cannot recover traces that were never sampled.

Partitioning Strategies

High-volume trace stores require careful partitioning:

# Elasticsearch index per time window
indices:
  pattern: "traces-{service}-{yyyy.MM.dd}"
  rollovers:
    - max_age: 7d
      max_docs: 50 million
  aliases:
    write: "traces-write"
    read: "traces-read"

Query Performance at Scale

As trace volume grows, query performance degrades without proper optimization:

// Optimize trace queries with date filtering
async function queryTraces(service: string, startTime: Date, endTime: Date) {
  // Always filter by time range first - reduces scan scope
  const query = {
    index: `traces-${service}-*`,
    body: {
      query: {
        bool: {
          must: [
            { range: { timestamp: { gte: startTime, lte: endTime } } },
            { term: { "service.name": service } },
          ],
        },
      },
      sort: [{ timestamp: "desc" }],
      size: 100, // Limit results
    },
  };

  return elasticsearch.search(query);
}

Reliability

Data Lifecycle Management

Automate data lifecycle to prevent unbounded growth:

# OTel Collector with lifecycle management
exporters:
  otlp/jaeger:
    endpoint: jaeger:4317
    retry_on_failure:
      enabled: true
      initial_interval: 5s
      max_interval: 30s
      max_elapsed_time: 5m

processors:
  # Tag spans with expiration metadata
  resource:
    attributes:
      - action: upsert
        key: data_category
        value: tracing

  # Batch and compress before export
  batch:
    timeout: 10s
    send_batch_size: 8192

Backup and Recovery Considerations

Trace data recovery is often overlooked:

  • Regular backups: Schedule Elasticsearch snapshots or managed service backups
  • Point-in-time recovery: Test restoration procedures periodically
  • Cross-region replication: Replicate critical trace data to secondary region
  • RTO/RPO planning: Define acceptable downtime and data loss windows for tracing infrastructure

Cost Optimization Patterns

Trace storage costs scale with volume. Optimize with these approaches:

# Cost optimization configuration
processors:
  # Prune low-value attributes before storage
  transform:
    trace_state: "(trace_state):lens(include: [service.name, operation.name, error])"

  # Aggregate redundant data
  groupbyattrs:
    keys: ["service.name", "operation.name", "http.status_code"]
    mode: sum

  # Compress spans with limited attributes
  memory_limiter:
    check_interval: 1s
    limit_mib: 1000
    spike_limit_mib: 200

Multi-Tenant Trace Storage

When serving multiple customers from shared infrastructure:

# Multi-tenant storage isolation
tenants:
  - name: customer-a
    index_prefix: "traces-a"
    retention_days: 30
    quota:
      storage_gb: 100
      queries_per_minute: 60

  - name: customer-b
    index_prefix: "traces-b"
    retention_days: 90
    quota:
      storage_gb: 500
      queries_per_minute: 120

Implement tenant isolation at the query layer to prevent cross-tenant data leakage.

Production Failure Scenarios

Symptom Confirm with First response
A request becomes a separate trace after one service Compare trace IDs and traceparent at the boundary; check whether injection and extraction both ran Repair the specific HTTP or messaging instrumentation path, then verify with a controlled request
An error request has no trace, or a trace ends before the error Check SDK sampling flags, collector tail-sampling policies, and dropped-span counters Preserve a baseline sample and confirm error policy coverage end to end
Traces are delayed or incomplete across many services Check collector queue depth, memory limiter events, exporter retries, and backend ingestion latency Restore collector capacity or backend connectivity; expect queued spans to expire or be dropped under sustained pressure
A slow trace shows impossible or negative durations Compare service clocks and span start/end times; inspect whether the spans came from different hosts Correct clock synchronization and use parent-child timing carefully across hosts
Searches return traces but cannot isolate the failing operation Inspect resource attributes, operation names, status, and error events on representative spans Add stable, low-cardinality attributes and record errors at the failing boundary
A trace has duplicate or misleading service spans Compare SDK and service-mesh instrumentation for the same request Keep both only when they answer different questions; otherwise disable the duplicate instrumentation

Common Pitfalls / Anti-Patterns

1. Creating Spans for Everything

Every span has overhead. Do not create spans for every loop iteration or minor function call:

// Bad: Spans for everything
async function processItems(items: Item[]) {
  const span = tracer.startSpan("processItems");
  for (const item of items) {
    const itemSpan = tracer.startSpan("processItem"); // Too granular
    await processItem(item);
    itemSpan.end();
  }
  span.end();
}

// Good: Batch operations as single span
async function processItems(items: Item[]) {
  const span = tracer.startSpan("processItems");
  const results = await Promise.all(items.map((item) => processItem(item)));
  span.setAttribute("items.count", items.length);
  span.end();
  return results;
}

2. Forgetting to End Spans

Unfinished spans remain open and appear as ongoing operations:

// Bad: Span not ended on error path
async function riskyOperation() {
  const span = tracer.startSpan('risky');
  if (condition) {
    throw new Error('condition failed');
  }
  span.end(); // May never execute
}

// Good: Use try/finally
async function riskyOperation() {
  const span = tracer.startSpan('risky');
  try {
    // Work
    span.setStatus({ code: SpanStatusCode.OK });
  } catch (e) {
    span.recordException(e);
    span.setStatus({ code: SpanStatusCode.ERROR });
    throw;
  } finally {
    span.end();
  }
}

3. Not Propagating Context Across Async Boundaries

Async operations lose trace context without explicit propagation:

// Bad: Context lost
async function outer() {
  const span = tracer.startSpan("outer");
  await inner(); // Span context not passed
  span.end();
}

async function inner() {
  const span = tracer.startSpan("inner"); // Orphan span
  span.end();
}

// Good: startActiveSpan makes each operation's span current while its callback runs.
// This relies on an SDK context manager configured for asynchronous code.
async function outer(): Promise<void> {
  await tracer.startActiveSpan("outer", async (outerSpan) => {
    try {
      await inner();
    } finally {
      outerSpan.end();
    }
  });
}

async function inner(): Promise<void> {
  await tracer.startActiveSpan("inner", async (innerSpan) => {
    try {
      await doWork();
    } finally {
      innerSpan.end();
    }
  });
}

4. Storing Too Much Data in Span Attributes

Span attributes are not a data store. Keep them small and queryable:

// Bad: Large data in attributes
span.setAttribute("response_body", JSON.stringify(largeObject));

// Good: Reference data by ID
span.setAttribute("order_id", order.id);
span.setAttribute("items_count", order.items.length);

5. Ignoring Sampling in High-Volume Services

Unsampled tracing at high volume creates massive overhead:

# OTel Collector tail sampling
processors:
  tail_sampling:
    decision_wait: 10s
    policies:
      - name: errors
        type: status_code
        status_code: { status_codes: [ERROR] }
      - name: slow-traces
        type: latency
        latency: { threshold_ms: 2000 }
      - name: probabilistic
        type: probabilistic
        probabilistic: { sampling_percentage: 1 }

Observability Checklist

Tracing Coverage

  • HTTP request/response spans for all API endpoints
  • Database query spans with statement and duration
  • External API call spans with URL and status
  • Message queue publish/consume spans
  • Background job spans with job ID and outcome
  • Custom business operation spans with relevant context

Span Attributes

  • Service name and version
  • Operation name
  • Trace ID and span ID
  • Start time and duration
  • HTTP: method, URL, status code
  • DB: system, statement, rows affected
  • Business: entity IDs, customer tier, transaction amount

Correlation

Correlation is what makes traces useful in practice. Without it, a trace is a self-contained timing diagram for a single request. With it, you can jump from a slow trace to the specific log lines that explain what happened, or from an error rate spike in your metrics dashboard to the traces of the requests that are failing. The key is consistent trace ID propagation: every log line, every metric emission, and every span must carry the same trace ID.

The trace ID should appear in your structured log output as a top-level field, not buried inside a message string. If your logs look like {“message”: “checkout failed trace_id=abc123”}, you cannot query efficiently. If they look like {“trace_id”: “abc123”, “message”: “checkout failed”}, you can filter all logs for a trace in one query. This is a small difference in log format that produces a large difference in debugging workflow.

For metrics, include the trace ID only where it adds value. Attaching trace IDs to every metric from every service creates cardinality explosion. The right pattern is to include trace-derived dimensions — like service name, endpoint, or error type — on metrics that aggregate across many traces, and include the full trace ID only on metrics emitted from within a span context where you have a specific investigation to conduct.

  • Trace ID included in all log entries
  • Trace ID included in metric labels (where appropriate)
  • Log entries linkable from span events
  • Metrics aggregatable by trace-derived dimensions

Sampling Configuration

Sampling configuration decides which traces you keep and which get dropped. Get it wrong and you either lose visibility when you need it most, or you burn through storage so fast that finance starts asking questions. The two main approaches are head-based sampling (decision at request start) and tail-based sampling (decision after the trace finishes).

Head-based sampling is straightforward. Pick a percentage and every request competes for that slot equally. You get a consistent baseline view of your system, but you cannot tell the difference between a trace that errored out at the last millisecond and one that sailed through cleanly. That error trace might be the most important one to keep. Tail-based sampling fixes this by holding traces briefly, then deciding based on what actually happened. Error traces get kept no matter how short. Slow traces get kept even if the head-based sampler missed them.

Most production systems run both. Head-based at low percentage (1-10%) keeps storage predictable. Tail-based picks up the errors and outliers that probabilistic sampling misses. The combination is affordable and comprehensive.

The checklist below is the minimum viable setup. Adjust percentages based on your traffic volume and storage budget. High-traffic services often drop head-based to 0.1% and let tail-based do more work. Payment and authentication paths should be sampled at 100%, no exceptions — you do not get to skip debugging those failures.

  • Head-based sampling for consistent baseline (1-10%)
  • Tail-based sampling for errors (100% of errors)
  • Tail-based sampling for slow traces (>threshold)
  • Always sample for tagged critical requests

Incident Triage

  • Start from the failing request’s trace ID, time range, and affected service; check whether the trace is absent, partial, delayed, or complete.
  • If spans are missing, compare propagation headers at the last visible boundary before investigating backend storage.
  • If the trace is delayed or incomplete across services, inspect collector queue depth, memory pressure, export retries, and backend ingestion health.
  • If an error trace is absent, verify head-sampling flags and tail-sampling policies, then check for evictions or late spans.
  • Compare the trace’s critical path with service metrics and structured logs before assigning the incident to a dependency.
  • Record the trace ID, the last complete span, the first missing boundary, and any dropped-span signal in the incident notes.

Security and Compliance Notes

Trace context helps services correlate work; it does not establish who made a request or whether that request is allowed. Treat incoming traceparent, tracestate, and baggage as untrusted input, especially at public endpoints and partner boundaries. Validate their format and size, discard invalid values, and create a new trace when policy requires an untrusted trace to stop at the boundary. Do not put identity, authorization decisions, secrets, or customer data in baggage, and do not use a trace ID as an access-control key.

Propagation also needs an explicit boundary policy. Forward context only to services and partners that are expected to receive it. Strip or replace headers when sending requests to unrelated external hosts, and ensure proxies do not accept caller-supplied context as evidence that a request came from an internal service. Keep internal topology out of exports sent to third parties unless the recipient and contract permit it.

Span attributes and events can contain sensitive data even when instrumentation did not intend to collect it. Avoid passwords, access tokens, email addresses, payment data, raw request or response bodies, full URLs with query parameters, and SQL statements containing literal values. Prefer low-cardinality, non-identifying fields. Apply allowlists and redaction or hashing in the SDK or collector before export, and review exception messages and stack traces for embedded secrets. Redaction should happen before data reaches a backend, since deleting it later may not remove indexed copies or backups.

Restrict trace access with least-privilege roles, tenant isolation, and audited administrative access. Encrypt telemetry in transit and at rest, and limit collector credentials to the destinations and operations they need. Set retention by data class and business purpose, including deletion from replicas and backups where supported; longer retention increases exposure and may conflict with privacy or data-residency requirements.

Sampling changes what evidence remains. Head sampling can discard a trace before an error or security event occurs, while tail sampling can miss spans if traces arrive late, are incomplete, or exceed collector capacity. Keep a baseline sample, test policies against incident scenarios, and monitor dropped or undecided traces. Traces are useful for investigations, but they are not a complete audit log and should not be the sole record for compliance or security events.

Use this checklist when reviewing a tracing deployment:

  • Validate incoming trace context and define where propagation stops.
  • Keep credentials, direct identifiers, and sensitive payloads out of span data; redact before export.
  • Limit trace access by role and tenant, and audit privileged access.
  • Encrypt trace data in transit and at rest; scope collector credentials.
  • Set and enforce retention and deletion rules for storage, replicas, and backups.
  • Test sampling policies for security incidents and document what telemetry they can omit.
  • Do not rely on traces as the only compliance or security audit record.

Quick Recap Checklist

  • Extract and inject trace context at every HTTP, RPC, and messaging boundary; test async paths separately.
  • Use spans for meaningful operations and keep their names and attributes stable enough to query.
  • Put trace IDs in structured logs, then compare trace timing with metrics during an investigation.
  • Keep sampling decisions consistent across services; route complete traces together when tail sampling.
  • Set retention separately from sampling and confirm expiry covers replicas and backups.
  • Keep secrets and personal data out of span attributes and events, and do not treat traces as an audit log.

Interview Questions

1. What is the difference between a trace and a span in distributed tracing?

Expected answer points:

  • A trace represents the complete end-to-end journey of a single request through all services
  • A span is a single unit of work within that trace, representing one operation or service call
  • Spans are organized hierarchically with parent-child relationships forming the trace tree
  • Each span captures timing, attributes, events, and status about that specific operation
2. How does W3C Trace Context propagation work across service boundaries?

Expected answer points:

  • Trace context propagates via HTTP headers, most importantly the traceparent header
  • The traceparent header contains: version (2 chars), trace ID (32 hex chars), parent ID (16 hex chars), and flags
  • When service A calls service B, it injects the trace context into outgoing request headers
  • Service B extracts the context and creates a child span, linking to the parent
  • The tracestate header allows for vendor-specific propagation data
3. What are the main components of the OpenTelemetry architecture?

Expected answer points:

  • Application Code / SDK: Language-specific instrumentation libraries that create spans
  • Auto-instrumentation: Framework-specific agents that instrument common operations automatically
  • Collector: Middleware that receives, processes, and exports telemetry data
  • Exporters: Connectors that send data to backends like Jaeger, Zipkin, or cloud providers
  • The OTel SDK is vendor-neutral, allowing you to switch backends without code changes
4. Explain head-based sampling vs tail-based sampling. When would you use each?

Expected answer points:

  • Head-based sampling decides at trace start whether to capture, using probabilistic or rule-based selection
  • Tail-based sampling captures all spans temporarily, then decides what to keep after the trace completes
  • Head-based sampling is simpler and has lower memory overhead since you discard early
  • Tail-based sampling enables intelligent decisions like "keep all errors" or "keep slow traces" after seeing the full picture
  • Production systems often use both: head-based for consistent baseline sampling, tail-based for targeted capture of important traces
5. How do you propagate trace context through asynchronous message queues like Kafka?

Expected answer points:

  • Producer injects trace context into message headers before sending
  • Context is serialized into headers like traceparent using W3C format
  • Consumer extracts context from message headers and creates a linked span
  • Use context.with(extractedContext, () => { ... }) to run handlers within the correct context
  • This ensures traces span across async boundaries, showing the full request flow even through queues
6. What are semantic conventions for span attributes and why are they important?

Expected answer points:

  • Semantic conventions are standardized attribute names for common operations (HTTP, DB, messaging)
  • Examples: http.method, http.status_code, db.system, db.statement
  • They enable interoperability between instrumentation from different libraries
  • Backend systems can interpret attributes consistently regardless of instrumentation source
  • They make traces queryable across your entire system using consistent filter names
7. How would you handle trace context propagation for external API calls that you cannot modify?

Expected answer points:

  • Use W3C traceparent header to propagate context to external services
  • If the external service supports W3C tracing, spans will be linked automatically
  • For services that don't propagate headers, create a span representing the external call with relevant attributes
  • Include the downstream service URL, response status, and duration as span attributes
  • Add custom attributes for business context even when you cannot instrument the remote service
8. What is the relationship between distributed tracing and the RED method (Rate, Errors, Duration)?

Expected answer points:

  • RED metrics are derived from trace data aggregated across similar spans
  • Rate: Request count per second, derived by counting spans per operation over time
  • Errors: Error rate calculated from spans with error status codes
  • Duration: Latency percentiles (p50, p95, p99) calculated from span durations
  • Traces provide the granular data; metrics are the rollup of that data for alerting
  • Use traces for debugging specific issues, use RED metrics for alerting and dashboards
9. What are the security considerations when implementing distributed tracing?

Expected answer points:

  • Never include passwords, tokens, or secrets in span attributes or events
  • Sanitize PII from span attributes before export
  • Encrypt trace data in transit using TLS
  • Implement access controls and audit logging for trace data access
  • Configure sampling to avoid capturing sensitive high-traffic endpoints excessively
  • Scrub or exclude headers like Authorization before creating spans
10. How would you debug a scenario where traces are being created but not linked across services?

Expected answer points:

  • Check if trace context is being extracted at service entry points (HTTP middleware)
  • Verify that context is being injected into outgoing requests
  • Look for async boundaries where context might be lost (missing context.with)
  • Check if message queue producers are injecting headers and consumers are extracting them
  • Verify all HTTP clients and message frameworks are instrumented
  • Check collector logs for context propagation failures
  • Ensure sampling decisions are consistent across the trace propagation path
11. What storage backend options exist for distributed traces, and how do you choose between them?

Expected answer points:

  • Jaeger (Cassandra, Elasticsearch, badger) - good for self-hosted with flexible querying
  • Zipkin (Cassandra, Elasticsearch, MySQL) - simpler alternative with basic search
  • AWS X-Ray (managed) - tight integration with AWS services but vendor lock-in
  • GCP Cloud Trace (managed) - seamless integration with Google Cloud, scales automatically
  • Azure Application Insights (managed) - comprehensive APM with built-in analytics
  • Choice depends on: existing cloud provider, query flexibility needs, operational overhead, cost
  • For multi-cloud: prefer vendor-neutral backends like Jaeger or self-hosted OTel-compatible storage
12. How do you determine appropriate retention periods for trace data?

Expected answer points:

  • Retention depends on use case: debugging (hours to days), compliance (months to years), analytics (aggregated indefinitely)
  • Hot storage (fast query): typically 7-30 days for recent traces
  • Cold storage (archive): months to years for historical analysis
  • Consider sampling older data - keep 100% for recent, sample for historical
  • Cost implications: trace data is voluminous; compression and tiered storage help
  • Compliance requirements may mandate minimum retention periods
  • Balance between investigative value and storage costs
13. What are the trade-offs between centralized trace storage and distributed edge storage?

Expected answer points:

  • Centralized (Jaeger, Zipkin): simpler operations, single query endpoint, potential network latency for upload
  • Edge storage (X-Ray daemon buffers): resilience to network partitions, reduced upload bandwidth, more complex retrieval
  • Hybrid approach: buffer at edge, batch upload to central, local fallback during outages
  • Consider data locality requirements - some regulations mandate data stays in certain regions
  • Edge buffering prevents data loss during collector downtime but requires disk management
  • Centralized storage simplifies debugging across services but creates dependency on network
14. How does the OTel Collector handle backpressure when the trace backend is unavailable?

Expected answer points:

  • OTel Collector has built-in sender functionality with retry mechanisms
  • When backend is down, spans queue in memory - risk of memory exhaustion under sustained load
  • Configure memory_limiter processor to drop spans when memory pressure exceeds threshold
  • Use persistent queue (disk-backed) for better resilience during backend outages
  • Exponential backoff with jitter prevents thundering herd when backend recovers
  • Dead letter queue / retry_stale configuration handles spans that cannot be exported
  • Monitor queue depth metrics to anticipate potential data loss
15. What strategies exist for reducing trace storage costs at scale?

Expected answer points:

  • Adaptive sampling: lower overall rate, 100% for errors and slow traces
  • Attribute pruning: remove low-value attributes before storage
  • Span deduplication: compress similar spans in batch operations
  • Data tiering: hot storage for recent data, archive/aggregate older data
  • Compression: use columnar formats (Parquet) that compress well
  • Trace summarization: keep full traces for errors, aggregated metrics for success paths
  • TTL enforcement: automatically expire old data based on retention policy
16. How do you implement multi-tenancy in a trace storage system?

Expected answer points:

  • Tenant isolation via separate indices/tables per customer (Jaeger with Elasticsearch)
  • Tag-based filtering: all spans tagged with tenant ID, query layer filters
  • Separate collectors or collector groups per tenant for strict data isolation
  • Consider data residency requirements - tenants may need data in specific regions
  • Resource quota enforcement to prevent one tenant from monopolizing storage
  • Access control: ensure tenants can only query their own trace data
  • Cost attribution: track storage and query costs per tenant for billing
17. What are the performance implications of trace collection and how do you optimize it?

Expected answer points:

  • Trace collection adds latency: OTel SDK overhead ~1-5ms per span creation
  • Batching exporters reduce network overhead by amortizing connection costs
  • Async export prevents blocking the main request path
  • SimpleSpanProcessor vs BatchSpanProcessor: batch is more efficient at scale
  • Collector pipeline: use processors to aggregate and reduce data before export
  • Network: consider gRPC vs HTTP exporters; gRPC has lower overhead for high volume
  • Profile in staging to understand actual overhead before production deployment
18. How would you design a trace data pipeline for a globally distributed system?

Expected answer points:

  • Regional collectors ingest locally, then forward to central aggregation
  • Use load balancing across collectors for horizontal scalability
  • Implement trace context propagation across regional boundaries
  • Consider data residency - some regions may require local storage before aggregation
  • Global sampling: each region samples independently, increasing total capture rate
  • Global view requires stitching traces from multiple regions - use consistent trace ID generation
  • Network design: dedicated links for trace traffic prevent interference with application traffic
19. What monitoring metrics should you track for your trace collection infrastructure?

Expected answer points:

  • Spans started vs ended (detector for leaks)
  • Export success/failure rate per backend
  • Queue depth and memory usage for exporters
  • Collector receive latency (p50, p95, p99)
  • Dropped spans count and reason (sampling, queue full, export failure)
  • Context propagation success/failure rate
  • Backend query latency for trace retrieval
  • Set SLOs/SLIs on these metrics and alert on violations
20. How does distributed tracing interact with event-driven architectures and saga patterns?

Expected answer points:

  • Saga orchestrator creates parent span; each saga step is a child span
  • Compensation operations (rollbacks) should be spans linked to the original transaction
  • Event-driven: inject context into message headers, extract in consumers
  • Choreography-based sagas: use correlation ID linking all related spans
  • Long-running sagas require sustained context propagation across hours or days
  • Consider span linking vs parent-based models for saga step relationships
  • Trace visualization helps identify bottleneck steps in saga execution

Further Reading

Conclusion

Distributed tracing follows a request across services by linking each operation to a trace with shared context. OpenTelemetry provides the APIs and collector pipeline for recording, sampling, and exporting those spans. When a trace is incomplete, check propagation at service and messaging boundaries, then compare span timing with logs and metrics.

Category

Related Posts

Jaeger: Distributed Tracing for Microservices

Learn Jaeger for distributed tracing visualization. Covers trace analysis, dependency mapping, and integration with OpenTelemetry.

#jaeger #tracing #observability

Debugging Backend Applications

Use a repeatable backend debugging workflow to reproduce failures, inspect evidence, test one hypothesis at a time, and verify fixes safely in production.

#debugging #backend #observability

Serverless Architecture: Boundaries, State, and Cold Starts

Understand serverless responsibility boundaries, event-driven design, cold starts, state, and observability before moving workloads to managed functions.

#software-architecture #serverless #cloud