Logging Best Practices: Structured Logs, Levels, Aggregation
Learn production logging with structured formats, useful log levels, correlation IDs, and scalable aggregation. Includes secure patterns for containerized apps.
Production logging works best when records are structured, carry request context, and reach a searchable store reliably. This guide covers log levels, correlation IDs, sensitive-data redaction, aggregation with tools such as Fluent Bit and Vector, and retention policies. It also explains how to manage logging cost and performance with asynchronous shipping, sampling, and pipeline health metrics. Use these patterns to make incidents easier to investigate while keeping log volume and access under control.
Logging Best Practices: Structured Logs, Levels, and Aggregation
Introduction
Logs help explain what happened in production after a user reports a problem. Plain text makes that investigation slow, especially when requests cross services and a team has to search unrelated files to reconstruct one event.
This guide covers structured formats, useful log levels, correlation IDs, aggregation, retention, and performance. It also explains how to protect sensitive values and keep a logging pipeline healthy at scale.
Log Levels
Log levels help filter noise. Not every log entry needs to be visible during normal operations.
Standard Log Levels
| Level | Purpose | When to Use |
|---|---|---|
| DEBUG | Detailed diagnostic information | During development and troubleshooting |
| INFO | Confirmation that things work as expected | Significant business events |
| WARN | Unexpected but handled situations | Recoverable errors, degraded states |
| ERROR | Errors that need attention | Failures that affect requests |
| FATAL | System cannot continue | Critical failures requiring immediate action |
Choosing the Right Level
This feels intuitive but gets harder at scale. A few guidelines:
- INFO for business events like orders placed, users registered. You want these for analytics and auditing.
- WARN for situations that require attention but the system continues: retries succeeded, cache misses, degraded mode.
- ERROR for failures that affect the current request: database timeout, external API failure, validation error.
- DEBUG for information that helps during development but would overwhelm production: loop iterations, intermediate values.
Don’t log DEBUG in production unless you can enable it selectively for specific requests. A debug log in a hot path can generate gigabytes per hour.
Correlation IDs
When a request flows through multiple services, correlation IDs let you follow it across all logs.
Propagating Trace Context
// Middleware to extract or generate correlation ID
function correlationMiddleware(req, res, next) {
const traceId = req.headers["x-trace-id"] || generateUUID();
req.correlationId = traceId;
res.setHeader("x-trace-id", traceId);
// Add to logger context
log = log.with({ traceId });
next();
}
Propagate the correlation ID to all downstream calls:
// Outgoing HTTP request
fetch("https://api.example.com/users", {
headers: {
"X-Trace-ID": req.correlationId,
},
});
// Database queries
db.query("SELECT * FROM users WHERE id = $1", [userId], {
traceId: req.correlationId,
});
// Message queue messages
queue.send({
payload: orderData,
headers: {
"X-Trace-ID": req.correlationId,
},
});
Using Correlation IDs for Search
With structured logs and correlation IDs, debugging a user issue looks like this:
# Find all logs for a specific request
grep '"trace_id":"abc123def456"' /var/log/app.log
# Or in your log aggregator
query: trace_id = "abc123def456"
Search for the trace ID and you get the incoming request, database queries, cache hits, outgoing API calls, and the error that occurred.
What to Include in Logs
Context matters. The more relevant context you include, the easier debugging becomes.
Essential Fields
Every log entry needs at minimum:
- timestamp: ISO 8601 format in UTC
- level: Log severity level
- service: Which service generated this log
- message: Human-readable description
Request Context
For web services, include:
- request_id or trace_id
- user_id (if authenticated)
- HTTP method, path, status code
- Client IP address
- User agent
{
"timestamp": "2026-03-22T14:32:01.456Z",
"level": "INFO",
"service": "api-gateway",
"message": "Request completed",
"request_id": "req_abc123",
"method": "GET",
"path": "/api/users/usr_789",
"status": 200,
"duration_ms": 120,
"ip": "192.168.1.42",
"user_agent": "Mozilla/5.0..."
}
Business Events
For significant business events:
- Event type (login, purchase, registration)
- Entity IDs involved
- Outcome (success, failure)
- Duration if applicable
- Any relevant metadata
{
"timestamp": "2026-03-22T14:32:01.456Z",
"level": "INFO",
"service": "checkout-service",
"message": "Order placed",
"event": "order_placed",
"order_id": "ord_xyz789",
"customer_id": "cust_123",
"total_amount": 99.99,
"currency": "USD",
"item_count": 3
}
What NOT to Log
Logging sensitive data creates security and compliance problems. Never log:
- Passwords or password hashes
- Credit card numbers or CVV codes
- Social security numbers or national IDs
- API keys or secrets
- Full authorization tokens (log the type and last 4 chars only)
Redact Sensitive Data
function redactSensitiveFields(value: unknown): unknown {
if (Array.isArray(value)) {
return value.map(redactSensitiveFields);
}
if (value === null || typeof value !== "object") {
return value;
}
const sensitiveFields = ["password", "token", "secret", "creditCard", "ssn"];
const redacted: Record<string, unknown> = {};
for (const [key, fieldValue] of Object.entries(value)) {
if (sensitiveFields.some((f) => key.toLowerCase().includes(f))) {
redacted[key] = "[REDACTED]";
} else {
redacted[key] = redactSensitiveFields(fieldValue);
}
}
return redacted;
}
log.info(
"User authenticated",
redactSensitiveFields({ userId: "usr_123", password: "secret123" }),
);
// Logs: { userId: 'usr_123', password: '[REDACTED]' }
Log Aggregation Architecture
At scale, logs need to be collected, aggregated, and stored efficiently.
Common Architecture
graph LR
A[Application] -->|stdout/JSON| B[Container Runtime]
B --> C[Log Agent]
C --> D[Log Aggregator]
D --> E[Storage]
D --> F[Search Interface]
G[Analytics/BI] --> E
Container Logging
In containerized environments, applications write to stdout and stderr. The container runtime handles collection:
# Write logs to stdout, not files
# Bad: RUN echo "$(date) Log entry" >> /var/log/app.log
# Good: console.log(JSON.stringify({ timestamp, message }))
For applications that must write to files, use a sidecar log agent or mount a shared log directory:
# Pod with log volume
apiVersion: v1
kind: Pod
metadata:
name: myapp
spec:
containers:
- name: app
image: myapp:latest
volumeMounts:
- name: logs
mountPath: /var/log/myapp
- name: log-agent
image: log-agent:latest
volumeMounts:
- name: logs
mountPath: /var/log/myapp
- name: agent-config
mountPath: /etc/log-agent
volumes:
- name: logs
emptyDir: {}
- name: agent-config
configMap:
name: log-agent-config
Shipping Logs to Aggregators
Fluentd/Fluent Bit Configuration
# fluent-bit.conf
[SERVICE]
Flush 5
Daemon Off
Log_Level info
Parsers_File parsers.conf
[INPUT]
Name tail
Path /var/log/containers/*.log
Parser docker
Tag container.*
Refresh_Interval 5
[FILTER]
Name kubernetes
Match container.*
Kube_URL https://kubernetes.default.svc:443
Kube_CA_File /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
Kube_Token_File /var/run/secrets/kubernetes.io/serviceaccount/token
[OUTPUT]
Name es
Match container.*
Host elasticsearch.logging.svc
Port 9200
Logstash_Format On
Logstash_Prefix kubernetes
Retry_Limit False
Vector Configuration
Vector is a newer alternative with better performance and lower resource usage:
# vector.toml
[sources.docker]
type = "docker_logs"
[transforms.parse_json]
type = "remap"
inputs = ["docker"]
source = '.message = parse_json!(.message)'
[sinks.elasticsearch]
type = "elasticsearch"
inputs = ["parse_json"]
endpoint = "http://elasticsearch.logging.svc:9200"
index = "kubernetes-%Y.%m.%d"
Log Storage and Retention
Storage costs grow with log volume. Design retention policies carefully.
Retention Tiers
| Tier | Duration | Use Case |
|---|---|---|
| Hot | 0-7 days | Real-time troubleshooting |
| Warm | 7-30 days | Investigating recent issues |
| Cold | 30-90 days | Compliance, audit |
| Archive | 1+ years | Legal requirements |
Elasticsearch Index Lifecycle Management
Without management, Elasticsearch indices grow until the disk fills. Index Lifecycle Management (ILM) automates moving indices through hot, warm, cold, and delete phases based on age or size, so you stop manually deleting old indices on Fridays. The policy below sets rollover at 7 days or 50GB in the hot phase, then shrinks and force-merges in warm, freezes in cold, and deletes after a year.
{
"policy": {
"phases": {
"hot": {
"actions": {
"rollover": {
"max_age": "7d",
"max_size": "50gb"
},
"set_priority": 100
}
},
"warm": {
"min_age": "7d",
"actions": {
"shrink": { "number_of_shards": 1 },
"forcemerge": { "max_num_segments": 1 },
"set_priority": 50
}
},
"cold": {
"min_age": "30d",
"actions": {
"freeze": {},
"set_priority": 0
}
},
"delete": {
"min_age": "365d",
"actions": {
"delete": {}
}
}
}
}
}
Performance Considerations
Logging can become a bottleneck if you don’t design it carefully.
Asynchronous Logging
Write logs asynchronously so they don’t block your application:
import logging
import queue
from threading import Thread
class AsyncLogHandler(logging.Handler):
def __init__(self, batch_size=100, flush_interval=1.0):
super().__init__()
self.queue = queue.Queue(maxsize=10000)
self.batch_size = batch_size
self.flush_interval = flush_interval
self.worker = Thread(target=self._process_logs, daemon=True)
self.worker.start()
def emit(self, record):
try:
self.queue.put_nowait(self.format(record))
except queue.Full:
pass # Drop log if queue is full
def _process_logs(self):
batch = []
while True:
try:
item = self.queue.get(timeout=self.flush_interval)
batch.append(item)
while len(batch) < self.batch_size:
item = self.queue.get_nowait()
batch.append(item)
except queue.Empty:
pass
if batch:
self._send_batch(batch)
batch = []
def _send_batch(self, batch):
# Send to log aggregator
pass
Sampling High-Volume Logs
For debug-level logs in high-traffic paths, sample to reduce volume:
const sampler = new RateSampler({ rate: 0.1 }); // 10% sample rate
log.debug(
{
message: "Processing item",
itemId: item.id,
sampled: sampler.sample(),
},
"Item processing details",
);
// Only actually logs ~10% of the time
Monitoring Log Health
Logs themselves need monitoring. If logging stops, you lose visibility into your systems.
Metrics to Track
Tracking the right metrics keeps your logging pipeline healthy and gives you early warning before small problems become incidents.
Log ingestion rate measures how many log events your system processes per second. A sudden drop to zero is the clearest signal something is wrong — either the application stopped logging, the agent stopped collecting, or the aggregator stopped accepting events. Set a baseline during normal operations and alert when you deviate significantly. Spikes in ingestion rate usually mean a developer accidentally pushed verbose DEBUG logging or an application is stuck in a retry loop generating excessive errors.
Log volume by service and level helps you identify which services are noisy and whether that noise is signal or garbage. A service emitting 80% of its logs at DEBUG is just burning storage without adding value. Tracking volume by level over time tells you whether your log level policies are actually being followed. When total volume spikes unexpectedly, it usually points to a misconfigured service, an incident generating elevated error rates, or someone who forgot to disable debug logging after a debugging session.
Error rate in logs is your primary production health signal. You want the ratio of ERROR-level entries to total entries, tracked per service and over time windows. A rising error rate is often the first sign of a degradation that has not yet caused a full outage — a downstream service is slowing down, a dependency is returning partial responses, or a configuration drift is causing failures. Correlate error spikes with deployment times to catch bad releases quickly.
Log processing latency measures the time between a log event being emitted and that event being searchable in your aggregator. In a healthy pipeline this is under 10 seconds. When latency climbs, logs arrive late, which means during an incident you might search for an error and find it has not appeared yet, forcing you to wait or check raw server logs. Latency degradation usually comes from aggregator backpressure, network congestion, or the log agent falling behind because volume exceeded what it could handle.
Log agent errors and restarts matter because the log agent is infrastructure you forget exists until it breaks. Agents crash silently, run out of memory under load, or get OOM-killed by Kubernetes when the node is under pressure. When an agent restarts, there is a gap in coverage. Track restart counts and reason codes. If an agent restarts frequently, either increase its resource limits or reduce the log volume it handles.
Alert on Silence
# Prometheus alert for missing logs
- alert: LogIngestionSilence
expr: |
rate(fluentd_input_status_records_total[5m]) == 0
for: 5m
labels:
severity: critical
annotations:
summary: "No logs being ingested from Fluentd"
description: "Fluentd has not sent logs to Elasticsearch in 5 minutes"
When to Use Structured Logging
Use structured logging when:
- Debugging requires cross-referencing multiple log entries
- Requests span multiple services
- You need selective debugging in high-volume APIs
- Audit trails are required for compliance
- You need to correlate logs with traces or metrics
Don’t use structured logging when:
- Simple scripts or one-off utilities where stdout debugging suffices
- Very low-traffic applications where unstructured grep suffices
- Legacy systems where migration cost outweighs benefits
- Development environments where DEBUG-level verbosity is acceptable
Trade-off Analysis
| Aspect | Structured Logging | Plain Text Logging |
|---|---|---|
| Searchability | Field-level queries via log aggregators | grep/string matching only |
| Storage Cost | Higher (JSON overhead per line) | Lower (minimal formatting) |
| Parse Complexity | Zero (machine-readable by default) | Brittle (format changes break parsers) |
| Human Readability | Moderate (requires jq or aggregator UI) | High (direct reading in terminal) |
| Tooling Required | Log aggregator (ELK, Loki, Splunk) | None or basic text tools |
| Correlation | Automatic via shared fields | Manual trace ID injection |
| Performance Impact | Slight overhead for JSON serialization | Minimal |
SLI/SLO/Error Budget Templates for Logging
Log-Based SLI Template
# logging-sli-config.yaml
service: logging-observability
environment: production
slis:
- name: log_ingestion_success_rate
description: "Percentage of emitted logs successfully ingested"
query: |
sum(rate(log_ingested_total[5m]))
/
sum(rate(log_emitted_total[5m]))
- name: log_processing_latency_p95
description: "Time from log emit to searchable in aggregator"
query: |
histogram_quantile(0.95,
sum(rate(log_processing_latency_seconds_bucket[5m])) by (le)
)
- name: log_drop_rate
description: "Dropped logs as a percentage of emitted logs"
query: |
sum(rate(log_dropped_total[5m]))
/
sum(rate(log_emitted_total[5m])) * 100
Log SLO Template
# logging-slo-config.yaml
objectives:
- display_name: "Log Ingestion Availability"
sli: log_ingestion_success_rate
target: 99.5
window: 30d
description: "99.5% of emitted logs should be ingested"
- display_name: "Log Processing Latency"
sli: log_processing_latency_p95
target: 99.0
threshold_ms: 30000
window: 30d
description: "95% of logs should be searchable within 30 seconds"
Error Budget Calculator
# error-budget-calculator.py
def calculate_error_budget(slo_target, window_days=30):
"""
Calculate error budget in minutes for a given SLO target.
Example: 99.5% SLO over 30 days = 21.6 minutes of allowed errors
"""
window_seconds = window_days * 24 * 60 * 60
allowed_errors = window_seconds * (1 - slo_target)
return allowed_errors / 60 # Convert to minutes
# Standard SLO error budgets (30-day window)
slo_budgets = {
"99.0%": calculate_error_budget(0.990), # 432 minutes = 7.2 hours
"99.5%": calculate_error_budget(0.995), # 216 minutes = 3.6 hours
"99.9%": calculate_error_budget(0.999), # 43.2 minutes
"99.95%": calculate_error_budget(0.9995), # 21.6 minutes
"99.99%": calculate_error_budget(0.9999), # 4.32 minutes
}
for slo, budget in slo_budgets.items():
print(f"SLO {slo}: {budget:.2f} minutes error budget")
Multi-Window Burn-Rate Alerting for Log Ingestion
Burn-rate alerts detect when dropped logs consume the ingestion SLO’s error budget faster than expected. Use pipeline counters, not application ERROR log volume: an application can emit many errors while the logging pipeline is healthy.
1-Hour Window Burn-Rate Alert (Fast Burn)
# Burn-rate alerts for logging
groups:
- name: logging-burn-rate
rules:
# Fast burn: 1-hour window, 14.4x burn rate (burns 1% budget in 1 hour)
- alert: LogErrorBudgetFastBurn
expr: |
(
sum(rate(log_dropped_total[1h]))
/
sum(rate(log_emitted_total[1h]))
)
> (1 - 0.999) * 14.4
for: 5m
labels:
severity: critical
category: logging
window: 1h
annotations:
summary: "Log error budget burning fast (1h window)"
description: "Error rate is burning budget {{ $value | humanize }}x faster than sustainable. Budget may be depleted in ~7 hours."
6-Hour Window Burn-Rate Alert (Medium Burn)
# Medium burn: 6-hour window, 6x burn rate (burns 10% budget in 6 hours)
- alert: LogErrorBudgetMediumBurn
expr: |
(
sum(rate(log_dropped_total[6h]))
/
sum(rate(log_emitted_total[6h]))
)
> (1 - 0.999) * 6
for: 30m
labels:
severity: warning
category: logging
window: 6h
annotations:
summary: "Log error budget burning (6h window)"
description: "Error rate is burning budget {{ $value | humanize }}x faster than sustainable. Check for sustained error patterns."
Multi-Window Burn-Rate Alert Set
# Complete burn-rate alert set (multi-window)
- alert: LogErrorBudgetBurnAllWindows
expr: |
(
sum(rate(log_dropped_total[1h]))
/
sum(rate(log_emitted_total[1h]))
)
> (1 - 0.999) * 14.4
or
(
sum(rate(log_dropped_total[6h]))
/
sum(rate(log_emitted_total[6h]))
)
> (1 - 0.999) * 6
for: 5m
labels:
severity: critical
category: logging
annotations:
summary: "Log error budget burning across multiple time windows"
description: |
Multi-window burn-rate alert triggered.
The dropped-log ratio is above the multi-window burn-rate threshold.
Check agent buffers, exporter errors, and aggregator capacity.
SLO Error Budget Dashboard Panels
{
"dashboard": {
"title": "Logging SLO Error Budget",
"panels": [
{
"title": "Error Budget Remaining (30d)",
"type": "gauge",
"targets": [
{
"expr": "(1 - ((sum(rate(log_dropped_total[30d])) / sum(rate(log_emitted_total[30d]))) / (1 - 0.999))) * 100",
"legendFormat": "Budget Remaining %"
}
],
"fieldConfig": {
"defaults": {
"min": 0,
"max": 100,
"thresholds": {
"steps": [
{ "value": 0, "color": "red" },
{ "value": 50, "color": "yellow" },
{ "value": 90, "color": "green" }
]
}
}
}
},
{
"title": "Burn Rate (1h)",
"type": "graph",
"targets": [
{
"expr": "(sum(rate(log_dropped_total[1h])) / sum(rate(log_emitted_total[1h]))) / (1 - 0.999)",
"legendFormat": "Burn Rate"
}
]
},
{
"title": "Dropped Logs (1h)",
"type": "stat",
"targets": [
{
"expr": "sum(increase(log_dropped_total[1h]))",
"legendFormat": "Dropped log records"
}
]
}
]
}
}
OpenTelemetry Integration for Logging
OpenTelemetry hooks into your application at the SDK level and sends logs, traces, and metrics through the same pipeline. No per-vendor instrumentation, no rewriting your log format when you switch backends.
Auto-Instrumentation Setup
// OpenTelemetry collector configuration
import { NodeSDK } from "@opentelemetry/sdk-node";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
import { OTLPLogExporter } from "@opentelemetry/exporter-logs-otlp-http";
import { BatchLogRecordProcessor } from "@opentelemetry/sdk-logs";
const sdk = new NodeSDK({
traceExporter: new OTLPTraceExporter({
url: "http://otel-collector:4318/v1/traces",
}),
logRecordProcessor: new BatchLogRecordProcessor(
new OTLPLogExporter({
url: "http://otel-collector:4318/v1/logs",
}),
),
});
sdk.start();
Correlating Logs with Traces and Metrics
// Inject trace context into log records
import { trace, context } from "@opentelemetry/api";
function emitLog(level: string, message: string, attributes = {}) {
const span = trace.getSpan(context.active());
const record = {
timestamp: new Date().toISOString(),
level,
message,
trace_id: span?.spanContext().traceId,
span_id: span?.spanContext().spanId,
...attributes,
};
logger.emit(record);
}
Benefits of OpenTelemetry for Logging
| Benefit | Description |
|---|---|
| Vendor neutrality | Switch backends without re-instrumenting |
| Unified data model | Logs, traces, and metrics share the same correlation IDs |
| Automatic context prop | Trace context automatically injected into logs |
| Sampling coordination | Sample logs and traces together for consistent debugging |
Observability Hooks for Logging
This section defines what to log, measure, trace, and alert for logging systems themselves.
Log (What to Emit)
| Event | Fields | Level |
|---|---|---|
| Log ingestion started | service, host, agent_version | INFO |
| Log ingestion stopped | service, host, reason | WARN |
| Buffer approaching full | host, buffer_used_percent, buffer_limit | WARN |
| Malformed log detected | host, parse_error_type, sample | WARN |
| Retry attempt | host, destination, attempt, max_attempts | DEBUG |
| Batch sent successfully | host, destination, batch_size, bytes_sent | DEBUG |
| Authentication failure | host, client_ip, reason | WARN |
Measure (Metrics to Collect)
| Metric | Type | Description |
|---|---|---|
log_emitted_total |
Counter | Total logs emitted by service |
log_ingested_total |
Counter | Total logs ingested to aggregator |
log_dropped_total |
Counter | Logs dropped due to errors/full buffers |
log_processing_latency_seconds |
Histogram | Time from emit to searchable |
log_buffer_utilization_percent |
Gauge | Buffer fill percentage |
log_parsing_errors_total |
Counter | Malformed log entries |
log_bytes_sent_total |
Counter | Bytes sent to aggregators |
log_aggregator_queue_depth |
Gauge | Pending logs in aggregator queue |
Trace (Correlation Points)
| Operation | Trace Attribute | Purpose |
|---|---|---|
| Log emit | log.aggregate |
Track logs from emit through aggregation |
| Log parsing | log.parse.status |
Monitor parsing health |
| Log shipping | log.ship.destination |
Track delivery to aggregators |
| Batch processing | log.batch.size |
Monitor batch efficiency |
Alert (When to Page)
| Alert | Condition | Severity | Purpose |
|---|---|---|---|
| Log Silence | No logs received for 5 minutes | P1 Critical | Log pipeline failure |
| High Drop Rate | Drop rate > 1% for 5 minutes | P2 High | Pipeline health |
| Buffer Critical | Buffer > 90% full | P2 High | Prevent data loss |
| Parse Error Spike | Parse errors > 100/min | P3 Medium | Data quality |
| Latency High | Processing latency > 30s p95 | P3 Medium | Performance degradation |
Alerting Hook Template
# logging-observability-hooks.yaml
groups:
- name: logging-observability-hooks
rules:
# Alert on silence - no logs coming in
- alert: LoggingPipelineSilence
expr: rate(fluentd_input_status_records_total[5m]) == 0
for: 5m
labels:
severity: critical
annotations:
summary: "No logs being ingested (Alert on Silence)"
description: "Fluentd/Bit has not sent logs to Elasticsearch in 5 minutes. Either the log pipeline is down or all services have stopped logging."
# Alert on high drop rate
- alert: LoggingDropRateHigh
expr: |
sum(rate(fluentd_output_status_num_errors_total[5m]))
/
sum(rate(fluentd_input_status_records_total[5m])) > 0.01
for: 5m
labels:
severity: high
annotations:
summary: "Log drop rate above 1%"
description: "{{ $value | humanizePercentage }} of logs are being dropped. Check Fluentd/Bit error logs."
# Alert on buffer approaching full
- alert: LoggingBufferCritical
expr: fluentd_buffer_queue_length / fluentd_buffer_limit > 0.9
for: 5m
labels:
severity: high
annotations:
summary: "Log buffer above 90% capacity"
description: "Fluentd/Bit buffer is filling up. Risk of log loss if not addressed."
# Alert on high parsing errors
- alert: LoggingParseErrorSpike
expr: rate(log_parsing_errors_total[5m]) > 100
for: 5m
labels:
severity: warning
annotations:
summary: "High log parsing error rate"
description: "More than 100 parsing errors per minute. Review log format consistency."
# Alert on processing latency
- alert: LoggingProcessingLatencyHigh
expr: |
histogram_quantile(0.95,
sum(rate(log_processing_latency_seconds_bucket[5m])) by (le)
) > 30
for: 10m
labels:
severity: warning
annotations:
summary: "Log processing latency above 30 seconds"
description: "P95 log processing latency is {{ $value }}s. Logs may not be searchable in real-time."
# SLO error budget burn rate
- alert: LoggingErrorBudgetBurningFast
expr: |
(
sum(rate(log_dropped_total[1h]))
/
sum(rate(log_emitted_total[1h]))
) > (1 - 0.999) * 14.4
for: 5m
labels:
severity: critical
annotations:
summary: "Log error budget burning at unsustainable rate"
description: "Error budget is being consumed 14.4x faster than sustainable. Immediate investigation required."
Cost Optimization for Logging Pipeline
Logging costs creep up fast when you are not paying attention. Here is how to keep them under control.
Log Volume Budgeting
# Kubernetes resource quota for logging
apiVersion: v1
kind: ResourceQuota
metadata:
name: logging-budget
namespace: production
spec:
hard:
# Limit log storage per namespace
requests.storage: 100Gi
# Limit Fluentd memory
requests.memory: 2Gi
limits.memory: 4Gi
Cost Optimization Strategies
| Strategy | Impact | Implementation |
|---|---|---|
| Reduce DEBUG in production | 60-80% volume reduction | Runtime level control, feature flags |
| Index only essential fields | 40-60% storage reduction | Field mapping optimization in ES |
| Aggressive ILM policies | 50-70% cost reduction | Move old data to cold/archive tiers |
| Sampling high-volume paths | 90% volume reduction | Deterministic sampling for non-critical |
| Compress before shipping | 30-50% bandwidth savings | gzip compression in log agents |
Architecture for Cost-Effective Logging
graph TB
A[Application] --> B[Fluent Bit Agent]
B --> C{Local Buffer}
C -->|Normal hours| D[Hot Storage - 7 days]
C -->|Off-peak batch| E[Warm Storage - 30 days]
D --> F[Cold Storage - 90 days]
E --> F
F --> G[Archive - 1+ year]
G --> H[Glacier/Blob Storage]
egress costs in multi-region set-ups
Cross-region log shipping is one of those costs that looks small on paper and surprises you on the bill. Every byte leaving a region incurs egress fees, and at logging scale those fees add up fast. The Fluentd grep filter below drops DEBUG and TRACE entries before they cross the boundary. If you need more aggressive cuts, the sampler filter passes through a deterministic 10% of debug logs while keeping every error-level entry intact.
# Fluentd filter to drop low-priority logs before shipping
<filter container.**>
@type grep
<exclude>
key level
pattern /DEBUG|TRACE/
</exclude>
</filter>
# Alternative: drop based on sampling
<filter container.**>
@type sampler
@label @sampled
sample_rate 0.1 # Keep only 10% of debug logs
random_seed 12345
</filter>
Multi-Region Logging Strategies
Global systems need log collection that respects regional boundaries and does not add latency to user-facing paths.
Regional Log Aggregation
When services run across multiple regions, shipping all logs to a central aggregator adds latency on every query and charges you for every cross-region byte. Route logs to regional buckets first, then push only alerting-relevant data to a central cluster. The S3 sink config below routes logs to regional buckets by matching region names in the log tags.
# Regional Fluentd aggregator config
[sinks]
[sinks.s3_regional]
type = "s3"
bucket = "logs-us-east-1"
region = "us-east-1"
[sinks.s3_eu_central]
type = "s3"
bucket = "logs-eu-central-1"
region = "eu-central-1"
# Route logs to regional storage based on source
[transforms.route_by_region]
type = "route"
inputs = ["parse_json"]
route = '''
match /(?i)(eu|europe)/ => "s3_eu_central"
match * => "s3_regional"
'''
Cross-Region Log Correlation
Incidents that span regions require querying every regional index before you know what happened. The TypeScript function below fires all regional queries in parallel via Promise.all, then merges and deduplicates by trace_id. Without parallelism, a slow regional query holds up the entire incident response. Each log entry carries the same trace_id from the originating service, so deduplication by trace_id gives you a complete picture.
// Fan-out query across regions
async function searchLogsAcrossRegions(
query,
regions = ["us-east-1", "eu-central-1"],
) {
const results = await Promise.all(
regions.map((region) =>
elasticsearch.search({
index: `logs-${region}-*`,
body: {
query: {
bool: {
must: [query],
filter: [{ term: { region } }],
},
},
},
}),
),
);
// Merge and deduplicate by trace_id
return results
.flatMap((r) => r.hits.hits)
.reduce((acc, hit) => {
acc[hit._source.trace_id] = hit._source;
return acc;
}, {});
}
Compliance Considerations
| Requirement | Implementation |
|---|---|
| GDPR (EU data) | Regional aggregation, no cross-border log transfer |
| Data residency | Separate indices per region, regional access controls |
| Audit trails | Immutable WORM storage in each region |
| Incident response | Replicate critical error logs to a central alerting index |
Cross-Region Replication Configuration
{
"index": {
"number_of_shards": 3,
"number_of_replicas": 1,
"allocation": {
"include": {
"region": "us-east-1"
}
}
},
"cluster.routing": {
"allocation.awareness.attributes": "region"
}
}
Production Failure Scenarios
| Failure | Impact | Mitigation |
|---|---|---|
| Log aggregation pipeline downtime | No new logs searchable; teams blind to issues | Buffer logs locally; implement retry with backoff; alert on pipeline health |
| Elasticsearch cluster saturation | Log ingestion backs up; logs dropped | Monitor ES cluster health; implement backpressure; use ILM to manage indices |
| Corrupted log data | Searches return incomplete results; debugging misses context | Validate JSON structure at ingestion; use dead-letter queues for malformed logs |
| Sensitive data logged | Security/compliance breach; potential data exposure | Implement redaction middleware; scan logs before storage; educate developers |
| Excessive log volume | Storage costs spike; performance degradation | Implement sampling for DEBUG logs; enforce log level policies; archive aggressively |
| Missing correlation IDs | Cannot trace requests across services | Auto-inject correlation IDs in middleware; reject requests without trace context in high-security paths |
Real-world Failure Scenarios
Scenario 1: Log Data Loss During Incident
What happened: During a production incident, engineers discovered that logs from the primary application server were not being shipped to the central log aggregator. The log forwarder had crashed silently 2 hours prior.
Root cause: No health checks were configured for the log forwarder daemon. The process had exited but the orchestration system did not restart it because it was running as a sidecar rather than a managed service.
Impact: Engineers spent 45 minutes manually accessing individual server logs to piece together the sequence of events, delaying incident resolution.
Lesson learned: Monitor log forwarder processes and shipper queues. Implement heartbeat logging so missing heartbeats trigger an alert. Ship logs to multiple destinations for critical services.
Scenario 2: Structured Logging Breaking Search Dashboards
What happened: After migrating from plain-text to structured JSON logging, the Kibana dashboard used by the operations team stopped displaying log events. The team was flying blind for 3 hours until the issue was diagnosed.
Root cause: The Kibana index pattern was configured to look for a message field as the primary text field. Structured logs used field names like msg and event_text, so no events matched the default search.
Impact: All monitoring dashboards showed empty results. A customer-impacting database slowdown went undetected for longer than necessary.
Lesson learned: Validate dashboard queries against a test environment before migrating logging formats. Ensure field naming conventions match across the log pipeline and dashboards. Maintain backwards compatibility during format transitions.
Common Pitfalls / Anti-Patterns
Anti-Patterns to Avoid
1. Logging Everything at DEBUG in Production
DEBUG-level logging in high-throughput services generates gigabytes per hour. Use sampling for debug scenarios, or enable DEBUG selectively via feature flags for specific request IDs.
2. Plain Text Logging with String Concatenation
// Bad: Cannot search, parse, or aggregate
logger.info("User " + userId + " purchased " + item);
// Good: Structured, searchable, aggregatable
logger.info("User purchased item", { userId, itemId, itemName, price });
3. Missing Trace Context Propagation
Logs without correlation IDs are useless for tracing requests across services. Always propagate trace_id through HTTP headers, database connections, and message queues.
4. Logging Sensitive Data
Never log passwords, full tokens, credit card numbers, or PII. Implement redaction at the logger level, not the application level, to catch mistakes.
5. Synchronous Logging to Network Storage
Writing logs synchronously to a remote log server adds latency to every operation. Use async logging with local buffering and background shipping.
6. No Log Retention Policy
Without retention policies, storage costs grow unbounded. Define hot/warm/cold/archive tiers and automate data lifecycle management.
7. Logs as the Only Observability Signal
Relying solely on logs for debugging is insufficient at scale. Combine logs with metrics and traces for complete observability.
Observability Checklist
Key Log Metrics
- Log ingestion rate (logs/second) by service and level
- Log volume by service, level, and environment
- Error rate in logs (ERROR level count over time)
- Log processing latency (time from log emit to searchable)
- Log agent errors and restarts
- Storage utilization per index
Logs You Should Have
- Request logs with trace_id, user_id, method, path, status, duration_ms
- Authentication events (login attempts, failures, token refreshes)
- Business events (orders, payments, registrations) with entity IDs
- Database query logs for slow queries (>100ms threshold)
- External API call logs with request/response timing
- Background job start/complete/fail logs with job IDs
- Health check and readiness probe logs
- Configuration change logs (who changed what when)
Alerts You Need
- No logs received from a service for >5 minutes (Alert on Silence pattern)
- Error rate spike above baseline (unexpected errors)
- Log volume anomaly (sudden drop or spike)
- Log processing latency >30 seconds
- Elasticsearch cluster health degraded (yellow/red)
- Log agent restart detected
Security Checklist
- No passwords, API keys, or secrets in log output
- Credit card numbers, CVV, SSN never logged
- Authorization tokens logged as type + last 4 chars only (e.g., “Bearer ***abc123”)
- PII fields identified and redacted in redaction middleware
- Log access requires authentication and is audited
- Log aggregation pipeline uses TLS in transit
- Elasticsearch access restricted to authorized personnel
- Log retention complies with data retention policies
- Sensitive data cannot be searched in Kibana/ES by unauthorized users
Quick Recap Checklist
- Emit structured logs with consistent fields and severity levels.
- Propagate correlation IDs so a request can be followed across services.
- Redact secrets and sensitive personal or payment data before logs are written.
- Use sampling and retention policies to control volume and storage cost.
- Monitor ingestion, processing delay, storage health, and alert on silence.
Interview Questions
timeframe AND service.name: checkout-service AND "checkout failed". If no direct match, search for errors in the checkout service within the time window, then trace back via correlation IDs to find the root cause service.
doc_values: false on fields that are only used for filtering. Finally, enforce log volume budgets per service to prevent any single service from overwhelming the cluster.
x-b3-traceId in Zipkin, traceparent in W3C trace context) through every service call. Each service creates a span with the incoming trace ID and its own span ID, creating a parent-child tree of operations.
metric.name: jvm_memory_used AND metric.area: heap AND increase(metric.value[1h]) > threshold. Set a PagerDuty alert when heap usage exceeds 80% sustained for 15 minutes or when the GC reclaim rate falls below the allocation rate. Correlate with your application logs to identify which code paths are allocating the most objects.
Further Reading
- Metrics, Monitoring & Alerting — Learn how to turn service signals into useful alerts.
- Distributed Tracing — Follow requests across service boundaries and connect spans with logs.
- Secrets Management — Protect credentials and sensitive values across backend systems.
- Elasticsearch Guide: Logstash — Log processing pipelines
- Fluentd Documentation — Official documentation for log aggregation
- OpenTelemetry Logging SDK — Vendor-neutral logging specification and SDK documentation
- Loki Documentation — Log aggregation designed to work with Grafana
- Python structlog Documentation — Structured logging for Python applications
- Pino JavaScript Logger — Structured logging for Node.js
- Distributed Tracing with OpenTelemetry — Connect logs, traces, and metrics
- Log Observability and SLOs — Guidance for log-based SLOs and error budgets
- RFC 5424: The Syslog Protocol — The standard syslog message format
Conclusion
Key Takeaways:
- Structured JSON logs enable efficient searching and aggregation
- Correlation IDs connect logs across service boundaries
- Log levels filter noise: DEBUG for development, ERROR/WARN/INFO for production
- Never log sensitive data; always implement redaction
- Monitor your monitors: log aggregation needs its own observability
- Retention policies prevent unbounded storage growth
Copy/Paste Checklist:
# Verify structured logging format
grep -c '"timestamp".*"level".*"message".*"service"' /var/log/app.json
# Find logs for specific trace
grep '"trace_id":"abc123"' /var/log/app.json
# Count errors by service
jq 'select(.level == "ERROR") | .service' /var/log/app.json | sort | uniq -c
# Alert on log silence (Prometheus)
- alert: LogIngestionSilence
expr: rate(fluentd_input_status_records_total[5m]) == 0
for: 5m
labels:
severity: critical
# Redaction function (TypeScript)
const sensitiveFields = ['password', 'token', 'secret', 'creditCard', 'ssn'];
function redact(obj) {
return Object.fromEntries(
Object.entries(obj).map(([k, v]) =>
sensitiveFields.some(f => k.toLowerCase().includes(f)) ? [k, '[REDACTED]'] : [k, v]
)
);
}
Good logging practices pay off when you need them most: debugging production issues at 2am. Structured logs with correlation IDs let you trace requests across service boundaries. Appropriate log levels keep noise manageable. Retention policies balance cost with compliance requirements.
Start with JSON structured logging in your applications. Add correlation ID propagation early. Build log aggregation before you need it, not during an incident.
For deeper observability, combine logging with the Metrics, Monitoring & Alerting and Distributed Tracing practices covered in our other guides. These three pillars work together: logs show you what happened, metrics show you patterns, and traces show you why it happened.
Category
Related Posts
Alerting in Production: Building Alerts That Matter
Build alerting systems that catch real problems without fatigue. Learn alert design principles, severity levels, runbooks, and on-call best practices.
The Observability Engineering Mindset: Beyond Monitoring
Move beyond monitoring with structured logs, metrics, and traces. Learn how SLOs, sampling, OpenTelemetry, and team practices improve incident debugging.
Metrics, Monitoring, and Alerting: From SLIs to Alerts
Learn the RED and USE methods, SLIs/SLOs/SLAs, and how to build alerting systems that catch real problems. Includes examples for web services and databases.