Load Balancing: Traffic Control for Modern Infrastructure

Learn how load balancers distribute traffic across servers, the differences between L4 and L7 load balancing, and when to use software vs hardware solutions.

published: reading time: 44 min read author: GeekWorkBench updated: June 17, 2026
Quick Summary

Load balancers distribute connections and requests across backend servers, with routing shaped by protocol, health checks, and failure behavior. This guide compares L4 and L7 routing, session affinity, regional DNS, rate limiting, and deployment patterns across traditional and Kubernetes environments. Use the examples and checklists to choose an approach, protect backend connections, monitor uneven load, and account for delays during failover and draining.

Load Balancing: The Traffic Controller of Modern Infrastructure

Introduction

Load balancing exists because every system has a breaking point. Push enough traffic toward a single server and it buckles. The load balancer sits between users and your server pool, spreading requests around so nothing collapses.

I think of load balancers as air traffic control for your network. They do not just route packets. They make decisions based on real-time conditions, health status, and configured policies. Without them, scaling beyond a handful of servers becomes a nightmare of manual failover and prayer.

This article compares Layer 4 and Layer 7 routing, health checks, session affinity, and global load balancing. It also covers how balancers support rate limits, circuit breakers, and safer deployments.

Layer 4 vs Layer 7 Load Balancing

Network engineers talk about Layer 4 (transport) and Layer 7 (application) load balancing. The layer tells you how deep into the network stack the load balancer inspects when making routing decisions.

Routing decisions

Layer 4 Load Balancing Flow

Layer 4 load balancers operate at the transport layer. They route based on source and destination IP addresses plus port numbers, without looking inside the actual request. This makes them faster and able to handle more throughput since parsing happens at a lower level.

Picture L4 as a postal sorter that reads the network address, not the contents of the request. It avoids HTTP parsing, which can reduce per-connection overhead; actual throughput and latency depend on the implementation, hardware, and traffic pattern.

Layer 4 works well for TCP-based protocols like databases or SSH connections, or any protocol where raw throughput matters more than content inspection.

graph LR
    subgraph "Layer 4"
        L4Req[IP + Port] --> L4Dec[Routing Decision]
    end

Layer 7 Load Balancing Flow

Layer 7 load balancers operate at the application layer. They can inspect HTTP headers, URLs, cookies, and request bodies. This opens up sophisticated routing based on what the user actually requested.

With L7, you can route based on URL path, send API requests to one cluster and static assets to another, or direct mobile users to a different backend. You can also terminate TLS at the load balancer, inspecting encrypted traffic before forwarding it.

The cost is higher resource usage. Parsing HTTP is more expensive than reading IP and port numbers. Modern L7 balancers are heavily optimized, but this distinction matters when designing systems that need maximum performance.

graph LR
    subgraph "Layer 7"
        L7Req[HTTP Headers<br/>URL Path<br/>Cookies] --> L7Dec[Routing Decision]
    end

Backend operation and availability

Software vs Hardware Load Balancers

Back in the day, load balancers were expensive hardware appliances. Companies like F5 sold dedicated network devices for tens of thousands of dollars, with specialized ASICs for packet processing.

The industry shifted toward software. HAProxy, Nginx, and cloud offerings like AWS ALB or Google Cloud Load Balancing commoditized load balancing. You can deploy capable software balancers on commodity hardware or use managed cloud services.

Software wins on flexibility and cost. You can modify routing logic with code changes, integrate with container orchestration, and scale by deploying more instances. Hardware appliance licensing makes less sense in cloud environments.

But hardware still has a place. Regulated industries sometimes require dedicated appliances for compliance. Extremely high-throughput environments, like major video streaming platforms, still use custom ASIC-based solutions. For most web applications though, software approaches work fine.

Sticky Sessions

Session affinity creates problems. If User A logs into Server 1 and their next request goes to Server 2, that server has no memory of the login. The user appears logged out.

Sticky sessions route a particular user’s requests to the same backend server. The load balancer tracks which client maps to which server, using cookies, client IP, or some other identifier.

Cookie-based sticky sessions insert a tracking cookie that identifies the target server. IP-based affinity hashes the client IP to always return the same backend. Header-based approaches use a custom header set by an upstream service.

Sticky sessions cause their own headaches though. They complicate maintenance windows since you cannot take down a server without disconnecting active users. They make horizontal scaling harder because you cannot freely redistribute load. Many applications work better with session state stored in a distributed cache like Redis.

Health Checks and Failover

Load balancers continuously check that backend servers can handle traffic. Health checks run at configurable intervals, testing whether each server responds correctly. A server that fails too many health checks gets marked unhealthy and removed from rotation.

Health checks range from simple TCP connection tests to HTTP requests with expected response validation. Deeper checks catch more application failures but add probe load and can depend on downstream services. Choose a check that answers whether the backend should receive traffic, and avoid making every probe depend on a fragile database or third-party service.

graph TD
    LB[Load Balancer] --> HC[Health Checker]
    HC -->|TCP ping| S1[Server 1]
    HC -->|TCP ping| S2[Server 2]
    HC -->|TCP ping| S3[Server 3]
    S1 -->|Healthy| S1Status[✓ Healthy]
    S2 -->|Timeout| S2Status[✗ Unhealthy]
    S3 -->|Healthy| S3Status[✓ Healthy]
    S2Status -->|Remove| Pool[Removed from pool]

When a server fails health checks, the load balancer stops routing new traffic to it after the configured failure threshold. Existing connections may continue, drain, or be terminated depending on the product and policy. A recovered server usually needs to pass a configured number of checks before it rejoins the pool.

TLS and implementation choice

SSL Termination

Handling HTTPS at the load balancer layer has practical benefits. TLS termination means the load balancer decrypts incoming HTTPS traffic and can inspect HTTP when it operates at L7. It can then forward plain HTTP or re-encrypt traffic to the backend. Offloading TLS reduces cryptographic work on application servers, but makes the load balancer a trust boundary.

You manage certificates in one place rather than on every backend. Backend communication can use plain HTTP, reducing CPU overhead on application servers. Your load balancer can inject security headers, rewrite URLs, and perform other transformations on decrypted traffic.

TLS termination creates a trust boundary at the load balancer. If it forwards plain HTTP, the backend leg is unencrypted even on a private cloud network. Re-encrypt that leg when the network is not trusted or policy requires it; TLS passthrough keeps encryption to the backend but limits HTTP inspection and routing.

Choosing the Right Load Balancer

Choosing a load balancer depends on your requirements. For simple HTTP traffic with moderate scale, Nginx or HAProxy on a couple of virtual machines works well. They are battle-tested, documented, and free.

Cloud providers offer managed load balancers that integrate with their ecosystems. AWS Application Load Balancer handles L7 routing with rule-based decisions. Network Load Balancer provides ultra-low-latency L4 forwarding for TCP workloads. These services scale automatically and reduce operational overhead.

If you run Kubernetes, an Ingress controller or Gateway implementation can handle external HTTP(S) routing. Choose a maintained implementation that fits your cluster; the community ingress-nginx controller was retired in March 2026, so existing users should plan a migration.

For microservices, service meshes like Istio or Linkerd include load balancing as part of their service-to-service communication layer. These handle traffic shaping, circuit breaking, and retries alongside basic load distribution.


When to Use Each Approach

When to Use Load Balancing

Load balancing is essential when:

  • You run multiple backend servers serving the same application
  • You need high availability (single server failure should not cause outage)
  • You want to scale horizontally by adding more servers
  • You need to perform maintenance without downtime
  • Traffic volume exceeds what a single server can handle
  • You want to protect against server overload and failures

When to Use Layer 4 (L4) Load Balancing

L4 may be the right choice when:

  • You need high throughput with low latency, after measuring the selected implementation
  • You are load balancing TCP/UDP protocols beyond HTTP
  • You do not need to inspect application-layer data
  • You are routing database connections or streaming data
  • Raw performance matters more than routing intelligence

When to Use Layer 7 (L7) Load Balancing

L7 is a better fit when:

  • You need content-based routing (URL path, headers, cookies)
  • You want to terminate TLS at the load balancer
  • You need to implement sticky sessions
  • You are serving multiple applications on the same IP
  • You want to rewrite URLs or redirect requests

Global Server Load Balancing

When an application spans regions, a regional traffic policy can help reduce latency and route around outages. DNS-based GSLB can return a preferred healthy region using location, health, or load signals, but the answer may reflect a recursive resolver rather than the end user and can remain cached until its TTL expires.

graph TD
    UserAP[User - APAC] --> DNS[GSLB DNS]
    UserEU[User - EU] --> DNS
    UserUS[User - US] --> DNS
    DNS -->|Return APAC IP| LB_APAC[Load Balancer - Tokyo]
    DNS -->|Return EU IP| LB_EU[Load Balancer - Frankfurt]
    DNS -->|Return US IP| LB_US[Load Balancer - Virginia]

Anycast vs GSLB

Anycast advertises the same IP address from multiple locations. BGP routes packets according to network paths and routing policy, which often sends traffic to a nearby site but does not guarantee the geographically closest one. This selection happens at the network layer, without application-level health or routing decisions.

DNS-based GSLB can apply weights and health-check results to DNS answers, and some providers offer latency or location policies. Health checks may probe application endpoints, but the client still connects to the returned address and DNS caching can delay a change. Examples include Amazon Route 53, Azure Traffic Manager, and Google Cloud’s global load-balancing services.

A design can combine Anycast and DNS-based GSLB, but each has different failure and routing behavior. Anycast selects a network path; GSLB can apply DNS policies and health signals, subject to resolver caching. Test regional failover from multiple networks and confirm clients retry when an address becomes unhealthy.

Health Checking Across Regions

GSLB health checks must account for entire regions going dark. A health check that only verifies server TCP connectivity misses regional outages. Proper GSLB checks server-level health, regional health, latency thresholds, and capacity limits.

Most GSLB implementations probe from multiple vantage points to distinguish between a slow region and a slow internet path. Use probes from multiple locations and define a failure threshold that fits the service’s recovery goals; a single probe location can mistake a local network problem for a regional outage.


Rate Limiting at the Load Balancer

An L7 load balancer or reverse proxy can enforce request limits before traffic reaches the application. That helps control abusive clients and runaway scripts, but it does not replace application authorization or upstream DDoS protection.

Leaky Bucket and Token Bucket Algorithms

NGINX’s limit_req module uses a leaky-bucket method to smooth request processing. Token bucket is a related but distinct algorithm: each client receives tokens at a steady rate, each request consumes one, and accumulated tokens allow a configured burst. When the bucket is empty, requests are rejected or delayed.

# NGINX limit_req leaky-bucket rate limiting
limit_req_zone $binary_remote_addr zone=api_limit:10m rate=100r/s;

server {
    location /api/ {
        limit_req zone=api_limit burst=200 nodelay;
    }
}

Sliding Window Counter

A sliding-window counter tracks or estimates requests over a moving time interval, smoothing the abrupt boundary of a fixed window. Burst behavior depends on the implementation; a sliding-window log can enforce a stricter limit but uses more memory. Choose the variant based on how much burst traffic your service can handle.

# HAProxy sliding-window counter example
backend per_ip_rates
    stick-table type ip size 100k expire 60s store http_req_rate(60s)

frontend app
    bind :80
    mode http
    http-request track-sc0 src table per_ip_rates
    http-request deny deny_status 429 if { sc_http_req_rate(0) gt 100 }
    default_backend app_servers

Limiting by Different Keys

Rate limits can use different keys, but each has trade-offs:

Key Use Case
Source IP Single user/clients behind NAT
API Key Authenticated API consumers
Header (User-Agent) Weak bot signal; easy to spoof
JWT claim Per-user service tier limits

Rate limiting and reputation features vary by provider and product. For example, AWS WAF can apply rate-based rules to an Application Load Balancer, while Google Cloud Armor provides rate limiting policies with supported load balancers. Check the selected service’s limits and network-layer DDoS protections separately.


Circuit Breaker Pattern Integration

Health checks remove failed backends from rotation. Circuit breakers work one level deeper. They prevent your load balancer from hammering a struggling backend that has not fully failed but is showing signs of strain.

How Circuit Breakers Interact with Load Balancers

A well-designed circuit breaker sits between the load balancer and the backend, monitoring error rates and latency. When a backend starts returning errors or exceeding latency thresholds, the circuit breaker “opens” and trips, immediately returning failures without forwarding the request.

stateDiagram-v2
    Closed --> Open : Error threshold exceeded
    Open --> HalfOpen : Cool-down period elapsed
    HalfOpen --> Closed : Probe request succeeds
    HalfOpen --> Open : Probe request fails

The load balancer still performs health checks on an open circuit breaker. Once the backend recovers and passes enough health checks, the circuit breaker allows traffic through again.

Implementation Approaches

For Kubernetes deployments, service meshes like Istio implement circuit breaking at the sidecar proxy level. You configure outlier detection on your DestinationRule:

apiVersion: networking.istio.io/v1alpha3
kind: DestinationRule
metadata:
  name: backend-service
spec:
  host: backend-service
  trafficPolicy:
    outlierDetection:
      consecutiveGatewayErrors: 5
      interval: 30s
      baseEjectionTime: 30s
      maxEjectionPercent: 50

For traditional deployments, libraries like Hystrix (Java), Resilience4j, or Polly (.NET) implement circuit breakers in your application code. The load balancer health checks still matter — they handle complete failures while circuit breakers handle degraded states.


Canary and Blue-Green Deployments

Load balancers make deployment strategies like canary releases and blue-green deployments possible without downtime. Instead of replacing all servers at once, you shift traffic gradually and monitor for problems.

Blue-Green Deployment

Blue-green keeps two identical environments. The current production environment (blue) handles live traffic while the new version (green) sits idle. When you are ready to deploy, you shift all traffic from blue to green at the load balancer level with a single routing rule change.

graph LR
    subgraph "Blue Environment (Current)"
        B1[Server 1]
        B2[Server 2]
    end
    subgraph "Green Environment (New)"
        G1[Server 1']
        G2[Server 2']
    end
    LB[Load Balancer] -->|Route to Blue| B1
    LB -->|Route to Blue| B2
    LB -.->|Ready to switch| G1
    LB -.->|Ready to switch| G2

If something goes wrong, you flip traffic back to blue instantly. The old environment stays warm during the deployment window so rollback never requires rebuilding anything.

Canary Deployment

Canary releases shift a small percentage of traffic to the new version while the rest stays on the current version. You route 5% of users to the new build, monitor error rates and latency, and gradually increase the percentage.

# HAProxy canary configuration
backend canary_backend
    server new_app_1 10.0.1.101:8080 weight 5  # 5% of traffic
    server current_app_1 10.0.1.201:8080 weight 95

# Increase weight gradually as confidence builds
# 5% -> 10% -> 25% -> 50% -> 100%

Canary deployments require more sophisticated monitoring. You need to compare error rates and latency between the canary and production groups. If the canary shows 1% higher error rate while serving 5% of traffic, that is a real signal worth investigating. Load balancers that export metrics to Prometheus or Datadog make this kind of comparison straightforward.

Load Balancer Role in Deployment Safety

The load balancer is the control point for both strategies:

  • Instant rollback: Change backend weights to route 100% back to the old version
  • Gradual rollout: Shift traffic incrementally to the new version
  • Health monitoring integration: Automatically pause rollout if backend error rates spike
  • Connection draining: Move users off servers being decommissioned gracefully

Both strategies depend on having enough capacity to run two environments simultaneously. With auto-scaling, provision capacity for both environments, shift traffic, verify application and data compatibility, then scale down the old environment.


Kubernetes Ingress and Service Mesh Load Balancing

In Kubernetes, a Service provides a stable endpoint for a set of Pods. By default, kube-proxy runs as a separate node-level component and programs packet-forwarding rules for Service traffic; some network implementations provide their own replacement. This service proxying is distinct from the external traffic handled by an Ingress or Gateway implementation.

Ingress Controllers

For external traffic entering the cluster, Ingress resources define how HTTP/HTTPS routing works. An Ingress controller implements those rules; TLS termination and weighted traffic features depend on the controller. The community ingress-nginx project was retired on March 24, 2026 and no longer receives fixes. Existing deployments continue to run, but teams should plan migration to Gateway API or a maintained controller. Do not confuse the community project with F5 NGINX Ingress Controller. See the Kubernetes retirement notice.

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: app-ingress
spec:
  rules:
    - host: api.example.com
      http:
        paths:
          - path: /users
            pathType: Prefix
            backend:
              service:
                name: users-service
                port:
                  number: 80
          - path: /products
            pathType: Prefix
            backend:
              service:
                name: products-service
                port:
                  number: 80

Ingress controllers commonly handle HTTP(S) routing by host and path, and may support TLS termination and weighted traffic features depending on the controller. The Kubernetes Ingress API remains supported; Gateway API is another option with a broader routing model.

Service Mesh Load Balancing

Service meshes like Istio and Linkerd move load balancing from the application layer to the sidecar proxy running alongside each pod. Every outbound request goes through the sidecar, which makes routing decisions based on service mesh policies.

With a service mesh, you get per-request routing, chaos injection for testing, automatic retries with backoff, and fine-grained traffic shaping. Your application code has no idea any of this is happening.

apiVersion: networking.istio.io/v1alpha3
kind: VirtualService
metadata:
  name: reviews
spec:
  hosts:
    - reviews
  http:
    - route:
        - destination:
            host: reviews
            subset: v1
          weight: 90
        - destination:
            host: reviews
            subset: v2
          weight: 10

The sidecar proxy handles health checking, load balancing algorithms, and circuit breaking transparently. Your application code has no awareness of which version is handling any given request.

Comparison: Ingress vs Service Mesh

Aspect Ingress Controller Service Mesh
Scope North-south traffic (into cluster) East-west traffic (within cluster)
Deployment Cluster-level controller Per-pod sidecars
Protocol HTTP/HTTPS primarily HTTP, gRPC, TCP, or TLS features vary by mesh
Complexity Lower Higher
Use Case Exposing services externally Service-to-service communication

Use both only when the deployment needs both roles: an ingress or Gateway implementation for north-south traffic and a service mesh for east-west policies. A Kubernetes Service alone may be enough for simpler cluster-internal traffic.


Topic-Specific Deep Dives

When Not to Rely on Load Balancing Alone

Load balancing distributes traffic across servers, but it introduces its own latency and failure modes. Knowing when it falls short helps you design for real-world outages.

Health checks run at intervals, typically every 5-30 seconds. When a backend fails, there is a window where requests still route to the dead server before the health check detects the failure and removes it from the pool. Detection time depends on probe interval, timeout, failure threshold, and provider behavior; it is not an instant failover guarantee. DNS-based failover adds resolver caching delays, while database replication and leader election need their own split-brain protections. Choose these mechanisms against a stated recovery objective.

Load balancers route requests, they do not synchronize state. If your application maintains server-local state that cannot be externalized to shared storage, taking down a server means losing that state for any user whose requests got routed there. This is why sticky sessions create fragility. Store session state in Redis, Memcached, or a distributed database instead.

Load balancers handle traffic distribution well, but they cannot conjure server capacity out of nowhere. If you expect sudden traffic spikes, you need auto-scaling triggered by load metrics, not reactive scaling after servers are already overwhelmed. Combine load balancing with container orchestration that spins up new instances before demand exceeds capacity. The load balancer then routes to the newly added instances.

Load balancing does not replace state replication, capacity planning, or application-level recovery. Pair it with health checks, shared session storage where needed, and scaling or circuit-breaker policies that match the service’s failure modes.


Trade-off Analysis

Layer 4 vs Layer 7 Trade-offs

Aspect Layer 4 Layer 7
Performance Higher throughput, lower latency Higher latency, more CPU usage
Routing Intelligence IP + port only Content-based (URL, headers, cookies)
TLS Handling Often forwards TLS without HTTP inspection; termination depends on the product Can terminate TLS and inspect HTTP when configured
Protocol Support Any TCP/UDP protocol HTTP/HTTPS primarily
Memory Usage Lower Higher (stateful inspection)
Complexity Simpler configuration More configuration options

Software vs Hardware Trade-offs

Aspect Software Load Balancers Hardware Load Balancers
Cost Low / open source High ($50K+ appliances)
Flexibility Easily modified, code changes Fixed functionality, firmware updates
Scalability Horizontal (add more VMs) Vertical (bigger appliance) or clustering
Performance Good for most web apps Extreme throughput (ASICs)
Compliance Evaluate controls and evidence for the deployment Evaluate controls and evidence for the deployment
Operations Requires maintenance Managed appliance support

Sticky Sessions Trade-offs

Aspect Without Sticky Sessions With Sticky Sessions
Load Distribution No affinity constraints, but not always equal work Can be uneven when clients have different behavior
Session State Must use shared storage (Redis) Simpler if server-local state is acceptable
Scaling Easy horizontal scaling Harder — cannot freely move clients
Failure Handling User sessions survive server failure User loses session if sticky server dies
Complexity Requires external session store Simpler initially

Health Check Depth Trade-offs

Check Type What It Catches Cost / Trade-off
TCP Connect Port open, basic reachability Fast, low overhead, but misses app failures
HTTP Request Application responding More accurate, slightly higher latency
Deep HTTP Response body validation Catches more failures, highest latency
Custom Script Business logic verification Most accurate, operational complexity

Rate Limiting Algorithm Trade-offs

Algorithm Burst Handling Memory Usage Predictability
Token Bucket Allows bursts Low Variable
Sliding Window Smoother limits; behavior depends on variant Higher More consistent than fixed windows
Fixed Window Allows bursts Lowest Inconsistent at boundaries
Leaky Bucket Smooths output Low Consistent

Observability and Security

Metrics

  • Request rate (requests per second by backend)
  • Response latency (p50, p95, p99 per backend)
  • Backend health status (healthy/unhealthy/draining)
  • Active connections per backend
  • Connection rate (new connections per second)
  • TLS handshake rate and latency (if terminating SSL)
  • Backend error rate (5xx responses from backends)
  • Health check success/failure rate
  • Backend response time trends

Logs

  • Backend health check failures with details
  • SSL handshake failures (certificate errors, protocol mismatches)
  • Connection timeouts from backends
  • Routing decisions for L7 (which rule matched)
  • Backend server added/removed events
  • Rate limiting events
  • Connection errors and disconnections

Alerts

  • Any backend unhealthy for more than 30 seconds
  • All backends unhealthy (complete outage)
  • Request error rate exceeds threshold
  • p99 latency exceeds service level objective
  • Active connections approach limits
  • Health check failure rate increases
  • Unusual traffic patterns (potential attack)

Security Checklist

  • Restrict access to load balancer management interface
  • Use TLS between the load balancer and backends when the network is not trusted or policy requires encryption
  • Implement access controls on health check endpoints
  • Monitor for traffic anomalies indicating attack
  • Use private VIPs for internal load balancers (not internet-facing)
  • Rotate SSL certificates if terminated at load balancer
  • Implement rate limiting at the load balancer layer and trust forwarded client IP headers only from known proxies
  • Log all administrative changes to load balancer config
  • Use network ACLs to restrict which clients can reach load balancer
  • Enable audit logging for compliance
  • Use provider or upstream DDoS protection for volumetric attacks; the load balancer alone is not a DDoS control
  • Verify backend servers are not directly accessible (all traffic through LB)

Production Failure Scenarios

Failure Impact Mitigation
Load balancer itself fails Complete service outage Deploy redundant load balancers; use VRRP/keepalived
Backend server fails silently Requests routed to dead server; errors for users Implement health checks; remove failed servers quickly
Health check misconfiguration False positives remove healthy servers Use multiple check types; set appropriate thresholds
Sticky session overload One server gets all traffic; cascade failure Minimize sticky sessions; use session storage (Redis)
SSL termination bottleneck Load balancer CPU maxes out on encryption Use SSL offloading hardware; scale horizontally
Connection exhaustion No new connections accepted; service hangs Monitor connection counts; implement connection limits
ARP/cache issues with VIP Traffic routing breaks; intermittent failures Use keepalived with proper priority; monitor ARP tables
Misconfigured routing rules Traffic goes to wrong backend; data issues Test rules in staging; implement gradual rollout

Common Pitfalls / Anti-Patterns

Single Point of Failure

A single load balancer can become a single point of failure. If it stops accepting traffic, healthy backends are unreachable through that path. Backend-only monitoring will not detect this, so monitor the load balancer and the client-facing route as well.

graph TD
    A[Client] --> B[Single LB]
    B --> C[Server 1]
    B --> D[Server 2]

Load balancers are often deployed as a single instance because they were considered reliable. But software bugs, memory corruption, kernel panics, network partitioning, and operator misconfiguration can all kill a load balancer. Hardware load balancers fail too, sometimes due to power supply issues or fan failures in dense rack-mounted appliances.

Active-passive uses a standby that monitors the active via VRRP. When the active fails, the standby takes over the virtual IP within seconds. keepalived implements this on Linux. Active-active runs multiple load balancers simultaneously, distributing traffic across them for both redundancy and capacity.

Managed cloud load balancers can provide multi-zone redundancy when configured for the provider’s supported availability model. Verify the selected service’s zone coverage, health behavior, quotas, and failure guarantees. This reduces the work of operating the proxy layer but can limit control and increase provider coupling.

For on-premises deployments, evaluate the cost of a standby or active-active pair against the service’s recovery target and the impact of an outage.

Ignoring Health Check Tuning

Health check tuning is often an afterthought, but getting it wrong in either direction causes real problems. Too aggressive and you bounce healthy servers in and out of rotation. Too lenient and failed servers stay in the pool too long, accumulating errors for users.

When health checks run every second with a 1-second timeout and only one failure required to remove a server, transient network hiccups or brief CPU spikes cause unnecessary churn. A server that recovers within 2 seconds still gets removed and must pass several consecutive checks to rejoin. Users see errors during these transitions. Flapping generates load spikes on remaining backends as they absorb traffic from ejected servers.

A 30-second interval with a five-failure threshold can leave a failed backend in rotation for a long time; exact detection depends on probe scheduling, timeouts, and provider semantics. Calculate the delay from the actual configuration and compare it with the service recovery objective.

As a starting experiment, try a 10-second interval, 3-second timeout, three consecutive failures to remove a backend, and two successes to restore it. That can detect a failure in roughly 20–30 seconds, depending on when probes run and how the load balancer counts failures. Tune the values to your recovery objective and test for flapping under load.

Interval controls how often checks run. Timeout determines how long to wait for a response. Failure threshold sets how many consecutive failures trigger removal. Success threshold determines how many passing checks re-enable a server. Some systems let you configure check types with different weights, running cheap TCP checks frequently and expensive HTTP checks occasionally.

Test your health check configuration under load. A server under heavy request pressure may briefly slow responses enough to trigger aggressive health checks. Monitor your flapping rate and adjust thresholds until you eliminate unnecessary server churn.

Not Planning for Connection Draining

When you need to take a backend server offline for maintenance or deployment, the naive approach is to just stop the server. The load balancer sends new traffic elsewhere, but existing connections break mid-request. Users see errors. If you are doing rolling deployments where you replace servers one by one, this means every server replacement creates a brief outage.

Connection draining allows existing requests on a server being taken offline to complete while preventing new connections from being routed to it. The load balancer stops sending new traffic to the target server but allows in-flight requests to finish. Once all existing connections close or the drain timeout expires, the server can be safely shut down.

Set the drain period to cover the requests and connections the service is expected to finish, while respecting application and platform timeouts. Long uploads or streaming requests may need a different policy from short HTTP requests. During draining, the pool has less capacity, so account for that in rollout limits.

These controls are product-specific. NGINX’s worker_shutdown_timeout limits graceful worker shutdown; it is not a backend drain directive. HAProxy can put a server into drain state, which removes it from normal balancing while preserving health checks and existing connections. Cloud load-balancer drain periods vary by service. In Kubernetes, coordinate readiness changes, any preStop hook, and terminationGracePeriodSeconds; a rolling update alone does not guarantee every in-flight request will finish.

Connection draining reduces failed in-flight work during deployments, but it does not by itself guarantee zero downtime. Mark old targets as draining, wait for the configured drain condition or deadline, and ensure the application can recover from clients that reconnect. Without draining, a deployment can interrupt active requests or long-lived connections.

Test draining with long requests and persistent connections before relying on it during a production rollout.

Overusing Sticky Sessions

Sticky sessions can help legacy applications that keep session state in process, but they add operational constraints. Teams often remove the affinity once session state can live in shared storage.

When requests from the same user always go to the same server, you cannot freely add or remove servers from the pool. If Server 3 is handling 1000 sticky users, removing it disconnects all of them. Your load balancer cannot redistribute those users evenly because their sessions are bound to specific servers. This forces you to keep servers running even when capacity exceeds demand, or to coordinate complex session migration procedures.

Sticky sessions create a scenario where a single server failure affects only the users routed to that server. That sounds like isolation, but it means failures are invisible until they happen. A server approaching memory limits continues receiving requests until it crashes, affecting only its sticky users. Without sticky sessions, a struggling server would naturally receive fewer requests as load redistributes.

Maintenance windows become surgical procedures. You must track which users are on which servers, drain connections carefully, and avoid disrupting active sessions. Rolling deployments require careful sequencing. With hundreds of servers and thousands of users, this coordination becomes error-prone.

Sticky sessions are reasonable for legacy applications where session state cannot be externalized without significant refactoring. They can reduce load on shared session storage for read-heavy workloads where local cache suffices. But even in these cases, treat sticky sessions as a temporary solution, not a permanent architecture.

External session storage works better. Redis provides sub-millisecond read latency for session data. Memcached works well for simple key-value session storage. Your load balancer routes requests freely across all backends, and any backend can serve any user because session state lives in shared storage.

# Problem: All of User A's requests go to Server 1
# If Server 1 fails, User A loses session

# Better: Store sessions in Redis
# All servers can serve User A
session_store: redis

Migrate sessions to Redis first. Once session state is external, disable sticky sessions and watch your load distribution improve automatically.

Not Monitoring Backend Load

A load balancer showing all backends equally utilized is not necessarily distributing load well. Simple metrics like connection count or request count hide the reality that different requests require different amounts of server resources. One backend might be running complex database queries while another handles lightweight static file requests.

Round-robin and least-connections algorithms work at the connection level, not the request level. A backend processing a long-polling request holds a connection for minutes while a backend serving fast static files cycles through many connections per second. Both show similar connection counts, but the long-polling server is much more loaded.

Monitor actual server resource utilization: CPU usage, memory pressure, disk I/O, and database query times. A backend at 80% CPU is more loaded than one at 20% even if both handle the same request rate. Track request latency per backend, not just overall latency. If one backend consistently shows higher latency, it is handling harder requests or is overloaded.

least_conn routes to the backend with fewest active connections, which is a better proxy for actual load on request-response workloads where connection time correlates with work done. weighted_round_robin accounts for different backend capacities by assigning weight to each backend. adaptive algorithms adjust based on real-time CPU or memory metrics, though these require more complex instrumentation.

Watch the gap between load balancer metrics and backend metrics. If backends are at high CPU while load balancer shows balanced traffic, your algorithm is misconfigured or your backends have different capacities. Auto-scaling should trigger on backend resource metrics, not load balancer metrics, to catch these hidden imbalances.

# Simple connection count is not enough
balance roundrobin  # Equal connections, not equal load

# Better: least_conn or weighted by actual load
balance least_conn

Set up dashboards that show backend CPU, memory, and latency alongside load balancer metrics. The moment backend utilization diverges while load balancer metrics look balanced, you have a problem to investigate.


Operational Practices

Test Failure Modes Regularly

In staging, terminate a backend and verify that the load balancer removes it, alerts fire, and users recover. Repeat the exercise after changes to health checks or routing rules.

Keep Load Balancers Focused

Use the load balancer to route traffic and enforce traffic policies. Keep compression, authentication, and business logic in application services unless the chosen proxy specifically needs to handle them.


Quick Recap Checklist

  • Load balancers distribute traffic across servers to improve availability and scalability.
  • L4 routes at the transport layer using IP and port for performance.
  • L7 routes at the application layer for content-aware decisions.
  • Health checks monitor backend availability and remove failed servers from rotation.
  • Sticky sessions preserve affinity but reduce routing flexibility.
  • TLS termination reduces backend cryptographic work and changes the trust boundary.
  • Software load balancers work well for many use cases; managed services reduce operational work.
  • Use redundant load balancers to avoid a single point of failure.

Interview Questions

1. What is the difference between Layer 4 and Layer 7 load balancing? When would you choose one over the other?

Layer 4 operates at the transport layer, making routing decisions based on IP address and port number without inspecting the actual request content. Layer 7 operates at the application layer, enabling content-based routing using HTTP headers, URLs, cookies, and request bodies.

Choose L4 when you need maximum throughput with minimal latency, are load balancing non-HTTP protocols like databases or SSH, or do not need application-layer inspection. Choose L7 when you need content-based routing, SSL termination, sticky sessions, or URL rewriting. L7 adds latency due to parsing but provides much greater routing intelligence.

2. What are sticky sessions and what problems do they create in distributed systems?

Sticky sessions route a particular user's requests to the same backend server using cookies, client IP hashing, or custom headers. The load balancer tracks which client maps to which server.

Problems they create: complicates horizontal scaling since you cannot freely redistribute load, makes maintenance windows difficult since taking down a server disconnects active users, creates session affinity that defeats load balancing benefits, and causes user-facing errors when a sticky server fails. Best practice is to store session state in a distributed cache like Redis instead of relying on sticky sessions.

3. How do health checks work in load balancers, and what are the different types?

Health checks are periodic tests the load balancer runs against backends to verify they can handle traffic. Types include:

  • TCP connect: Verifies the port is open and accepting connections
  • HTTP/HTTPS: Makes a full request and validates the response code and optionally the response body
  • Custom: Application-specific checks that verify actual service functionality

Health checks run at configurable intervals. A server failing too many checks gets marked unhealthy and removed from rotation. Lighter checks run more frequently; occasional deep validation catches real failures. Configure thresholds to avoid flapping — too aggressive removes healthy servers, too lenient means slow failure detection.

4. What is the difference between active-passive and active-active load balancer configurations?

In active-passive, one load balancer handles traffic while a standby monitors it through a mechanism such as VRRP. If the active fails, the standby takes over the virtual IP. This reserves standby capacity for failover rather than using both nodes to serve production traffic.

In active-active, multiple load balancers handle traffic simultaneously, which can use capacity more fully but requires coordination. Managed cloud services may provide multi-zone redundancy when configured for the provider’s supported availability model; verify the chosen service’s behavior and limits.

5. How does SSL termination work at the load balancer, and what are its trade-offs?

TLS termination lets the load balancer decrypt HTTPS traffic, inspect HTTP when operating at L7, and centralize certificate management. It shifts cryptographic work to the load balancer and creates a trust boundary.

If the backend leg uses plain HTTP, that traffic is unencrypted even inside a private network. Re-encrypt it when required by the network threat model or policy. TLS passthrough preserves encryption to the backend but prevents the load balancer from inspecting HTTP content.

6. What is Global Server Load Balancing (GSLB) and when do you need it?

GSLB commonly uses DNS policies to return an address for a healthy or preferred region. Unlike simple round-robin, a provider may consider health, configured weights, latency, or location. DNS answers can be cached by recursive resolvers until their TTL expires, so failover is not immediate for every client. GSLB is useful when an application spans regions and regional routing improves latency or availability.

GSLB health checks should cover regional failure modes, not just a single server. Probe from multiple vantage points, define what counts as an unhealthy region, and test how clients behave while cached DNS answers expire. Provider options include Amazon Route 53, Azure Traffic Manager, Google Cloud global load balancing, or dedicated GSLB products.

7. How do circuit breakers relate to load balancer health checks?

Load balancer health checks handle complete server failures — when a server stops responding entirely, health checks remove it from rotation. Circuit breakers work at a finer granularity, monitoring error rates and latency to detect degraded backends that are still responding but poorly.

When a circuit breaker opens, it immediately returns failures without forwarding requests to a struggling backend, giving that backend time to recover. Health checks continue monitoring an open circuit — once the backend recovers, the circuit breaker closes and allows traffic again. Service meshes like Istio implement circuit breaking at the sidecar proxy level with configurable outlier detection policies.

8. How would you implement a canary deployment using a load balancer?

A canary deployment routes a small percentage of production traffic to a new version while the majority stays on the current version. At the load balancer, you set backend weights — start with 5% to the new version, monitor error rates and latency for both groups, and gradually increase the weight as confidence builds.

For example, in HAProxy you would define both versions as backends and adjust weights: 95% to current, 5% to new. Incrementally shift to 90/10, 75/25, 50/50, and finally 0/100 when the canary is proven. If the canary shows elevated error rates or latency, immediately shift traffic back to the stable version. This requires proper observability — compare error rates and latency between groups, not just absolute numbers.

9. What is the difference between token bucket and sliding window rate limiting?

Token bucket allows bursts. Each client gets a bucket that fills with tokens at a steady rate — for example, 100 tokens per second. A request consumes one token. When the bucket is empty, requests are rejected or delayed. If traffic is below the rate limit, tokens accumulate, allowing brief bursts up to the bucket size.

A sliding-window counter tracks or estimates requests over a moving interval, smoothing fixed-window boundaries. Burst behavior depends on the implementation; a sliding-window log can enforce a stricter limit but uses more memory. Token bucket explicitly allows a configured burst, so choose based on your traffic and rejection policy.

10. What metrics should you monitor on a production load balancer?

Critical metrics include: request rate per backend (requests/second), response latency percentiles (p50, p95, p99) per backend, backend health status (healthy/unhealthy/draining), active connections per backend, backend error rate (5xx responses from backends), TLS handshake rate and latency if terminating SSL, and health check success/failure rate.

Also monitor load balancer CPU and memory utilization, connection exhaustion indicators (no new connections accepted), unusual traffic patterns (potential DDoS or abuse), and backend response time trends to catch degradation before it causes errors. Set up alerts on backend error rate increases, p99 latency exceeding SLO, and all backends becoming unhealthy simultaneously.

11. Describe how a load balancer performs health checking and what happens when a backend fails health checks.

Load balancers perform health checks at configurable intervals by sending probe requests to backend servers. Common types include TCP connect checks (verifies port is open), HTTP checks (makes request and validates response), and custom application-level checks.

When a backend fails consecutive health checks, the load balancer marks it unhealthy and removes it from the server pool. Existing connections may be terminated or allowed to drain depending on configuration. The load balancer stops routing new traffic to the failed backend. Once the backend recovers and passes enough consecutive health checks (typically 2-3 successes), it automatically rejoins the pool.

12. What is the difference between round-robin, least connections, and IP hash load balancing algorithms?

Round-robin assigns new connections or requests in rotation, depending on load-balancer mode. It is simple but does not account for server capacity or the amount of work in each request.

Least connections routes a new connection to the backend with the fewest active connections. This can help when connection duration tracks work, but long-lived or idle connections can make connection count a poor load signal.

IP hash computes a hash of the client IP to select a backend, which can provide affinity while the backend set remains stable. Clients behind NAT can share one address, and adding or removing backends may change mappings. It is not a substitute for shared session storage.

13. How does a blue-green deployment work with load balancers, and what are its advantages over rolling deployments?

Blue-green deployment maintains two identical environments. The current production environment (blue) handles live traffic while the new version (green) sits idle. Deployment involves shifting all traffic from blue to green at the load balancer level with a single routing rule change.

Blue-green makes rollback a routing change while the old environment remains healthy, and avoids a mixed-version rollout at the load-balancer layer. It requires capacity for both environments and compatible data changes. Use an expand-and-contract database migration so both versions can operate during the switch and rollback window.

14. What is anycast routing and how does it differ from GSLB-based load balancing?

Anycast advertises the same IP address from multiple locations. BGP selects a route according to network paths and policy; the selected site is often nearby but not necessarily geographically closest. This works at the network layer and is commonly used by CDNs.

GSLB can apply DNS policies using configured weights, health checks, and sometimes latency or location data. The location signal may come from a recursive resolver rather than the client, and DNS caching delays changes until TTLs expire. Unlike Anycast, GSLB can select among DNS records based on application health checks, subject to each provider’s probe and failover behavior.

15. What are the trade-offs of performing TLS termination at the load balancer versus SSL passthrough?

TLS termination at the load balancer: Reduces cryptographic work on application servers, centralizes certificate management, and allows HTTP inspection and URL rewriting when the load balancer operates at L7. It adds CPU load at the balancer and creates a trust boundary.

TLS passthrough: Keeps traffic encrypted to the backend, which handles decryption. The load balancer cannot inspect HTTP content or make URL-based routing decisions. If TLS terminates at the backend, certificate management remains distributed across those servers.

16. How does connection draining work and why is it important for maintenance windows?

Connection draining allows existing requests to complete on a backend server being taken offline while preventing new connections from being routed to it. The load balancer stops sending new traffic but allows in-flight requests to finish.

Without connection draining, abruptly removing a server can drop active connections. With draining configured, the balancer stops assigning eligible new traffic to the target while existing work has time to finish. Set the drain period from request-duration data and the product’s behavior; clients may still need retry logic if the deadline expires.

17. What is the role of load balancers in DDoS protection and rate limiting?

An L7 load balancer or reverse proxy can apply request limits before traffic reaches an application. Token bucket allows a configured burst; sliding-window behavior depends on the implementation. Rate-limit keys can include a source IP, API key, or JWT claim. Treat User-Agent as a weak signal because clients can change it.

A load balancer can spread traffic only within its own capacity and configured limits; it does not by itself stop a volumetric attack. Use provider or upstream DDoS protection for network-layer floods, and a WAF or application policy for request-rate abuse. Reputation scoring and rate-based rules depend on the provider and product.

18. Explain the relationship between ingress controllers and service mesh in Kubernetes load balancing.

Ingress or Gateway implementations handle north-south traffic into a cluster, commonly routing HTTP(S) by host and path. Community ingress-nginx retired on March 24, 2026 and no longer receives fixes; teams using it should plan migration. See the Kubernetes retirement notice.

Service meshes can handle east-west traffic through sidecar or other proxy architectures. Their supported protocols and features vary; policies may include retries, traffic shaping, and outlier detection, which add operational overhead.

Use both only when the deployment needs both roles. A Kubernetes Service may be enough for simpler cluster-internal traffic.

19. How would you design a highly available load balancer setup and what are the key considerations?

Design for redundancy from the start with at least two load balancers in active-passive or active-active configuration. Active-passive uses VRRP/keepalived where the standby monitors via heartbeat and takes over the virtual IP when active fails. Active-active distributes load across multiple LBs for better utilization.

Key considerations: Deploy load balancers in different availability zones, use managed cloud load balancers with built-in redundancy, implement connection draining for planned maintenance, configure proper health checks with appropriate thresholds to avoid flapping, and monitor load balancer health metrics alongside backend metrics.

20. What factors would influence your choice between using a managed cloud load balancer versus deploying software load balancers on VMs?

Choose managed cloud load balancer when: You want automatic scaling and high availability built-in, you prefer reducing operational overhead, you need tight integration with other cloud services (auto-scaling groups, security groups, WAF), or you are building cloud-native applications.

Choose software load balancers on VMs when: You need specific configuration options not available in the managed offering, cost is a major factor (managed LBs have per-hour pricing), you need consistent load balancing across multi-cloud or hybrid environments, or you want more control over routing logic and integration with external tools.


Further Reading


Conclusion

Copy/Paste Checklist

# Check HAProxy backend status (via socket)
echo "show stat" | socat stdio /var/run/haproxy.sock

# Query the configured NGINX stub_status location (example path)
curl http://localhost:8080/nginx_status

# Test health check endpoint
curl -I http://backend1:8080/health

# Check active connections
ss -s

# View HAProxy metrics
echo "show info" | socat stdio /var/run/haproxy.sock

# Test backend directly (bypass load balancer)
curl -H "Host: example.com" http://backend1:8080/

Category

Related Posts

CDN Deep Dive: Content Delivery Networks Explained

A comprehensive guide to CDNs — how they work, PoP architecture, anycast routing, cache invalidation strategies, SSL/TLS termination, and real-world performance trade-offs.

#system-design #networking #cdn

Forward and Reverse Proxies: Routing, Trust, and Use Cases

Learn how forward and reverse proxies handle HTTP traffic, CONNECT tunnels, TLS termination, caching, routing, trusted headers, and production failures.

#networking #proxies #http

Network Performance: Latency, Throughput, Jitter & Loss

Understand bandwidth, throughput, latency, RTT, jitter, and packet loss with practical measurements, tail percentiles, and production diagnostic guidance.

#networking #performance #latency