TCP Congestion Control: Flow, Loss, and Fairness
Learn how TCP congestion control limits traffic, responds to ACKs and loss, and affects throughput, latency, and fairness, with Linux inspection commands.
This guide explains how TCP balances a sender's congestion window against network feedback, and how that differs from receiver flow control. It compares loss-based control, CUBIC, BBR-style pacing, ECN, and the bandwidth-delay product, then shows which Linux socket and counter tools help investigate a slow transfer. Use the measurement and observability guidance to distinguish congestion from receiver limits, non-congestion loss, host pressure, or application delay before tuning production settings.
TCP Congestion Control: Flow, Loss, and Fairness
Introduction
TCP congestion control adjusts how much unacknowledged data a sender puts into the network, using signals such as loss and acknowledgements to respond to path conditions. It is separate from receiver flow control: congestion control protects the network path, while flow control prevents overwhelming the receiver.
This guide explains the congestion window, common signals and algorithms, and practical Linux inspection. It connects those mechanics to throughput and latency symptoms so teams can investigate before assuming every slowdown is a congestion problem.
Flow control and congestion control
TCP flow control uses the receiver window, usually written as rwnd. The receiver advertises how much additional data it can accept based on its receive buffer. If an application reads slowly, the advertised window may shrink or reach zero. This is a receiver-side limit.
Congestion control uses a sender-side congestion window, cwnd. It limits the amount of data the sender may have outstanding before it receives acknowledgments. The sender’s usable flight is constrained by both windows: in broad terms, it cannot have more outstanding data than the smaller of rwnd and cwnd permits. Socket memory, pacing, and other implementation limits can also affect what is actually sent.
For example, a server may have a 256 KiB receive window available while the path’s current congestion window is 64 KiB. The congestion window is then the tighter bound. If the client stops reading and advertises only 8 KiB, receiver flow control becomes the tighter bound instead. A transfer can alternate between these constraints.
The congestion window in practice
The sender does not normally receive a direct report of available path capacity. It infers a safe sending rate from delivery acknowledgments, loss, explicit congestion signals where enabled, and its own timing estimates. The congestion window is an amount of data, not a rate. The rate also depends on RTT: a rough upper estimate for a continuously backlogged flow is cwnd / RTT.
A related measure is the bandwidth-delay product (BDP), the amount of data that can be in flight to keep a path busy. For a path with 100 Mbps capacity and 80 ms RTT, the BDP is about 1 MB (100 megabits per second × 0.08 seconds, divided by 8 bits per byte). If the sender’s usable window is far below that amount, it may not fill the path even when neither endpoint is busy. This is one reason a high-latency connection needs a larger window to reach high throughput.
Slow start and congestion avoidance
In classic TCP descriptions, slow start begins with a relatively small congestion window and increases it quickly as ACKs arrive. The name is historical: growth is exponential by round trip while the window remains below the slow-start threshold (ssthresh). The sender is testing the path, not sending slowly in the everyday sense.
After the threshold, congestion avoidance grows the window more cautiously, classically by roughly one maximum-sized segment per RTT. This additive increase gives a flow room to use spare capacity without doubling its outstanding data every round trip. Real implementations may use different growth rules and pacing.
Loss detection and recovery
A gap in sequence numbers can cause the receiver to send duplicate ACKs for data it did receive after the gap. A sender that sees enough evidence of a missing segment may retransmit it before its retransmission timer expires. Selective acknowledgments (SACK) can describe which later blocks arrived, helping recovery retransmit only missing ranges.
A retransmission timeout (RTO) is a different signal: the sender did not receive an acknowledgment in time. A timeout can indicate severe delay, a lost ACK, or lost data. Classic loss-based control responds to timeout more sharply than to a loss inferred from duplicate ACKs, because the absence of feedback suggests that little data is getting through. The retransmission timer is based on measured RTT and variation, with safeguards against repeated rapid retransmissions.
On detecting congestion-related loss, a sender typically reduces its sending allowance, then grows it again as delivery continues. Classic TCP uses a multiplicative decrease; the details depend on the recovery algorithm and event. ACKs are evidence that data reached the receiver, but ACK arrival patterns can be distorted by delayed ACK behavior, ACK compression, or reordering. They are useful signals, not a perfect live map of the path.
graph TD
A[Application has data] --> B[Check receiver window and congestion window]
B --> C[Send permitted data]
C --> D[Network delivers segments]
D --> E[ACKs report received data]
E --> F[Update sending allowance]
F --> B
D --> G[Loss or congestion signal]
G --> H[Recover missing data and reduce sending]
H --> B
ACKs, loss, and congestion signals
In the classic model, packet loss is treated as evidence that a queue or link could not keep up. That inference works well across many wired paths, but it is not logically equivalent to congestion. Radio interference, a failing cable, overloaded receiver, faulty hardware, route changes, packet corruption, or policers can also drop packets. Reordering can look like loss temporarily until later packets arrive.
Explicit Congestion Notification (ECN) provides another option. When ECN is negotiated and supported along the path, a congested router can mark packets instead of dropping them; the receiver reports the mark and the sender reduces its rate. ECN does not remove the need to detect loss, and it only helps where endpoints and network equipment support and preserve the signal.
Different algorithms interpret delivery and congestion differently. Reno-style loss-based control and CUBIC use loss as an important congestion signal. CUBIC’s window growth is designed for high-speed and long-distance paths and is standardized in RFC 9438. BBR estimates bottleneck bandwidth and round-trip propagation time and uses pacing; it responds to delivery behavior differently from a loss-only model. These are examples, not a complete catalog. Defaults and available choices vary by operating system, kernel version, and administrator policy.
Production Failure Scenarios
| Scenario | What operators may see | Useful response |
|---|---|---|
| Receive application reads slowly | rwnd limits the flight size; the sender may show a persist timer or a small window |
Inspect receiver application queues and socket buffers; fix the slow consumer before changing congestion settings. |
| Loss on a long-distance path | Retransmissions rise and a classic loss-based flow’s throughput falls, even when nominal link capacity is high | Compare paths and endpoints, inspect retransmission causes, and involve the network provider if drops occur in transit. |
| Bursty wireless loss | TCP throughput fluctuates as losses trigger recovery; nearby Wi-Fi clients may fare worse than wired hosts | Compare wired and wireless paths, access point counters, and application timing; don’t assume every loss event came from a full router queue. |
| Bufferbloat under upload load | RTT and request latency rise while a bulk transfer runs; throughput may still look healthy | Compare idle and loaded RTT, then consider queue management or traffic shaping at the bottleneck. |
| Flow competes with different algorithms | One flow gets a different share or latency profile than another under the same bottleneck | Identify endpoint algorithms and queue policy; test a representative mix before changing defaults globally. |
| CPU or virtual NIC saturation | Retransmits, drops, or throughput plateaus coincide with host CPU pressure or interface counters | Inspect host and NIC counters, CPU scheduling, and virtualization limits along with path measurements. |
A common production pattern is a service that completes small requests normally but takes much longer to stream a large response to one region. RTT and loss may both be higher on that path, and the flow may need several RTTs to build a useful window again after loss. Compare per-region transfer timings and TCP retransmission counters before raising the application’s timeout. A longer timeout can hide a transport problem while increasing the number of occupied workers.
When to Use
Use congestion-control concepts when a TCP transfer fails to approach the expected path rate, latency grows during bulk traffic, retransmissions increase, or different flows receive uneven service. They help separate a receiver that is not reading from a sender that is limiting its flight size, and both from a path that is dropping or delaying packets.
They are also useful when tuning data transfer services, replication, backups, and long-lived HTTP connections. Start by observing the existing workload and path. A lab iperf3 result can help compare controlled endpoints, but it does not reproduce every production route, competing flow, or application access pattern.
When NOT to Use
Do not infer the cause of a slow request from one cwnd snapshot, a single ping, or a retransmission count without a time window and context. These measurements describe parts of a connection or path; they do not show whether time was spent in application code, storage, TLS setup, or a downstream dependency.
Do not change a host’s congestion-control algorithm globally just because one test improves. Check how that change affects mixed traffic, fairness, tail latency, CPU use, and other tenants. Avoid saturating a shared production link with throughput tests during an incident. Use a bounded test between authorized endpoints and label its traffic.
Trade-Off Table
| Approach | Strength | Cost or limitation |
|---|---|---|
| Classic loss-based control | Widely understood response to loss; interoperates across many network paths | Non-congestion loss can reduce sending, and queues can build before drops appear. |
| CUBIC | Window growth suited to high bandwidth-delay paths; standardized in RFC 9438 | Loss remains a central signal; a loss event can reduce throughput even when the drop has another cause. |
| BBR-style model-based control | Uses delivery-rate and RTT estimates to pace traffic; can work well on some high-BDP paths | Behavior with competing flows and bottleneck queue policies depends on algorithm version and deployment; measure the actual mix. |
| ECN-capable path | Can signal congestion without dropping marked packets | Requires negotiation and support along the path; marks may be disabled or erased. |
| Larger socket and receive windows | Can allow a flow to use a high-BDP path | Does not create capacity; memory use and queueing can rise, and auto-tuning behavior varies. |
No algorithm guarantees equal sharing for every combination of RTT, loss pattern, pacing, and queue management. Fairness is an observed property of competing flows at a bottleneck, not a label that can be assumed from an algorithm name. Measure representative combinations before deploying a change broadly.
Practical Linux inspection and measurement
Start with the active algorithm configured on Linux:
sysctl net.ipv4.tcp_congestion_control
sysctl net.ipv4.tcp_available_congestion_control
These values describe the host’s current setting and available choices. They do not prove which algorithm a particular connection used throughout its life, and distributions or containers may expose different settings and privileges.
Inspect TCP sockets and their reported state with ss:
ss -tin
ss -tin dst 203.0.113.20
The extended output may include RTT estimates, congestion window (cwnd), retransmission information, and the congestion-control name. Output varies by kernel and socket state. Take repeated snapshots while a representative transfer is active; a one-time sample can miss a short recovery or rate change.
The nstat utility can show host TCP counters. Counter names and availability vary, so inspect the list first and take before/after samples around a controlled interval:
nstat -az
sleep 30
nstat -az
For a quick filtered view on systems with the standard counter names:
nstat -az | grep -E 'TcpRetransSegs|TcpExtTCPTimeouts|TcpExtTCPLossProbes'
These are host-wide cumulative counters, not per-connection attribution. Compare deltas rather than totals, and account for unrelated traffic. A packet capture can help explain retransmission and reordering patterns, but capture filters, offload behavior, and encrypted payloads limit what it reveals:
sudo tcpdump -i eth0 -nn -s 128 -w tcp-sample.pcap 'host 203.0.113.20 and tcp'
Use an authorized destination and stop the capture after a short sample. Packet headers and timing can still contain sensitive operational data. For controlled throughput testing, run iperf3 only between endpoints you own or have permission to use. Capture the test duration, direction, parallel stream count, RTT, host load, and competing traffic so results can be repeated meaningfully.
For a rough BDP estimate, multiply measured path bandwidth in bits per second by RTT in seconds, then divide by eight to get bytes. A window below the BDP can limit a single flow, but increasing buffers without checking queueing may raise latency and memory use. Inspect receiver-window behavior and application read rates alongside cwnd.
This connects with the site’s guide to network performance measurements, which covers throughput, RTT, loss, and loaded-latency comparisons.
Observability Checklist
- Record the endpoint pair, direction, duration, RTT, test method, and whether traffic was competing with other flows.
- Track transfer goodput and application completion time; they answer different questions.
- Sample
cwnd, RTT, retransmits, and receiver-window behavior during the transfer when tooling permits. - Use counter deltas over a defined interval; label host-wide counters as host-wide.
- Compare idle and loaded RTT to identify queueing delay.
- Correlate TCP measurements with CPU, NIC drops, application read/write rates, and storage or dependency latency.
- Segment results by region, access network, protocol, and workload size where privacy and sample volume allow.
- Include the active algorithm, kernel version, and relevant queue policy in reproducible performance reports.
- Keep tests bounded and avoid comparing results from different routes or load conditions as if they were equivalent.
Security and Compliance Notes
Only capture traffic or run load tests on systems and paths you are authorized to inspect. High-rate TCP tests can consume shared capacity and resemble abusive traffic. Restrict access to packet captures and remove them when their diagnostic purpose ends; headers can expose internal addresses, connection timing, and service usage even when application payloads are encrypted.
Follow organizational retention and privacy rules for network telemetry. Avoid collecting payloads when timing and TCP headers are sufficient. Keep diagnostic files and command output free of credentials, and avoid publishing internal addresses, customer identifiers, or topology details in incident reports.
Common Pitfalls / Anti-Patterns
- Treating
rwndandcwndas interchangeable. The first reflects receiver capacity; the second reflects the sender’s congestion limit. - Treating packet loss as proof that a router queue filled. Wireless errors, host drops, and reordering can also appear as loss.
- Assuming ACK arrival is a precise measure of current bandwidth. ACKs can be delayed or compressed.
- Calling
cwnda sending rate. It is an outstanding-data limit; RTT and pacing also matter. - Concluding that the configured host default describes every connection and every peer.
- Comparing CUBIC and BBR based on different routes, RTTs, test lengths, or competing workloads.
- Increasing socket buffers to fix low throughput without checking BDP, receiver behavior, and queueing latency.
- Interpreting cumulative host counters as proof about one process or connection.
- Running unconstrained bandwidth tests on the same shared path customers rely on.
- Assuming fairness from an algorithm’s reputation rather than measuring representative competing flows.
Quick Recap Checklist
-
rwndprotects the receiver;cwndlimits the sender based on network feedback. - The smaller active window constrains outstanding data, subject to other stack limits.
- Slow start increases the sending allowance quickly; congestion avoidance grows it more cautiously.
- ACKs report delivery; duplicate ACKs, timeouts, SACK, and ECN contribute different evidence.
- Loss can signal congestion, but it can also come from a link, host, or reordering event.
- BDP relates bandwidth and RTT to the in-flight data needed to fill a path.
- Algorithm behavior and defaults differ across stacks; observe the actual connection and workload.
Interview Questions
Flow control uses the receiver-advertised window, rwnd, to prevent the sender from overrunning the receiver's available buffer. Congestion control uses the sender's congestion window, cwnd, to limit outstanding data based on network feedback. The usable flight is generally constrained by the smaller window.
Data remains in flight for longer before ACKs return. The bandwidth-delay product estimates how much data must be in flight to keep the path busy. If the sender's effective window is much smaller than that amount, it cannot fill the path even when the link has spare capacity.
Many loss-based algorithms interpret missing data as a sign that the network is overloaded and reduce their sending allowance. Wireless corruption, faulty links, host or NIC drops, and packet reordering can also cause loss-like evidence. The sender may reduce its rate before the operator identifies the non-congestion cause.
The sender waited for an acknowledgment longer than its current retransmission timer allowed. Data or its ACK may have been lost, or feedback may have been delayed. A timeout is not a diagnosis of the exact cause; correlate it with RTT, retransmits, receiver behavior, host drops, and path measurements.
Reproduce the transfer between representative, authorized endpoints and record direction, duration, RTT, goodput, and host load. Inspect active socket details with ss -tin, compare host counter deltas with nstat, and check receiver-window limits, retransmissions, NIC drops, and application read rates. Repeat under controlled load before changing an algorithm or buffer setting.
Expected answer points:
- Duplicate ACKs can indicate that later data arrived while a segment is missing, letting the sender retransmit before the timer expires.
- An RTO means the sender did not receive an ACK before its retransmission timer expired.
- A timeout is generally stronger evidence that feedback has stalled, so recovery is typically more conservative than for a fast-retransmit event.
Expected answer points:
- A capable router can mark an ECN-capable packet instead of dropping it when congestion builds.
- The receiver reports the mark so the sender can reduce its sending rate.
- Endpoints must negotiate and support ECN, and the path must preserve the markings and feedback.
Expected answer points:
- Run representative competing flows across the same controlled bottleneck, including relevant RTT and workload differences.
- Measure per-flow goodput, latency, loss, and queue behavior rather than relying on one isolated throughput result.
- Record algorithm versions and queue policy; the observed share depends on the deployment and traffic mix.
Expected answer points:
- A larger effective window can keep more data in flight and help fill a high-bandwidth, high-RTT path.
- Large queues can also hold traffic for longer under load, increasing RTT and delaying interactive requests.
- Measure throughput and loaded latency together, and check the actual bottleneck before raising buffer limits.
Further Reading
- RFC 5681: TCP Congestion Control describes classic slow start, congestion avoidance, fast retransmit, and fast recovery. It has been updated by RFC 9438.
- RFC 9438: CUBIC for Fast and Long-Distance Networks specifies the CUBIC congestion-control algorithm.
- RFC 3168: The Addition of Explicit Congestion Notification (ECN) to IP describes ECN support for IP and transport protocols.
- Linux
ss(8)manual documents socket inspection options. - Continue with the site’s network performance guide for practical RTT, throughput, loss, and latency measurements.
- Review TCP, IP, and UDP for how transport protocols fit into the network stack.
- See IP Routing and NAT for next-hop selection and packet paths across network boundaries.
- Read Network Latency, Timeouts, and Failure for deadlines, retries, and handling slow or failed requests.
Conclusion
TCP congestion control limits how much data a sender keeps in flight and adjusts that allowance as ACKs, loss, and other signals arrive. Receive-window flow control protects the destination application, while the congestion window responds to the path. BDP helps explain why a large window may be needed on a fast, high-latency route; loss, however, does not always mean congestion. On Linux, combine repeated ss snapshots, host counter deltas, endpoint metrics, and bounded path tests before tuning a production stack.
Category
Related Posts
Ethernet, ARP, and Neighbor Discovery Explained
Learn how Ethernet frames, switches, VLANs, ARP, and IPv6 Neighbor Discovery deliver packets on a local link, with Linux diagnostics and failure examples.
IP Routing and NAT: How Packets Cross Networks
Learn how routers choose paths with routing tables, longest-prefix match, and gateways, then see how NAT and port translation change packets at network edges.
Network Performance: Latency, Throughput, Jitter & Loss
Understand bandwidth, throughput, latency, RTT, jitter, and packet loss with practical measurements, tail percentiles, and production diagnostic guidance.