CWND (Congestion Window)
The congestion window (cwnd) is the TCP sender’s own limit on how much unacknowledged data it may have in flight. It is not the receiver’s advertised window — the smaller of the two governs — and because throughput is roughly cwnd ÷ RTT, it caps how fast one connection can go.
Also known as Congestion window, cwnd.
Full Explanation
The congestion window (CWND, usually written cwnd) is the sender-side state variable in TCP that limits how much unacknowledged data may be in flight at one time. It is the sender’s own running estimate of what the path will absorb, and it is not the receive window the far end advertises. RFC 5681 requires a sender to stay inside the minimum of cwnd and rwnd, so whichever of the two is smaller is the real ceiling. Nor is it a configured rate limit. It is a quantity of bytes. That quantity only becomes a rate when you divide by the time needed to acknowledge them. One flow’s throughput is approximately cwnd ÷ RTT (RFC 9438 section 3.3). Every connection starts with a small window and has to probe upward. That is why, for the small objects a CDN mostly serves, the window rather than the link usually decides how quickly the bytes arrive.
How it works
Cwnd is per-connection sender state, recomputed continuously from what the ACK stream reports. Almost everything you will see in a packet capture comes down to four questions. Where does the window start? How does it grow below the slow-start threshold? How does it grow above that threshold? How far does it drop when the path complains? The sender’s congestion-control algorithm chooses the answers.
- What it limits. RFC 5681 section 2 is blunt about the constraint: “a TCP MUST NOT send data with a sequence number higher than the sum of the highest acknowledged sequence number and the minimum of cwnd and rwnd.” Stacks disagree on the units. The same RFC notes that “some implementations maintain cwnd in units of bytes, while others in units of full-sized segments” (section 3.1). That is why Linux takes an initial window as a segment count, while RFC 9002 specifies QUIC’s window in bytes.
- Where it starts. The Standards Track initial window (IW) in RFC 5681 section 3.1 is only 2 to 4 segments, depending on the sender’s maximum segment size. The Experimental RFC 6928 section 2 raised the permitted upper bound to min(10*MSS, max(2*MSS, 14600)), about 14,600 bytes with a 1,460-byte MSS. The RFC is explicit that this is a permission, not a requirement: “This increase is optional: a TCP MAY start with an initial window that is smaller than 10 segments.” Linux takes the permission. include/net/tcp.h defines TCP_INIT_CWND 10, commented “as per rfc6928”.
- Slow start. While cwnd is below the slow-start threshold (ssthresh), “a TCP increments cwnd by at most SMSS bytes for each ACK received that cumulatively acknowledges new data” (RFC 5681 section 3.1). That is an upper bound of one doubling per round trip, not a guarantee of one. Receivers are told to acknowledge only “at least every second full-sized segment” (section 4.2). One ACK per two segments adds one segment for every two in flight. That grows the window by about half per round trip rather than doubling it.
- Congestion avoidance. Once cwnd passes ssthresh the growth becomes linear: “during congestion avoidance, cwnd is incremented by roughly 1 full-sized segment per round-trip time”. It MUST NOT be increased by more than SMSS bytes per RTT (RFC 5681 section 3.1).
- Backing off. Loss is the signal the standard algorithms react to. ECN “could also be used” instead (RFC 5681 section 3). On three duplicate ACKs, fast recovery sets ssthresh to no more than max(FlightSize / 2, 2*SMSS), equation (4) of section 3.1, applied by section 3.2. That roughly halves the usable window. A retransmission timeout is far harsher: cwnd “MUST be set to no more than the loss window, LW, which equals 1 full-sized segment (regardless of the value of IW)”, and slow start begins again.
- The algorithm decides the shape. CUBIC is the default in “the Linux, Windows, and Apple stacks” per RFC 9438, which obsoleted RFC 8312 in 2023. CUBIC “sets the multiplicative window decrease factor to 0.7, whereas Reno uses 0.5” (section 3.4), and states that factor as a SHOULD rather than an invariant (section 4.6), before growing the window back along a cubic curve. BBR does not use additive-increase/multiplicative-decrease at all. BBRv3 “uses recent measurements of a transport connection’s delivery rate, round-trip time, and packet loss rate to build an explicit model of the network path”. It computes a bandwidth-delay estimate of BBR.bw × BBR.min_rtt, and drives both its pacing rate and its in-flight cap from that model (draft-ietf-ccwg-bbr).
- The ceiling. Because throughput is about cwnd ÷ RTT, filling a path requires a window of roughly bandwidth × RTT, the bandwidth-delay product. RFC 7323 section 1.1 makes the mirror image concrete for the receive side. Without the Window Scale option, the largest usable window is 64 KiB. RFC 7323 states that “for LFN paths where the bandwidth * delay product exceeds 64 KiB, the receive window limits the maximum throughput of the TCP connection over the path”. The same arithmetic bounds cwnd.
- It is not permanent. A window earned once is not kept for ever. If a sender “has not sent data in an interval exceeding the retransmission timeout”, RFC 5681 says it SHOULD set cwnd to no more than the restart window: RW = min(IW, cwnd) (section 4.1).
- QUIC does the same job with different bookkeeping. QUIC’s controller “is per path”, keeps “the controller’s congestion window in bytes”, and an endpoint “MUST NOT send a packet if it would cause bytes_in_flight … to be larger than the congestion window” (RFC 9002 section 7). That is one window for the connection, shared by every stream on it. Endpoints “SHOULD use an initial congestion window of ten times the maximum datagram size … while limiting the window to the larger of 14,720 bytes or twice the maximum datagram size”, and the recommended minimum window is 2 × max_datagram_size (section 7.2). Loss or a rising ECN-CE count halves it (section 7.3.2). Persistent congestion drops it to that minimum, “similar to a TCP sender’s response on an RTO” (section 7.6.2).
Why it matters for a CDN
A CDN’s core promise is short round trips from a nearby edge. That promise is worth what it is worth largely because of cwnd. The window grows in units of round trips, so shortening the round trip is the only way to make it grow faster in wall-clock time.
- Small objects live or die inside the initial window. With a 1,460-byte MSS, IW10 is 14,600 bytes. RFC 6928’s increase “applies to the initial window of the connection in the first round-trip time (RTT) of data transmission during or following the TCP three-way handshake” (RFC 6928 section 2). A 15 KB HTML document therefore does not fit. The tail waits for the first ACK, so delivery costs at least two round trips no matter how much bandwidth the path has. That arithmetic is a floor under TTFB and under any render-blocking first flight. Low latency to the edge is what shrinks it.
- Large objects are limited by how long the ramp takes. Reaching a given window costs a number of round trips that no amount of bandwidth shortens. Only the RTT decides what each one costs. Slow start needs at best ten doublings to take a window from 10 segments past 10,000. That is roughly a second at 100 ms RTT, and about 50 ms at 5 ms RTT. An edge close to the user therefore reaches steady-state throughput sooner even when the bottleneck bandwidth is identical.
- Connection reuse concentrates the window. HTTP/2 clients “SHOULD NOT open more than one HTTP/2 connection to a given host and port pair” (RFC 9113 section 9.1), where “HTTP/1.0 and HTTP/1.1 clients use multiple connections to a server to make concurrent requests”. RFC 9113 notes in the same breath that verbose, repetitive header fields cause “the initial TCP congestion window to quickly fill” (section 1). So one already-grown window carries every object from that origin, instead of several independent windows each starting from IW. That is better for the ramp, and it is why one loss event on that connection throttles every response at once.
- Saving the handshake does not enlarge the first flight. With TCP Fast Open the server may send data during the handshake. But it “MUST follow [RFC5681] (based on [RFC3390]) to set the initial congestion window” (RFC 7413 section 4.2.2) and can send “up to an initial window of data” (section 7.2). TCP Fast Open and QUIC 0-RTT buy a round trip. Cwnd still decides how many bytes ride in it.
What CDNs do
Cwnd lives in the edge’s transport stack: the kernel’s TCP implementation, or a userspace QUIC library. Cwnd does not live in the caching layer. So what differs between vendors is which algorithm the sender runs and which knobs the operator has turned. Behaviour that current primary sources establish:
- Linux, which most edges run on. The initial window is fixed at ten segments in the kernel header (TCP_INIT_CWND 10), and CUBIC is the default algorithm (RFC 9438). Both are overridable per destination. ip route … initcwnd N sets “the initial congestion window size for connections to this destination”, where “actual window size is this value multiplied by the MSS”. congctl selects the algorithm per route (ip-route(8)).
- Cloudflare describes BBR as “a popular algorithm … which we have been using for much of our traffic”. As of its September 2025 post, it is running network-aware congestion-control experiments “on all of our free tier QUIC traffic”, reporting results averaging 10% faster than the prior baseline. It states that expansion to TCP traffic and to all customers is planned “over 2026 and beyond”. So the newer tuning is not yet universal across plans or protocols (Cloudflare blog).
- quiche, Cloudflare’s own QUIC library, makes the choice an explicit configuration knob: its CongestionControlAlgorithm offers Reno, CUBIC (documented as the default) and a BBRv2 implementation (quiche API documentation). The window’s behaviour is a per-endpoint implementation decision, not a property of the protocol.
Watch out for
- A timeout is not a halving. Fast recovery roughly halves the usable window. A retransmission timeout collapses cwnd to the loss window of one full-sized segment and restarts slow start (RFC 5681 section 3.1). QUIC’s equivalent is persistent congestion, which drops the window to the recommended minimum of 2 × max_datagram_size (RFC 9002 section 7.6.2). That is deliberately two packets rather than TCP’s one.
- “Cut in half” is algorithm-specific. Reno uses 0.5, CUBIC 0.7, and RFC 9438 phrases even that as a SHOULD (section 4.6). A model-based sender such as BBR may not reduce its in-flight cap on an isolated loss at all. Do not read a fixed factor out of a trace without knowing which sender produced it.
- A big cwnd is useless against a small receive window. The minimum of cwnd and rwnd governs (RFC 5681 section 2). Without the Window Scale option, rwnd cannot exceed 64 KiB. That alone caps throughput on long, fat paths (RFC 7323 section 1.1).
- An idle connection is not a warm connection. A pooled connection that has sent nothing for longer than the RTO SHOULD have cwnd reduced to min(IW, cwnd) before it sends again (RFC 5681 section 4.1). So keep-alive preserves the handshake and the TLS session, not necessarily the window.
- Raising the initial window is not free. RFC 6928 itself reports that “at initial windows larger than 10, the results are mixed” and that “this document does not include any supporting evidence for values of IW larger than 10” (section 1). A window that overshoots a slow access link simply converts into drops and retransmits.
- One QUIC or HTTP/2 connection means one window for every stream. QUIC removes head-of-line blocking at the stream layer. It does not give each stream its own congestion window. Since bytes_in_flight is accounted against a single per-path window (RFC 9002 section 7), a loss that reduces it slows every concurrent response on that connection.
- Units differ between stacks. Some implementations hold cwnd in bytes, others in full-sized segments (RFC 5681 section 3.1). Linux’s per-route initcwnd is a segment count multiplied by the MSS (ip-route(8)). Compare a captured cwnd against the right unit before concluding anything from it.
Best practice
- Serve from the closest edge. The window needs the same number of round trips wherever it runs. So cutting RTT is the only lever that makes it grow faster in real time.
- Aim to fit the critical first response inside one initial window, about 14,600 bytes at a 1,460-byte MSS. Do not rely on compression alone to get under it.
- Reuse connections and keep them working. Pool them for the handshake saving. But expect the window to be handed back after an idle gap longer than the RTO.
- Measure before tuning. Read cwnd and ssthresh per socket with ss -ti during a real transfer. Check whether the window is approaching bandwidth × RTT. If it is not, the limit is the receive window, the application, or a lossy hop. Enlarging cwnd will not help in that case.
- Change initcwnd only per destination and only with evidence. Keep pacing enabled: senders “SHOULD limit bursts to the initial congestion window”, and must either pace or limit bursts (RFC 9002 section 7.7).
- Choose the algorithm for the paths you actually serve. CUBIC is the default nearly everywhere. A model-based sender such as BBR is worth testing on paths with random loss or deep buffers. There, a loss-based window gives up throughput it did not need to.
Examples
# Check initial CWND on Linux
ip route show | grep initcwnd
# default via 10.0.0.1 dev eth0 initcwnd 10
# Set initial CWND to 15 segments
sudo ip route change default via 10.0.0.1 dev eth0 initcwnd 15
# Monitor CWND for active connections
ss -ti | grep -A5 'ESTAB'
# cwnd:10 ssthresh:65535 rtt:5.2/0.3
# Watch CWND growth in real time
ss -tin dst :443 | grep cwnd
# Repeat to see CWND increase during transfers
# Calculate minimum time to deliver N bytes
# IW=10 segments, segment=1460 bytes, RTT=50ms
# RTT 0: send 14,600 bytes (10 segments)
# RTT 1: send 29,200 bytes (20 segments)
# RTT 2: send 58,400 bytes (40 segments)
# Total after 3 RTTs (150ms): ~102 KB
# tcpdump to observe slow start
sudo tcpdump -i eth0 -n port 443 | \
awk '/length [0-9]/{print $1, $NF}'
# Watch packet sizes increase over time
Frequently Asked Questions
The congestion window (cwnd) is the TCP sender’s own limit on how much unacknowledged data it may have in flight. It is not the receiver’s advertised window — the smaller of the two governs — and because throughput is roughly cwnd ÷ RTT, it caps how fast one connection can go.
# Check initial CWND on Linux
ip route show | grep initcwnd
# default via 10.0.0.1 dev eth0 initcwnd 10
# Set initial CWND to 15 segments
sudo ip route change default via 10.0.0.1 dev eth0 initcwnd 15
# Monitor CWND for active connections
ss -ti | grep -A5 'ESTAB'
# cwnd:10 ssthresh:65535 rtt:5.2/0.3
# Watch CWND growth in real time
ss -tin dst :443 | grep cwnd
# Repeat to see CWND increase during transfers
# Calculate minimum time to deliver N bytes
# IW=10 segments, segment=1460 bytes, RTT=50ms
# RTT 0: send 14,600 bytes (10 segments)
# RTT 1: send 29,200 bytes (20 segments)
# RTT 2: send 58,400 bytes (40 segments)
# Total after 3 RTTs (150ms): ~102 KB
# tcpdump to observe slow start
sudo tcpdump -i eth0 -n port 443 | \
awk '/length [0-9]/{print $1, $NF}'
# Watch packet sizes increase over time
Yes. CWND (Congestion Window) is also known as Congestion window, cwnd. The congestion window (cwnd) is the TCP sender’s own limit on how much unacknowledged data it may have in flight. It is not the receiver’s advertised window — the smaller of the two governs — and because throughput is roughly cwnd ÷ RTT, it caps how fast one connection can go.
Related CDN concepts include:
- Latency — Latency is the time data takes to travel from one point on a network to …
- TCP (TCP) — TCP is the connection-oriented transport specified in RFC 9293: it delivers application data as one …
- Throughput — Throughput is the rate at which data actually crosses a link in a measured interval …
- BBR (Bottleneck Bandwidth and RTT) (BBR) — A congestion control algorithm from Google that measures a path's bottleneck bandwidth and minimum round-trip …
- TCP Fast Open (TFO) — TCP Fast Open (RFC 7413) is an experimental TCP extension that carries application data in …