TCP
TCP is the connection-oriented transport specified in RFC 9293: it delivers application data as one reliable, in-order byte stream, numbering every byte and retransmitting losses. It is not UDP, and not HTTP/3’s transport — that is QUIC. Every connection opens with a three-way handshake.
Also known as Transmission Control Protocol.
Full Explanation
TCP (Transmission Control Protocol) is the connection-oriented transport protocol. It gives an application one reliable, in-order stream of bytes across a packet-switched network. The current specification is RFC 9293 (STD 7). It obsoletes the original RFC 793. HTTP/1.1 and HTTP/2 both run over TCP. So TCP still carries most of the web.
TCP is not UDP. UDP is transaction oriented. RFC 768 says plainly that with UDP “delivery and duplicate protection are not guaranteed”. It also says that “applications requiring ordered reliable delivery of streams of data should use the Transmission Control Protocol” (RFC 768). TCP is also not the transport under HTTP/3. QUIC is the transport there instead. QUIC’s “packets are carried in UDP datagrams” (RFC 9000, section 1). TCP is not a security layer either. Encryption comes from TLS running on top of it.
The short version: TCP numbers every byte, acknowledges cumulatively, and retransmits what is lost. It holds the sender inside the receiver’s advertised window, and also inside its own estimate of what the path will carry. It opens every connection with a three-way handshake. That handshake is TCP’s defining cost: no application data moves until it finishes.
How it works
Four mechanisms and a teardown turn a lossy packet network into a byte stream. That stream looks local to the application.
- Establish the connection. “The ‘three-way handshake’ is the procedure used to establish a connection” (RFC 9293, section 3.5). The client sends SYN, the server answers SYN-ACK, and the client answers ACK. It is called three-way because “steps 2 and 3 can be combined in a single message” (section 3.4.1). The price is fixed: “Current TCP only permits data exchange after the three-way handshake (3WHS), which adds one RTT to network latency” (RFC 7413, section 1). That is one round trip of the path, before the first request byte. TCP Fast Open is the experimental exception. It carries data in the SYN itself.
- Number and acknowledge. “Since every octet is sequenced, each of them can be acknowledged” (section 3.4). The acknowledgment is cumulative. So an ACK of sequence number X reports “that all octets up to but not including X have been received” (section 3.4). What goes missing is sent again: “Because segments may be lost due to errors (checksum test failure) or network congestion, TCP uses retransmission to ensure delivery of every segment” (section 3.8).
- Flow control: the receiver’s window. Each segment carries a window. The window “indicates the range of sequence numbers the sender of the window (the data receiver) is currently prepared to accept” (section 3.8.6). So a fast sender cannot overrun a slow peer. This is a receiver-imposed limit. It is a separate mechanism from congestion control.
- Congestion control: the sender’s window. “A TCP endpoint MUST implement the basic congestion control algorithms slow start, congestion avoidance, and exponential backoff of RTO to avoid creating congestion collapse conditions” (section 3.8.2). The sender’s own limit is the congestion window. It deliberately starts small. RFC 5681 caps the initial window at two to four segments depending on segment size (RFC 5681, section 3.1). RFC 6928 raises that upper bound to “10 segments (maximum 14600 B)” (RFC 6928, section 1). RFC 6928 is an Experimental specification. A stack may choose to stay under this bound. Slow start then grows the window as acknowledgments return. CUBIC is today’s default. “CUBIC has been adopted as the default TCP congestion control algorithm in the Linux, Windows, and Apple stacks”. It is “to be regarded as the currently most widely deployed standard for TCP congestion control” (RFC 9438, section 1). BBR instead uses “recent measurements of a transport connection’s delivery rate, round-trip time, and packet loss rate to build an explicit model of the network path”. BBR is still an Experimental IETF draft. It has released implementations for Linux TCP and QUIC (draft-ietf-ccwg-bbr).
- Close it. Each side sends a FIN when it has no more data to send (section 3.6). Whichever side closes actively “MUST linger in the TIME-WAIT state for a time 2xMSL (Maximum Segment Lifetime)” (section 3.6.1). RFC 9293 does not treat MSL as a universal constant: “For this specification the MSL is taken to be 2 minutes. This is an engineering choice, and may be changed if experience indicates it is desirable to do so” (section 3.4.2). So read 2xMSL as this specification’s default rather than a fixed four minutes on every stack. Teardown, not just setup, has a cost.
Two framing properties matter here (section 2.2). The connection is full duplex: “Data flow is supported bidirectionally over TCP connections”. It is also point to point: “TCP supports unicast delivery of data”. The same section adds one more qualification. It matters most for CDN routing: “There are anycast applications that can successfully use TCP without modifications, though there is some risk of instability due to changes of lower-layer forwarding behavior.” Consider a re-route that moves a client to a different anycast node mid-connection. It delivers the client’s segments to a machine with no state for that connection. Port numbers identify which application service on a host a connection belongs to.
Why it matters for a CDN
For a CDN, TCP is the price of the first byte. That price is charged in round trips before any content moves.
- Setup dominates the first request. One round trip pays for the TCP handshake. Then the TLS handshake adds more, on top, at an HTTPS edge server. RFC 8470 describes early data as “avoiding the one or two round-trip delays needed for the TLS handshake” (RFC 8470, section 1). That is one round trip for TLS 1.3, and two for TLS 1.2. On a cold connection, TTFB is normally setup-bound rather than bandwidth-bound.
- Distance multiplies every round trip. The handshake, the TLS exchange, and each slow-start round all cost one round trip of the path. So a point of presence near the audience shrinks all of them at once. That is the mechanism behind most of a CDN’s latency win.
- Small objects finish before the connection is up to speed. Slow start exists because “beginning transmission into a network with unknown conditions requires TCP to slowly probe the network to determine the available capacity” (RFC 5681, section 3.1). A response that fits inside the first few windows completes while the sender is still ramping. So link capacity barely affects it.
- Fat request headers cost time on new connections. Repetitive, verbose HTTP fields cause “the initial TCP congestion window to quickly fill” (RFC 9113, section 1). This “can result in excessive latency when multiple requests are made on a new TCP connection” (RFC 9113, section 1).
What CDNs do
A CDN terminates the client’s TCP connection at the edge. It opens its own connection towards the origin. Each leg can then be managed separately.
- Reuse connections on both legs. The client leg uses Keep-alive. The origin leg uses a pool of warm edge-to-origin connections. Together, they mean the handshake is paid per session, not per request. HTTP/2 sharpens the same point: “fewer TCP connections can be used in comparison to HTTP/1.x. This means less competition with other flows and longer-lived connections, which in turn lead to better utilization of available network capacity” (RFC 9113, section 1).
- Cloudflare: HTTP/3 on the client hop. HTTP/3 with QUIC “is available to all plans (though it does require an SSL certificate at Cloudflare’s edge network)”. You switch it on in the dashboard under Speed → Settings → Protocol Optimization. The scope is stated outright: “This setting is for connection between the user and Cloudflare. HTTP/3 connection to the origin is not yet supported” (Cloudflare, HTTP/3 (with QUIC)).
- Akamai: HTTP/3 as an opt-in behavior. You “add the behavior and set the Enable slider to On” in a property. You can narrow it by hostname or by a percentage of clients. HTTP/3 “moves away from the traditional transmission control protocol (TCP) transport layer” and “uses the IETF QUIC protocol”. The certificate “needs to have transport layer security (TLS) 1.3 enabled in its deployment settings”. An Alt-Svc header advertising HTTP/3 is generated automatically, with a preset max-age of 93600 seconds. HTTP/3 does not replace HTTP/2. So the HTTP/2 behavior must stay in the property, to keep serving HTTP/2 clients (Akamai, HTTP/3).
- Keep the origin leg on TCP. Cloudflare documents that HTTP/3 to the origin is not yet supported. Akamai documents its HTTP/3 behavior only for “connections between requesting clients and the Akamai edge”. So origin-side HTTP/3 is simply not documented. Assume the edge-to-origin hop is TCP, normally TLS over TCP, unless your provider documents otherwise.
Watch out for
- Head-of-line blocking belongs to the transport, not to HTTP. Because “the parallel nature of HTTP/2’s multiplexing is not visible to TCP’s loss recovery mechanisms, a lost or reordered packet causes all active transactions to experience a stall regardless of whether that transaction was directly impacted by the lost packet” (RFC 9114, section 1.1). HTTP/2 concedes the point: “Note, however, that TCP head-of-line blocking is not addressed by this protocol” (RFC 9113, section 1). Consolidating many requests onto one TCP connection concentrates the exposure. Only QUIC’s per-stream reliability removes it.
- Every cold connection pays setup again. A first-time visitor, an idle connection that a client or a middlebox dropped, and a failover to a different edge all restart at SYN. Connection reuse hides that cost. It never removes it. TLS 1.3 early data shortens the restart. But it constrains what may ride in it. Clients “MAY send requests with safe HTTP methods … in early data when it is available and MUST NOT send unsafe methods (or methods whose safety is not known) in early data” (RFC 8470, section 2). That is because an attacker “might capture and replay the request(s) it contains” (section 1).
- Unicast means per-viewer state. TCP delivers to exactly one peer. So a live event is served as one connection per viewer. Fan-out comes from caching and request coalescing at the edge, never from the transport. So budget edge sockets and memory per concurrent viewer.
- Connection churn leaves residue. An actively closed connection lingers in TIME-WAIT for twice the maximum segment lifetime. So an edge or origin that cycles through short connections can run out of ephemeral ports and socket state. This happens long before it runs out of bandwidth. Long-lived pooled connections are the fix.
- The handshake is an attack surface. RFC 9293 lists SYN flooding among the attacks “focused on exhausting the resources of a TCP server” (RFC 9293, section 7). A CDN absorbs most of it in front of the origin. But the origin still needs its own connection limits.
Best practice
- Reuse connections on every leg: keep-alive to the client, a warm pool to the origin, and idle timeouts long enough that an ordinary page load never re-handshakes.
- Enable HTTP/3 for the client hop where your provider offers it. Keep HTTP/2 over TCP enabled behind it. HTTP/3 does not replace HTTP/2. Clients without QUIC, or on paths that block UDP, still arrive over TCP.
- Serve from a point of presence near the audience. Nothing else cuts handshake and slow-start time by as much.
- Measure setup separately from transfer, and subtract correctly. In curl, time_connect is “the time, in seconds, it took from the start until the TCP connect to the remote host (or proxy) was completed”. And time_appconnect runs “from the start until the SSL/SSH/etc connect/handshake to the remote host was completed” (curl manual, --write-out). Both are cumulative from the start of the request. So the TLS handshake on its own is time_appconnect minus time_connect.
- Judge delivery by TTFB and total transfer time rather than by link bandwidth. For small objects the handshakes and the slow-start ramp, not the wire, decide the duration.
- Leave congestion control to a current algorithm instead of hand-tuning windows: CUBIC by default, BBR where you have measured a gain.
Examples
Check the TCP setup with curl:
# Show TCP connect time separately
curl -w "TCP connect: %{time_connect}s\nTLS done: %{time_appconnect}s\nTotal: %{time_total}s\n" -o /dev/null -s https://cdn.example.com/style.css
# TCP connect: 0.012s
# TLS done: 0.035s
# Total: 0.042s
Nginx keeps a pool of connections to the origin:
upstream origin {
server origin.example.com:443;
keepalive 64; # pool of idle connections
}
server {
location / {
proxy_pass https://origin;
proxy_http_version 1.1;
proxy_set_header Connection "";
}
}
Frequently Asked Questions
TCP is the connection-oriented transport specified in RFC 9293: it delivers application data as one reliable, in-order byte stream, numbering every byte and retransmitting losses. It is not UDP, and not HTTP/3’s transport — that is QUIC. Every connection opens with a three-way handshake.
Check the TCP setup with curl:
# Show TCP connect time separately
curl -w "TCP connect: %{time_connect}s\nTLS done: %{time_appconnect}s\nTotal: %{time_total}s\n" -o /dev/null -s https://cdn.example.com/style.css
# TCP connect: 0.012s
# TLS done: 0.035s
# Total: 0.042s
Nginx keeps a pool of connections to the origin:
upstream origin {
server origin.example.com:443;
keepalive 64; # pool of idle connections
}
server {
location / {
proxy_pass https://origin;
proxy_http_version 1.1;
proxy_set_header Connection "";
}
}
Yes. TCP is also known as Transmission Control Protocol. TCP is the connection-oriented transport specified in RFC 9293: it delivers application data as one reliable, in-order byte stream, numbering every byte and retransmitting losses. It is not UDP, and not HTTP/3’s transport — that is QUIC. Every connection opens with a three-way handshake.
Related CDN concepts include:
- UDP (UDP) — User Datagram Protocol: a connectionless, best-effort transport that puts each message in one IP packet …