Keep-Alive

Protocol

Reusing one TCP connection for many HTTP request/response exchanges instead of opening a new connection per request. The default in HTTP/1.1, it saves the TCP and TLS handshakes on every request after the first. Not the TCP keep-alive probe, and not pipelining.

Also known as persistent connections, keep-alive connections, HTTP keep-alive.

16 min read Updated Aug 30, 2026

Full Explanation

HTTP keep-alive is also called a persistent connection. It reuses one TCP connection for many HTTP request/response exchanges. This avoids opening a fresh connection for each one. It is the default in HTTP/1.1: “HTTP/1.1 defaults to the use of ‘persistent connections’, allowing multiple requests and responses to be carried over a single connection” (RFC 9112, section 9.3). It is not the same as the TCP keep-alive probe. That is a transport-layer liveness check. It is a separate mechanism and must default to off (RFC 9293, section 3.8.4). It is also not pipelining. Pipelining means sending several requests without waiting for each response (RFC 9112, section 9.3.2). Here is the helicopter view. After the first request on a connection, every later request on it skips the TCP and TLS handshakes. So all that is left to decide is how long an idle connection is kept, how many requests it may carry, and who closes it. On HTTP/2 and HTTP/3 the question changes shape rather than disappearing. Those connections are persistent by specification (RFC 9113, section 9.1; RFC 9114, section 3.3) and carry concurrent streams. So keep-alive there is about connection lifetime, not per-request setup.

How it works

Setting up a connection is the expensive part: no HTTP data can move until it is established. On HTTPS the HTTP client also acts as the TLS client. Only “when the TLS handshake has finished, the client may then initiate the first HTTP request” (RFC 9112, section 9.7). A TCP connection starts with the three-way handshake, “the procedure used to establish a connection” (RFC 9293, section 3.5). An HTTPS connection then adds a TLS handshake with round trips of its own before any HTTP bytes flow. TLS 1.3 can shave one of those round trips. It added a 0-RTT mode, “saving a round trip at connection setup for some application data, at the cost of certain security properties” (RFC 8446, section 1.2). That saves a round trip, not a whole handshake. Repeating that setup for every small resource on a page is pure overhead. Keep-alive removes that overhead.

Persistence is not negotiated by a header on HTTP/1.1. It is inferred. A recipient decides whether a connection persists using two things: the protocol version, and the Connection header field of the most recently received message, checked in this order (RFC 9112, section 9.3). If the “close” connection option is present, the connection will not persist. Else if the received protocol is HTTP/1.1 or later, it persists after the current response. Else if the protocol is HTTP/1.0, the “keep-alive” option is present, the recipient is not a proxy or the message is a response, and the recipient wishes to honour the HTTP/1.0 mechanism, then it persists. Otherwise it closes. A client “MAY send additional requests on a persistent connection until it sends or receives a ‘close’ connection option or receives an HTTP/1.0 response without a ‘keep-alive’ connection option”. A client that does not support persistent connections MUST send “close” in every request. A server that does not support them MUST send it in every non-1xx response.

Reuse only works if every message on the connection has a length the recipient can determine. It must not have to wait for the socket to close. “In order to remain persistent, all messages on a connection need to have a self-defined message length”. Also, “a server MUST read the entire request message body or close the connection after sending its response; otherwise, the remaining data on a persistent connection would be misinterpreted as the next request” (RFC 9112, section 9.3). A client that intends to reuse the connection MUST likewise read the entire response body. A proxy server also MUST NOT maintain a persistent connection with an HTTP/1.0 client.

Tearing down is explicit. The “close” connection option “is defined as a signal that the sender will close this connection after completion of the response” (RFC 9112, section 9.6). A client that sends it MUST NOT send further requests on that connection. A server that sends or receives it MUST initiate closure and MUST NOT process any further requests on it. Servers typically close in stages. They half-close the write side first, because an immediate close risks a TCP reset erasing the client’s unread buffers, so that the last response is lost.

Idle timeouts are entirely a matter of implementation. Servers “will usually have some timeout value beyond which they will no longer maintain an inactive connection”. And “the use of persistent connections places no requirements on the length (or existence) of this timeout for either the client or the server” (RFC 9112, section 9.5). Real defaults therefore differ by orders of magnitude: nginx keeps a client connection open for 75 seconds by default (nginx keepalive_timeout). It keeps an idle upstream connection for 60 seconds (nginx upstream keepalive_timeout). Amazon CloudFront defaults to 5 seconds toward a custom origin (CloudFront timeout quotas). Idle time is not the only limit: nginx closes a connection after 1000 requests or one hour of use by default. Fastly’s backends expose max_use and max_lifetime for the same purpose. So a connection can end while it is still busy.

Either side may also close at any moment. This makes reuse racy, since “connections can be closed at any time, with or without intention”. Implementations “ought to anticipate the need to recover from asynchronous close events” (RFC 9112, section 9.3.1). A client may be sending a request on what it believes is an idle-but-open connection. That can happen at the instant the server decides to reclaim it (section 9.5). So a keep-alive client needs retry logic. It has to respect idempotency when it retries.

HTTP/2 makes persistence structural instead of optional. “HTTP/2 connections are persistent”. Clients “SHOULD NOT open more than one HTTP/2 connection to a given host and port pair”. A peer that closes one SHOULD first send a GOAWAY frame, so both ends can tell which frames were processed (RFC 9113, section 9.1). Each request/response exchange gets its own stream. So many exchanges interleave on that single connection (RFC 9113, section 2), and “fewer TCP connections can be used in comparison to HTTP/1.x” (section 1). Liveness has its own tool. A PING frame is “a mechanism for measuring a minimal round-trip time from the sender, as well as determining whether an idle connection is still functional” (RFC 9113, section 6.7). The HTTP/1.1 vocabulary is banned outright. An endpoint MUST NOT generate an HTTP/2 message containing the Connection, Proxy-Connection, Keep-Alive, Transfer-Encoding or Upgrade fields. Any message containing them MUST be treated as malformed (RFC 9113, section 8.2.2).

HTTP/3 keeps the same shape but moves the timer into the transport. “HTTP/3 connections are persistent across multiple requests” (RFC 9114, section 3.3). Their lifetime is set by QUIC rather than by an HTTP setting, since “each QUIC endpoint declares an idle timeout during the handshake”. An implementation has to open a new connection once the existing one has been idle for longer than that. It SHOULD do so as it approaches the limit (RFC 9114, section 5.1). The same section also sanctions what a CDN edge does with connections it has not been asked for yet: “a gateway MAY maintain connections in anticipation of need rather than incur the latency cost of connection establishment to servers”, while “servers SHOULD NOT actively keep connections open”. It is a MAY, not an obligation. Vendors differ on whether they take it up.

Why it matters for a CDN

A proxied request is two connections, not one: viewer to edge, and, on a miss, edge to the origin that holds the content. Cloudflare describes the same split: “there are often two established TCP connections: the first is between the requesting client to Cloudflare and the second is between Cloudflare and the origin server” (Cloudflare connection limits). Keep-alive pays off on both. But it matters most on the origin hop, because every cache miss and every revalidation has to traverse it, while cache hits never do. AWS states the benefit plainly: maintaining a persistent connection “saves the time that is required to re-establish the TCP connection and perform another TLS handshake for subsequent requests” (CloudFront origin settings). Because the edge fronts many viewers, one reused origin connection can serve a burst of misses that would otherwise each open a socket.

The second effect is on the origin’s own capacity. Fewer connections is not only faster, it is lighter: “origin web servers close TCP connections if too many are open”. HTTP keep-alive “helps avoid connection resets for requests proxied by Cloudflare” (Cloudflare connection limits). The RFC makes the same point from the other side. Multiple connections avoid head-of-line blocking, but “each connection consumes server resources”. A server “might reject traffic that it deems abusive or characteristic of a denial-of-service attack, such as an excessive number of open connections from a single client” (RFC 9112, section 9.4).

Keep-alive also constrains how a CDN may forward at all, because its signalling is hop-by-hop. Intermediaries MUST parse a received Connection header field before forwarding. They MUST remove every field it names, and then remove the Connection field itself or replace it with their own control options. Keep-Alive is separately listed among fields intermediaries SHOULD remove before forwarding (RFC 9110, section 7.6.1). A viewer’s framing preferences are therefore none of the origin’s business. The edge terminates the viewer’s connection and decides independently how it talks to the origin.

The failure mode shows up on the miss path. If the origin closes idle connections sooner than the edge expects to reuse them, reuse simply never happens. Every miss re-pays the handshakes. First-byte time (TTFB) rises, while latency on hits stays flat. That is what makes the symptom diagnostic. The CDN-side knob alone cannot fix it: “for the Keep-alive timeout value to have an effect, your origin must be configured to allow persistent connections” (CloudFront origin settings).

What CDNs do

  • Cloudflare “maintains keep-alive connections to improve performance and reduce cost of recurring TCP connects” and “reuses open TCP connections up to the Proxy Idle Timeout limit after the last HTTP request”. It asks that keep-alive be enabled on the origin (Cloudflare connection limits). The documented limits toward the viewer are 400 seconds for both HTTP/1.1 Connection Keep-Alive and HTTP/2 Connection Idle, listed as not configurable. Toward the origin, the Proxy Idle Timeout is 900 seconds, documented as fixed, with a 520 if a connection closed for idleness is reused (Cloudflare HTTP/2 to origin). Alongside that it runs actual TCP keep-alive probes to the origin. The first probe comes after about 30 seconds of inactivity, the second 15 seconds later, and the connection resets after two unanswered probes.
  • Cloudflare HTTP/2 to origin is a separate lever: “at Cloudflare, HTTP/2 connection to the origin is enabled by default”. Multiple HTTP/2 streams then share one long-lived TCP connection instead of opening a connection per request. The multiplexing default is per plan, not universal. Free, Pro and Business zones multiplex by default. On Enterprise, “multiplexing starts effectively disabled (1 stream)” and has to be configured per zone (Cloudflare HTTP/2 to origin). Reuse there is opportunistic, not pre-warmed. Asked whether it prewarms connections to origins, Cloudflare answers “no. Connections are created on demand and reused where possible. There is no persistent idle pool” (Cloudflare HTTP/2 to origin). Fastly, by contrast, documents backend properties in terms of “a single, pooled HTTP keepalive connection” (Fastly Backend API). So do not assume a warm pool exists on your CDN: check.
  • Amazon CloudFront exposes a keep-alive timeout per origin, for custom and VPC origins only. It is “how long (in seconds) CloudFront tries to maintain a connection to your custom origin after it gets the last packet of a response” (CloudFront origin settings). The default is 5 seconds. Going beyond the account quota requires a quota increase applied to each origin (CloudFront timeout quotas). AWS frames the payoff as a metric to watch: “increasing the keep-alive timeout helps improve the request-per-connection metric for distributions”.
  • Fastly exposes it on the Backend API as keepalive_time, “how long (in seconds) to keep a persistent connection to the backend between requests”. It notes that “by default, Fastly keeps connections open as long as it can”. The same backend object caps reuse by count and age with max_use, “maximum number of requests allowed over a single, pooled HTTP keepalive connection to this backend”, and max_lifetime. It separately configures the transport probes with tcp_keepalive_enable, tcp_keepalive_time (default 300 seconds), tcp_keepalive_interval (default 10 seconds) and tcp_keepalive_probes (default 3) (Fastly Backend API). The two families sitting side by side in one API is the clearest illustration that HTTP keep-alive and TCP keep-alive are different things.

Watch out for

  • Two unrelated mechanisms share the name. TCP keep-alive is optional. Its probes are sent only when nothing is outstanding. If implemented, it MUST be switchable per connection, MUST default to off, and MUST use an interval defaulting to no less than two hours (RFC 9293, section 3.8.4). TCP keep-alive answers whether the peer is still there. HTTP keep-alive answers whether another request may be sent on this connection. Vendors override those transport defaults aggressively: Cloudflare probes after ~30 seconds, Fastly’s default is 300. So when a doc says “keep-alive”, check which layer it means.
  • Hop-by-hop fields leaking through. Connection and Keep-Alive are not end-to-end. The obligation is asymmetric and worth getting right: removing the fields named by Connection, and Connection itself, is a MUST, while removing Keep-Alive is a SHOULD (RFC 9110, section 7.6.1). If an origin sees a viewer’s Connection header, an intermediary failed to strip it. The classic consequence is a hung connection, since an HTTP/1.0 proxy that does not understand Connection “will erroneously forward that header field to the next inbound server, which would result in a hung connection” (RFC 9112, appendix C.2.2).
  • On HTTP/2 those fields are fatal, not merely useless. Any HTTP/2 message carrying Connection, Proxy-Connection, Keep-Alive, Transfer-Encoding or Upgrade MUST be treated as malformed. An intermediary translating HTTP/1.x to HTTP/2 MUST remove them. Forget that, and requests fail rather than merely losing persistence (RFC 9113, section 8.2.2).
  • nginx changed its defaults, so version matters. Before 1.29.7, proxy_http_version defaulted to 1.0 and upstream keep-alive was off. So keep-alive to an upstream required a keepalive directive in the upstream block, plus proxy_http_version 1.1, and clearing the Connection header: the recipe in the example above. Since 1.29.7, version 1.1 is the default, and the upstream connection cache is active by default with a limit of 32 idle connections per worker process (nginx proxy_http_version, nginx keepalive). Do not assume either state. Check the build you actually run.
  • Timeout mismatch between edge and origin. If the origin closes idle connections faster than the edge keeps them, the pool never gets reused. Every miss re-pays setup. Cloudflare says as much for its 900-second idle timeout: “origins that close connections faster than 900 seconds may experience connection churn” (Cloudflare HTTP/2 to origin). A CloudFront keep-alive timeout is also inert unless the origin allows persistent connections (CloudFront origin settings).
  • Reuse is racy. A pooled connection can be closed by the peer at the exact moment you write a request onto it. From the server’s point of view it closed an idle connection, while from the client’s a request was in progress (RFC 9112, section 9.5). Clients that pool connections need retries. Only idempotent requests can be retried blindly.
  • Not every closure is an idle timeout. Request-count and age caps end connections too: nginx defaults to 1000 requests and one hour per keep-alive connection. Fastly has max_use and max_lifetime for the same purpose. So connections that keep dropping are not automatically a timeout problem.
  • Framing bugs silently kill persistence, or worse. A message without a self-defined length cannot be followed by another on the same connection. Unread request bodies mean “the remaining data on a persistent connection would be misinterpreted as the next request” (RFC 9112, section 9.3). That misinterpretation is the substrate of request smuggling. Request smuggling “exploits differences in protocol parsing among various recipients to hide additional requests… within an apparently harmless request” (RFC 9112, section 11.2). This is a reason a CDN must be strict about framing, not just fast.
  • Bigger pools are not better. HTTP no longer names a maximum connection count. But it “encourages clients to be conservative when opening multiple connections” (RFC 9112, section 9.4), and nginx warns that the pool size “should be set to a number small enough to let upstream servers process new incoming connections as well” (nginx keepalive).

Best practice

  • Make the origin’s idle timeout at least as long as the edge’s reuse window. Read both numbers rather than assuming. An nginx origin holds an idle client connection for 75 seconds by default, well short of the 900 seconds Cloudflare may keep its side open. This produces churn and can surface as a 520 when a closed connection is reused. In the other direction, CloudFront gives up after 5 seconds by default. So the origin setting is not the binding constraint there. The CDN setting is.
  • Enable persistent connections on the origin explicitly. Every CDN-side keep-alive setting is inert if the origin refuses to hold connections open.
  • On an nginx origin or reverse proxy, verify proxy_http_version and the upstream keep-alive pool for your version. Before 1.29.7, both need setting by hand, together with clearing the Connection header.
  • Never relay a viewer’s Connection or Keep-Alive header to the origin. Terminate the viewer connection, then choose the origin protocol yourself. On HTTP/2, relaying those fields makes the message malformed.
  • Do not send “Connection: keep-alive” on HTTP/1.1. It is already the default. An intermediary will strip it. The spec says the HTTP/1.0 mechanism “ought not be used by clients at all when a proxy is being used” (RFC 9112, appendix C.2.2).
  • Where you control the pool, size it to real concurrency rather than to peak request rate. One reused connection serves many sequential requests. An oversized idle pool burns sockets on both ends for nothing.
  • Prefer HTTP/2 to the origin where both ends support it, since multiplexing removes the per-request connection question altogether. But confirm the multiplexing default for your plan instead of assuming it is on.
  • Monitor requests per connection alongside TTFB on misses. A falling requests-per-connection ratio localises the problem to connection reuse. A TTFB graph alone does not do that.

Examples

# HTTP/1.1 keep-alive is default
GET /page.html HTTP/1.1
Host: example.com
Connection: keep-alive  # Optional, already default

# Nginx: configure keep-alive
keepalive_timeout 65;     # Client-side timeout
keepalive_requests 1000;  # Max requests per connection

# Nginx: keep-alive to upstream (critical for CDN-origin)
upstream origin {
    server 10.0.1.100:443;
    keepalive 64;  # Pool of 64 persistent connections
}
location / {
    proxy_http_version 1.1;
    proxy_set_header Connection "";
}

Frequently Asked Questions

Reusing one TCP connection for many HTTP request/response exchanges instead of opening a new connection per request. The default in HTTP/1.1, it saves the TCP and TLS handshakes on every request after the first. Not the TCP keep-alive probe, and not pipelining.

# HTTP/1.1 keep-alive is default
GET /page.html HTTP/1.1
Host: example.com
Connection: keep-alive  # Optional, already default

# Nginx: configure keep-alive
keepalive_timeout 65;     # Client-side timeout
keepalive_requests 1000;  # Max requests per connection

# Nginx: keep-alive to upstream (critical for CDN-origin)
upstream origin {
    server 10.0.1.100:443;
    keepalive 64;  # Pool of 64 persistent connections
}
location / {
    proxy_http_version 1.1;
    proxy_set_header Connection "";
}

Yes. Keep-Alive is also known as persistent connections, keep-alive connections, HTTP keep-alive. Reusing one TCP connection for many HTTP request/response exchanges instead of opening a new connection per request. The default in HTTP/1.1, it saves the TCP and TLS handshakes on every request after the first. Not the TCP keep-alive probe, and not pipelining.

Related CDN concepts include:

  • HTTP/1.1 — HTTP/1.1 is the text-based version of HTTP, first published in January 1997 and defined today …
  • Latency — Latency is the time data takes to travel from one point on a network to …
  • RTT (Round-Trip Time) (RTT) — RTT (round-trip time) is the delay from sending a packet until the response it triggers …
  • TCP (TCP) — TCP is the connection-oriented transport specified in RFC 9293: it delivers application data as one …
  • TTFB (Time To First Byte) (TTFB) — TTFB is the time from the start of a request until the first byte of …