Latency
Latency is the time data takes to travel from one point on a network to another, in milliseconds; in practice it is quoted as round-trip time. It is delay, not capacity: distance and the number of round trips set it, so a nearby edge cuts it and extra bandwidth does not.
Also known as Network latency.
Full Explanation
Latency is how long data takes to get from one point on a network to another. Cloudflare's definition is “the time it takes for data to pass from one point on a network to another” (Cloudflare, what is latency). This is measured in milliseconds. Strictly, that is a one-way figure. But the number engineers quote is almost always the round trip. RTT “is equal to double the amount of latency, since data has to travel in both directions”. Cloudflare measures it that way “because every device has its own independent clock, so it’s hard to measure latency in just one direction” (Cloudflare, making home Internet faster). Latency is a measurement of time, not of volume. It is not bandwidth: that is the maximum data a link can carry at any given time. It is not throughput either: that is the amount that actually passes over a period. A wide link across an ocean has plenty of both, and is still slow per step. It is also not the whole of time to first byte. That metric adds the server’s own work to the trip. Two things set the floor. One is how far the signal must travel. The other is how many round trips a request needs before its first byte. Distance is physics. Round trips are engineering. That is the whole case for putting an edge server near the user rather than buying a fatter pipe.
How it works
Round-trip delay has four ingredients. They respond to completely different fixes.
- Propagation. For long-haul paths this dominates. Light travels through fibre at “roughly 203,000 kilometers per second, which is about two-thirds the speed of light through a vacuum” (TeleGeography). TeleGeography’s rule of thumb for a segment is that “1,000 kilometers equates to roughly 10 milliseconds” of round-trip delay. That is a rule of thumb for a link, not a guarantee for an Internet path. Cloudflare gives an illustration of the same effect. Requests from about 100 miles away reach a data centre in Columbus, Ohio “likely within 5-10 milliseconds”. Requests from Los Angeles, about 2,200 miles away, take “closer to 40-50 milliseconds” (Cloudflare).
- Transmission and per-hop processing. RFC 2681 puts the floor at propagation plus transmission. The minimum of the round-trip metric “provides an indication of the delay due only to propagation and transmission delay” (RFC 2681, section 1.1). On top of that, every hop costs work. Where packets cross between networks at an Internet exchange, routers “have to process and route the data packets, and at times routers may need to break them up into smaller packets, all of which adds a few milliseconds to RTT”.
- Queueing. Everything above that floor is load, not distance: “Values of this metric above the minimum provide an indication of the congestion present in the path” (RFC 2681, section 1.1). This is why the same path shows one number at 04:00 and another at peak.
- Round trips before the first byte. A cold request runs its setup in sequence. First comes a DNS lookup. Next is the TCP three-way handshake. RFC 9293 calls this “the procedure used to establish a connection” (RFC 9293, section 3.5). Then comes the TLS handshake. A full TLS 1.3 handshake costs one round trip before the client can send application data. RFC 8446 names that flow the “1-RTT handshake” (RFC 8446, section 2, section 2.3). Only then come the request and response themselves.
The compounding is what hurts. Each setup step waits a full round trip, so a cold request on a 100 ms path spends several hundred milliseconds before a single byte of content moves. Cloudflare states the effect plainly: an increase of a few milliseconds “is compounded by all the back-and-forth communication necessary for the client and server to establish a connection, the total size and load time of the page, and any problems with the network equipment the data passes through along the way”. Reduce the RTT, and every one of those steps shrinks at once. Remove a step, and you stop paying that RTT at all.
Note which quantity you are holding. RFC 2681 defines the round-trip metric like this: “This memo defines a metric for round-trip delay of packets across Internet paths” (RFC 2681, section 1). That is what ping, curl and CDN dashboards report. One-way delay is a separate metric. Halving an RTT does not reliably give it.
Why it matters for a CDN
A CDN is, at bottom, a latency product. It shortens the client-to-server leg and removes round trips. The evidence that this is the right lever is quantitative.
- Page loads are latency-bound, not bandwidth-bound. In the 2010 Google simulation Cloudflare cites, “above about 5 Mbps, the page doesn’t load much faster”. Later FCC measurement data moved that point of diminishing returns only “to about 20 Mbps”. Latency, by contrast, pays back linearly: “For every 20 milliseconds of reduced latency, the page load time improved by about 10%.” On Cloudflare’s own data, “For every 200 milliseconds of latency we can save, we cut the page load time by over 1 second”. That relationship holds both at 950 ms and at 50 ms (Cloudflare).
- Because dependencies serialise. A page is hundreds of objects discovered in stages. Cloudflare’s walk-through of one news home page counts 650 assets. The main HTML arrives more than a second in, before the browser even learns which scripts it needs. As Cloudflare puts it, “better latency makes every file load faster, which in turn unblocks other files faster, and so on”. And: “The protocols will use all the bandwidth available, but often complete a transfer before all the available bandwidth is consumed. It’s no wonder then that adding more bandwidth doesn’t speed up the page load, but better latency does.”
- A cache hit deletes the origin leg. On a hit, CloudFront “checks its cache for the requested object” and returns it from the edge. On a miss, it forwards the request to the origin. It then waits for the origin’s first byte before anything reaches the viewer (CloudFront, how it delivers content). Cache hit ratio is therefore a latency control, not only a cost control.
- High delay also caps a single transfer. “The larger the value of delay, the more difficult it is for transport-layer protocols to sustain high bandwidths” (RFC 2681, section 1.1). A distant path is slow in two ways at once. That is why latency and throughput have to be read together.
What CDNs do
Every large CDN shortens the same leg from user to edge. But they steer to a nearby point of presence in different ways. The difference shows up in how predictable the choice is.
- Cloudflare steers with anycast. One IP address is announced from many data centres. The selection “will typically be optimized to reduce latency by selecting the data center with the shortest distance from the requester” (Cloudflare on anycast). Note the hedge in the source. In a CDN context, anycast “typically routes incoming traffic to the nearest data center with the capacity to process the request efficiently”. So capacity and BGP policy can override raw proximity.
- Amazon CloudFront steers with DNS: “DNS routes the request to the CloudFront POP (edge location) that can best serve the request, typically the nearest CloudFront POP in terms of latency” (CloudFront). Again, typically, not always.
- Fastly offers both, and documents the trade-off. With a CNAME, “a Fastly authoritative name server” answers the query. It uses data from the Fastly insights program “to route the user to the POP that offers the most consistently high performance for that specific end user”. With an anycast A or AAAA record, required for apex domains, “the client will receive the anycast IP, which is advertised by multiple Fastly POPs”, and “Standard internet routing will take the user to a nearby POP”. Fastly notes this leaves it “less ability to direct traffic for best performance” (Fastly, routing traffic to Fastly).
- All of them cut round trips as well as distance. Persistent connections are the HTTP/1.1 default. RFC 9112 describes this as “allowing multiple requests and responses to be carried over a single connection” (RFC 9112, section 9.3). HTTP/2 “enables a more efficient use of network resources and a reduced latency by introducing field compression and allowing multiple concurrent exchanges on the same connection” (RFC 9113). HTTP/3 rides QUIC. QUIC “relies on a combined cryptographic and transport handshake to minimize connection establishment latency” (RFC 9000, section 7).
Watch out for
- Latency is a distribution, not a number. The minimum of a sample shows the path. Values above it show congestion (RFC 2681, section 1.1). RFC 2681 defines the Xth percentile of the round-trip sample as a statistic in its own right (section 4.1). So report the p95 and p99 beside the median. An average blends the congested samples into the uncongested ones and hides them. Variation itself is a separate metric. RFC 2681 warns that “erratic variation in delay makes it difficult (or impossible) to support many interactive real-time applications”. This is jitter, not latency.
- Idle latency is not the latency users get. Cloudflare distinguishes latency on an idle connection from “latency measured in working conditions when many connections share the network resources, which we call ‘working latency’ or ‘responsiveness’”. A download that overfills a buffer delays every other flow behind it. The bottleneck is usually the last mile. This is “either the wire that connects a home, or the modem or router in the home itself” (Cloudflare). A clean edge RTT measured from a data centre says little about a congested home link.
- A round trip is two paths, so RTT divided by two is not one-way delay. The forward and reverse routes may differ, “such that different sequences of routers are used for the forward and reverse paths. Therefore round-trip measurements actually measure the performance of two distinct paths together”, and even symmetric paths “may have radically different performance characteristics due to asymmetric queueing” (RFC 2681, section 1.1). If “performance of an application may depend mostly on the performance in one direction”, measure one-way delay instead.
- Vendor latency metrics have narrow definitions. CloudFront’s Origin latency is “the total time spent from when CloudFront receives a request to when it starts providing a response to the network (not the viewer), for requests that are served from the origin, not the CloudFront cache”. It is also known as first byte latency or time-to-first-byte. By that definition it excludes cache hits, and it is measured at the edge rather than at the viewer. It is one of the additional metrics that “must be turned on for each distribution separately” at extra cost (CloudFront metrics). Comparing it with a browser TTFB number is comparing two different clocks.
- Buying the last round trip has a security price. 0-RTT removes a round trip. But RFC 8446 is explicit that “the security properties for 0-RTT data are weaker than those for other kinds of TLS data”. The data is not forward secret, and “there are no guarantees of non-replay between connections” (RFC 8446, section 2.3). Use it for requests that are safe to replay.
- Bandwidth is not the fix, and neither is proximity alone. The two are separate axes. A nearby edge does nothing for a 4K video stream that needs capacity. A wider pipe does nothing for a page of hundreds of small, dependency-ordered objects.
Best practice
- Measure the split, not the total. Time DNS, TCP connect, TLS and first byte separately. The example below does this with curl. That split makes the dominant step visible before you change anything.
- Compare edge RTT against origin RTT. That gap is the CDN’s actual contribution. It is the honest number to use when evaluating providers or a new region.
- Remove round trips before shortening them. Reuse connections. Let one connection carry concurrent requests over HTTP/2 or HTTP/3 instead of opening one per asset. Resume TLS sessions. Keep the edge-to-origin connections warm so a cache miss does not pay a cold handshake.
- Raise cache hit ratio. A hit is the only change that removes the origin round trip entirely rather than making it shorter.
- Start the round trips earlier. Where a dependency is discovered late, warm the path in advance. Cloudflare notes that Early Hints inform “browsers of dependencies earlier, allowing them to pre-connect to servers or pre-fetch resources that don’t need to be strictly ordered” (Cloudflare). The practical form of that is a preconnect or prefetch hint on the resources the next step will need.
- Track percentiles and working latency, not the idle average. If the tail rises while the mean stays flat, suspect queueing and capacity, not the fibre.
- Do not quote a bare figure for a rule of thumb. The 10 ms per 1,000 km guide and any city-pair minimum are estimates for a segment. Real paths take longer. A measured number beats a computed one.
Examples
# Measure latency to a CDN edge
$ ping cdn.example.com
round-trip min/avg/max = 4.2/5.1/6.8 ms
# Measure full HTTP latency breakdown
$ curl -w '\nDNS: %{time_namelookup}s\nConnect: %{time_connect}s\nTLS: %{time_appconnect}s\nTTFB: %{time_starttransfer}s\nTotal: %{time_total}s\n' -o /dev/null -s https://cdn.example.com/
DNS: 0.012s
Connect: 0.018s
TLS: 0.032s
TTFB: 0.038s
Total: 0.045s
Frequently Asked Questions
Latency is the time data takes to travel from one point on a network to another, in milliseconds; in practice it is quoted as round-trip time. It is delay, not capacity: distance and the number of round trips set it, so a nearby edge cuts it and extra bandwidth does not.
# Measure latency to a CDN edge
$ ping cdn.example.com
round-trip min/avg/max = 4.2/5.1/6.8 ms
# Measure full HTTP latency breakdown
$ curl -w '\nDNS: %{time_namelookup}s\nConnect: %{time_connect}s\nTLS: %{time_appconnect}s\nTTFB: %{time_starttransfer}s\nTotal: %{time_total}s\n' -o /dev/null -s https://cdn.example.com/
DNS: 0.012s
Connect: 0.018s
TLS: 0.032s
TTFB: 0.038s
Total: 0.045s
Yes. Latency is also known as Network latency. Latency is the time data takes to travel from one point on a network to another, in milliseconds; in practice it is quoted as round-trip time. It is delay, not capacity: distance and the number of round trips set it, so a nearby edge cuts it and extra bandwidth does not.
Related CDN concepts include:
- Point of Presence (PoP) — A Point of Presence (PoP) is one location where a network keeps its own servers, …
- Anycast — Anycast announces one IP address from many locations at once; the routing system, usually BGP, …
- RTT (Round-Trip Time) (RTT) — RTT (round-trip time) is the delay from sending a packet until the response it triggers …
- TTFB (Time To First Byte) (TTFB) — TTFB is the time from the start of a request until the first byte of …