Request Coalescing
A cache behaviour that serves many concurrent requests for one cache key from a single upstream fetch: the first miss goes to the origin, the rest wait on a queue and are answered from that one response. It exists to stop a cache stampede. Also called request collapsing or collapsed forwarding.
Also known as request collapsing, collapsed forwarding.
Full Explanation
Request coalescing is a cache behaviour. On a miss, the cache sends one request upstream for a given cache key. It then answers every other request for that key from that same response while the fetch is in flight. Fastly calls it request collapsing and defines it as combining multiple requests for the same object into a single request to origin, and then potentially using the resulting response to satisfy all pending requests. The word potentially is load-bearing: the queue is only satisfied if the response turns out to be usable. Its purpose is to stop a cache stampede. Fastly again: it prevents the expiry of a very highly demanded object in the cache causing an immediate flood of requests to an origin server, which might otherwise overwhelm it or consume expensive resources. One fetch reaches the origin instead of thousands.
It is not a protocol feature. It is not a Cache-Control directive either. No client asks for it, and no origin switches it on with a header. It is an implementation choice inside a cache server. That is why every platform names it differently: request collapsing (Fastly, Cloudflare), collapsed forwarding (Squid), proxy_cache_lock (nginx), read-while-writer plus cache-read retry (Apache Traffic Server). Those names are close but not interchangeable: ATS documents that once its read-while-writer settings are enabled you have something that is very close, but not quite the same, to Squid’s Collapsed Forwarding. It is also not stale-while-revalidate. Coalescing makes the waiting clients wait for a fresh fetch. Stale-while-revalidate answers them immediately from the stale copy instead. The two compose, and in production they usually should. Finally, it is not global. The merge happens inside one cache node or location, not across a CDN.
How it works
Coalescing is one fetch plus a queue, scoped to a single cache key. nginx states the rule most plainly. With proxy_cache_lock on, only one request at a time will be allowed to populate a new cache element identified according to the proxy_cache_key directive by passing a request to a proxied server. The rest either wait for a response to appear in the cache or the cache lock for this element to be released. Fastly calls that queue a waiting list. Cloudflare calls the mechanism a cache lock.
- A request misses. The cache begins the cache fill: one fetch upstream.
- Further requests for the same key are held instead of forwarded. Cloudflare: only the first request is forwarded to the origin to fetch the asset.
- If the response is cacheable, it is stored. It is then used for the whole queue: the remaining requests wait for the first request to complete, after which the response is streamed to all waiting requests.
- If the response is not usable, the queue is dequeued instead of satisfied. That path is the dangerous one (see Watch out for).
- Requests that arrive after the object is stored are ordinary hits. Fastly: once that response completes, normal caching behavior takes over, and collapsing stops.
The window is narrow. It is defined by the fetch, not by the clock. Fastly: to collapse requests they must be concurrent with the request that initiated the fetch from origin. Squid is exact about what concurrent means: received after the first request headers were parsed and before the corresponding response headers were parsed. Two mechanisms widen that window past the response headers. Fastly’s streaming miss writes the partial response to cache as soon as the headers arrive. A later request then joins the in-progress stream if the object is still fresh. If the object has already gone stale, that request is a miss instead, and it starts a new fetch. ATS’s read-while-writer is the same idea: the ability to read a cached object while another connection is completing the write to cache for that same object. But it starts late on purpose. ATS does not begin allowing clients to read until after the complete HTTP response headers have been read and processed, because until then it cannot know the response will be cacheable.
Two preconditions therefore apply. First, the request must be able to produce a cache object at all. Fastly: PASS requests and many errors are uncacheable by default, which means that we will never be able to successfully collapse those requests. Second, the resulting object must be usable. That means cacheable with a positive remaining lifetime.
One distinction is easy to miss: a cold miss and a stale revalidation are not the same case. Squid collapses two kinds of requests: regular client requests received on one of the listening ports and internal “cache revalidation” requests which are triggered by those regular requests hitting a stale cached object. nginx’s lock, by its own wording, guards only the population of a new cache element. The expiry case is covered by separate directives instead: proxy_cache_use_stale updating and proxy_cache_background_update. On a hot object that expires every few seconds, that difference decides whether the herd is absorbed at all.
Why it matters for a CDN
The size of the herd is set by request rate multiplied by fetch duration, not by how many users exist. Fastly’s own arithmetic: for an object requested 50 times per second, with a 500ms fetch latency, the origin would be processing 25 concurrent requests for the same object before Fastly had the opportunity to store it in cache. That is one object, on one PoP, at a modest rate. Coalescing turns those 25 into one. That is why Fastly can say that for high traffic services, correct use of request collapsing will substantially reduce and smooth out traffic to origin servers. What it saves is origin concurrency, connections and egress, not fetch time.
The corollary is counter-intuitive. It is worth internalising before you tune anything. Fastly: somewhat counter-intuitively, faster origins and features that improve time to first byte will reduce how often we can collapse requests. The reverse also applies: you may also see a higher occurrence of request collapsing when an origin takes longer to respond. Coalescing is a shock absorber that engages exactly when the origin is struggling. A falling collapse rate is usually good news about the origin, not a broken feature.
The cost is paid by the clients in the queue. Varnish puts it bluntly: at high rates the queue of waiting requests can get huge. Two problems follow: a thundering herd on release, and nobody likes to wait. Coalescing converts an origin-capacity problem into a client tail-latency problem. It belongs next to a stale-serving strategy, not on its own.
Topology decides how many separate queues your origin faces. Each cache node collapses only its own traffic. Concentrating misses is what makes coalescing effective across a global network. Fastly notes that disabling clustering has the potential to significantly increase traffic to origin not just due to a poorer cache hit ratio, but also due to the inability to collapse concurrent requests that originate on different delivery servers. Fastly adds that with shielding, requests from POPs across the network are focused into a single POP, allowing it to perform request collapsing before forwarding a single request to origin. Cloudflare’s Tiered Cache does the same by hierarchy: if the upper-tier does not have the content, only the upper-tier can ask the origin for content.
What CDNs do
- nginx: nginx ships with proxy_cache_lock off. Two timers govern the queue. proxy_cache_lock_timeout (default 5s) is one: after it expires, the request will be passed to the proxied server, however, the response will not be cached. proxy_cache_lock_age (default 5s) is the other: if the request populating the element has not completed for the specified time, one more request may be passed to the proxied server. That is a may. It is the knob most people forget exists.
- Varnish is automatic and needs no configuration: when several clients are requesting the same page Varnish will send one request to the backend and place the others on hold while fetching one copy from the backend. In some products this is called request coalescing and Varnish does this automatically. Its stated rule for an expired object is that if the grace period has run out and there is an ongoing backend request, then the request will wait until the backend request finishes. Within grace, the stale object is served instead of waiting.
- Fastly: on by default, with a limit the summary usually drops: by default, cache misses will qualify for request collapsing in both VCL and Compute services, when using the readthrough or simple cache interfaces. In contrast, the core cache interface supports request collapsing but only when explicitly configured within a cache transaction. In VCL services it can be switched off per request by setting req.hash_ignore_busy to true in vcl_recv. Doing so also makes stale objects unusable. Fastly exposes request_collapse_usable_count and request_collapse_unusable_count so the behaviour is measurable rather than assumed.
- Cloudflare: part of default cache behaviour, implemented as a cache lock whose scope is one location: the cache lock ensures that Cloudflare only sends one request at a time to the origin for a given asset from a single location in Cloudflare’s network. Cloudflare does not document collapsing across data centres. The documented way to narrow the funnel is Tiered Cache, which limits the number of data centers that can ask the origin for content.
- Squid: collapsed_forwarding is a directive with default value: collapsed_forwarding off. It is aimed at accelerator (reverse-proxy) deployments. Squid documents one real scope limit: revalidation collapsing is currently disabled for Squid instances containing SMP-aware disk or memory caches and for Vary-controlled cached objects.
- Apache Traffic Server: read-while-writer is on by default (proxy.config.cache.enable_read_while_writer, default 1). That alone is not collapsing, though. ATS ships an explicit collapsing mode in proxy.config.http.cache.open_write_fail_action. Its option 5 together with proxy.config.cache.enable_read_while_writer configuration allows to collapse concurrent requests without a need for any plugin, and option 6 adds a stale fallback. That setting’s default is 0, so collapsing is opt-in. The related open-read-retry path is also off by default (proxy.config.http.cache.max_open_read_retries is -1). It exists because the open read retry configurations attempt to reduce the number of concurrent requests to the origin for a given object. Note that the ATS admin guide still says read-while-writer is off by default. This contradicts its own records reference: the shipped default in RecordsConfig.cc is 1.
This is the nginx side of it:
# Nginx request coalescing configuration
proxy_cache_path /var/cache/nginx levels=1:2 keys_zone=cdn:10m;
upstream origin {
server 10.0.0.1:80; # the name in proxy_pass must resolve or be a server group
}
server {
location / {
proxy_cache cdn;
proxy_cache_lock on; # Enable request coalescing
proxy_cache_lock_timeout 5s; # Max wait, then this request goes to origin and is not cached
proxy_cache_lock_age 5s; # If the first fetch stalls this long, one more request is let through
proxy_cache_use_stale updating; # Serve the stale copy while an expired entry is updated
proxy_pass http://origin;
}
}
Watch out for
- A timeout shorter than the origin brings the herd back. In nginx, when proxy_cache_lock_timeout expires, the request will be passed to the proxied server, however, the response will not be cached. A slow origin then gets both the stampede and no cache entry from those requests. Before 1.7.8, the response could be cached. Do not carry old tuning advice across that boundary.
- The uncacheable response is the real failure mode. It can be worse than no coalescing. Fastly warns that if the origin response is not cacheable and no hit-for-pass marker can be created, the next request in the queue will be sent to origin and the remaining requests will form a new queue, resulting in the requests being sent consecutively, not concurrently. Fastly also warns that in some cases this can create extreme response times of several minutes. This is why Squid ships the feature off: enabling collapsed forwarding needlessly delays forwarding requests that look cachable (when they are collapsed) but then need to be forwarded individually anyway because they end up being for uncachable content. A one-line summary that queued requests are simply re-sent independently understates it. Whether they go concurrently or one at a time depends on a hit-for-pass marker being created.
- Releasing the queue is itself a burst. Varnish: suddenly releasing a thousand threads to serve content might send the load sky high. Coalescing moves the spike from the origin to your own edge. That is the right trade, but not a free one.
- Scope is one cache node or location, not the CDN. Cloudflare’s lock is per asset from a single location in Cloudflare’s network. Fastly’s collapsing happens per fetch server. An uncoordinated topology therefore means one queue per collapse domain. Vendor docs establish where collapsing does happen. They do not license a blanket claim that nothing else is merged anywhere.
- Vary is not the cache key. Confusing the two gives the wrong prediction. Coalescing merges requests that resolve to the same lookup address. Anything you put in the key, such as a query string, cookie, or device flag, splits the queue. A Vary mismatch is different: it is resolved inside one address. That is why Fastly says hit-for-pass objects respect the Vary header just as normal cache objects do, so it’s possible to have a single cache address in which some variants are hit-for-pass, and others are normal objects. Squid names its concrete limit instead of implying a general rule: revalidation collapsing is disabled for Vary-controlled cached objects.
- Wait-then-serialise on uncacheable objects. ATS is explicit that its open-read-retry settings are inappropriate when objects are uncacheable. In those cases, requests for an object effectively become serialized. Each one waits at least the retry interval before being proxied. Delay plus no collapse is the worst of both.
Best practice
- Size the wait from measured origin latency, not a round number. The timeout must exceed the origin’s slow-tail fetch time for the object class being protected. Otherwise expiry converts a protected miss into an uncached forwarded request. Set proxy_cache_lock_age deliberately too: it decides when a second request is allowed through while the first is still running.
- Put a stale-serving strategy in front of the queue. stale-while-revalidate means caches MAY serve the response in which it appears after it becomes stale, up to the indicated number of seconds. RFC 5861 is candid about the limit, though: if the window is too small, or traffic is too sparse, some requests will fall outside of it, and block until the server can validate the cached response. In nginx that is proxy_cache_use_stale updating plus proxy_cache_background_update, which allows starting a background subrequest to update an expired cache item, while a stale cached response is returned to the client. In Varnish it is grace. Coalescing then protects the cold and evicted cases. Nobody waits on the routine ones.
- Keep per-user responses out of the queue entirely. Mark them private, or better, flag them to bypass the cache before the fetch. Fastly advises that if you know before making a backend fetch that the response will not be cacheable, exclude it from request collapsing. Coalescing pays off on cacheable, widely shared objects: static assets, images, video segments, and hot HTML with a short TTL. On personalised responses, it costs latency and returns nothing instead.
- Shrink the number of collapse domains. Add an origin shield or a tiered topology so misses funnel through few nodes rather than every PoP. Fastly notes that clustering and shielding together create up to four opportunities for Fastly to collapse requests. Cloudflare’s Tiered Cache concentrates connections to origin servers so they come from a small number of data centers rather than the full set of network locations. This matters most on a cold cache and after a global purge.
- Keep the cache key tight. Coalescing can only merge requests that share the key. Cookies, tracking parameters, or a device dimension in the key fragment the queue into as many origin fetches as there are variants.
- Verify it, then watch it. Fire a burst of concurrent requests at one uncached URL. Count the requests the origin actually logged: it should be one, not the burst size. If the count is high, check three things: the lock is on, the timeout against origin latency, and the key for accidental uniqueness. In production, track the collapse counters your platform exposes (Fastly’s request_collapse_usable_count and request_collapse_unusable_count). Read a falling collapse rate against origin latency before treating it as a regression.
Interactive Animation
Examples
This Varnish VCL configures request coalescing:
sub vcl_backend_fetch {
# Varnish enables coalescing by default.
# When multiple clients request the same uncached URL,
# only the first request goes to the backend.
# Others wait for the response.
# You can control waiting behavior with timeouts:
set bereq.between_bytes_timeout = 10s;
set bereq.first_byte_timeout = 15s;
}# Test coalescing with concurrent requests
# Send 50 simultaneous requests for the same URL
seq 50 | xargs -P50 -I{} curl -s -o /dev/null -w "%{http_code}\n" \
https://cdn.example.com/popular-page
# With coalescing: origin sees 1 request
# Without coalescing: origin sees 50 requests
Frequently Asked Questions
A cache behaviour that serves many concurrent requests for one cache key from a single upstream fetch: the first miss goes to the origin, the rest wait on a queue and are answered from that one response. It exists to stop a cache stampede. Also called request collapsing or collapsed forwarding.
This Varnish VCL configures request coalescing:
sub vcl_backend_fetch {
# Varnish enables coalescing by default.
# When multiple clients request the same uncached URL,
# only the first request goes to the backend.
# Others wait for the response.
# You can control waiting behavior with timeouts:
set bereq.between_bytes_timeout = 10s;
set bereq.first_byte_timeout = 15s;
}# Test coalescing with concurrent requests
# Send 50 simultaneous requests for the same URL
seq 50 | xargs -P50 -I{} curl -s -o /dev/null -w "%{http_code}\n" \
https://cdn.example.com/popular-page
# With coalescing: origin sees 1 request
# Without coalescing: origin sees 50 requests
Yes. Request Coalescing is also known as request collapsing, collapsed forwarding. A cache behaviour that serves many concurrent requests for one cache key from a single upstream fetch: the first miss goes to the origin, the rest wait on a queue and are answered from that one response. It exists to stop a cache stampede. Also called request collapsing or collapsed forwarding.