Request Coalescing

Caching

A cache behaviour that serves many concurrent requests for one cache key from a single upstream fetch: the first miss goes to the origin, the rest wait on a queue and are answered from that one response. It exists to stop a cache stampede. Also called request collapsing or collapsed forwarding.

Also known as request collapsing, collapsed forwarding.

15 min read Updated Aug 30, 2026

Full Explanation

Request coalescing is a cache behaviour. On a miss, the cache sends one request upstream for a given cache key. It then answers every other request for that key from that same response while the fetch is in flight. Fastly calls it request collapsing and defines it as combining multiple requests for the same object into a single request to origin, and then potentially using the resulting response to satisfy all pending requests. The word potentially is load-bearing: the queue is only satisfied if the response turns out to be usable. Its purpose is to stop a cache stampede. Fastly again: it prevents the expiry of a very highly demanded object in the cache causing an immediate flood of requests to an origin server, which might otherwise overwhelm it or consume expensive resources. One fetch reaches the origin instead of thousands.

It is not a protocol feature. It is not a Cache-Control directive either. No client asks for it, and no origin switches it on with a header. It is an implementation choice inside a cache server. That is why every platform names it differently: request collapsing (Fastly, Cloudflare), collapsed forwarding (Squid), proxy_cache_lock (nginx), read-while-writer plus cache-read retry (Apache Traffic Server). Those names are close but not interchangeable: ATS documents that once its read-while-writer settings are enabled you have something that is very close, but not quite the same, to Squid’s Collapsed Forwarding. It is also not stale-while-revalidate. Coalescing makes the waiting clients wait for a fresh fetch. Stale-while-revalidate answers them immediately from the stale copy instead. The two compose, and in production they usually should. Finally, it is not global. The merge happens inside one cache node or location, not across a CDN.

How it works

Coalescing is one fetch plus a queue, scoped to a single cache key. nginx states the rule most plainly. With proxy_cache_lock on, only one request at a time will be allowed to populate a new cache element identified according to the proxy_cache_key directive by passing a request to a proxied server. The rest either wait for a response to appear in the cache or the cache lock for this element to be released. Fastly calls that queue a waiting list. Cloudflare calls the mechanism a cache lock.

The window is narrow. It is defined by the fetch, not by the clock. Fastly: to collapse requests they must be concurrent with the request that initiated the fetch from origin. Squid is exact about what concurrent means: received after the first request headers were parsed and before the corresponding response headers were parsed. Two mechanisms widen that window past the response headers. Fastly’s streaming miss writes the partial response to cache as soon as the headers arrive. A later request then joins the in-progress stream if the object is still fresh. If the object has already gone stale, that request is a miss instead, and it starts a new fetch. ATS’s read-while-writer is the same idea: the ability to read a cached object while another connection is completing the write to cache for that same object. But it starts late on purpose. ATS does not begin allowing clients to read until after the complete HTTP response headers have been read and processed, because until then it cannot know the response will be cacheable.

Two preconditions therefore apply. First, the request must be able to produce a cache object at all. Fastly: PASS requests and many errors are uncacheable by default, which means that we will never be able to successfully collapse those requests. Second, the resulting object must be usable. That means cacheable with a positive remaining lifetime.

One distinction is easy to miss: a cold miss and a stale revalidation are not the same case. Squid collapses two kinds of requests: regular client requests received on one of the listening ports and internal “cache revalidation” requests which are triggered by those regular requests hitting a stale cached object. nginx’s lock, by its own wording, guards only the population of a new cache element. The expiry case is covered by separate directives instead: proxy_cache_use_stale updating and proxy_cache_background_update. On a hot object that expires every few seconds, that difference decides whether the herd is absorbed at all.

Why it matters for a CDN

The size of the herd is set by request rate multiplied by fetch duration, not by how many users exist. Fastly’s own arithmetic: for an object requested 50 times per second, with a 500ms fetch latency, the origin would be processing 25 concurrent requests for the same object before Fastly had the opportunity to store it in cache. That is one object, on one PoP, at a modest rate. Coalescing turns those 25 into one. That is why Fastly can say that for high traffic services, correct use of request collapsing will substantially reduce and smooth out traffic to origin servers. What it saves is origin concurrency, connections and egress, not fetch time.

The corollary is counter-intuitive. It is worth internalising before you tune anything. Fastly: somewhat counter-intuitively, faster origins and features that improve time to first byte will reduce how often we can collapse requests. The reverse also applies: you may also see a higher occurrence of request collapsing when an origin takes longer to respond. Coalescing is a shock absorber that engages exactly when the origin is struggling. A falling collapse rate is usually good news about the origin, not a broken feature.

The cost is paid by the clients in the queue. Varnish puts it bluntly: at high rates the queue of waiting requests can get huge. Two problems follow: a thundering herd on release, and nobody likes to wait. Coalescing converts an origin-capacity problem into a client tail-latency problem. It belongs next to a stale-serving strategy, not on its own.

Topology decides how many separate queues your origin faces. Each cache node collapses only its own traffic. Concentrating misses is what makes coalescing effective across a global network. Fastly notes that disabling clustering has the potential to significantly increase traffic to origin not just due to a poorer cache hit ratio, but also due to the inability to collapse concurrent requests that originate on different delivery servers. Fastly adds that with shielding, requests from POPs across the network are focused into a single POP, allowing it to perform request collapsing before forwarding a single request to origin. Cloudflare’s Tiered Cache does the same by hierarchy: if the upper-tier does not have the content, only the upper-tier can ask the origin for content.

What CDNs do

This is the nginx side of it:

# Nginx request coalescing configuration
proxy_cache_path /var/cache/nginx levels=1:2 keys_zone=cdn:10m;

upstream origin {
    server 10.0.0.1:80;   # the name in proxy_pass must resolve or be a server group
}

server {
    location / {
        proxy_cache cdn;
        proxy_cache_lock on;           # Enable request coalescing
        proxy_cache_lock_timeout 5s;   # Max wait, then this request goes to origin and is not cached
        proxy_cache_lock_age 5s;       # If the first fetch stalls this long, one more request is let through
        proxy_cache_use_stale updating;  # Serve the stale copy while an expired entry is updated
        proxy_pass http://origin;
    }
}

Watch out for

Best practice

Interactive Animation

Loading animation...

Examples

This Varnish VCL configures request coalescing:

sub vcl_backend_fetch {
    # Varnish enables coalescing by default.
    # When multiple clients request the same uncached URL,
    # only the first request goes to the backend.
    # Others wait for the response.
    
    # You can control waiting behavior with timeouts:
    set bereq.between_bytes_timeout = 10s;
    set bereq.first_byte_timeout = 15s;
}
# Test coalescing with concurrent requests
# Send 50 simultaneous requests for the same URL
seq 50 | xargs -P50 -I{} curl -s -o /dev/null -w "%{http_code}\n" \
  https://cdn.example.com/popular-page
# With coalescing: origin sees 1 request
# Without coalescing: origin sees 50 requests

Frequently Asked Questions

A cache behaviour that serves many concurrent requests for one cache key from a single upstream fetch: the first miss goes to the origin, the rest wait on a queue and are answered from that one response. It exists to stop a cache stampede. Also called request collapsing or collapsed forwarding.

This Varnish VCL configures request coalescing:

sub vcl_backend_fetch {
    # Varnish enables coalescing by default.
    # When multiple clients request the same uncached URL,
    # only the first request goes to the backend.
    # Others wait for the response.
    
    # You can control waiting behavior with timeouts:
    set bereq.between_bytes_timeout = 10s;
    set bereq.first_byte_timeout = 15s;
}
# Test coalescing with concurrent requests
# Send 50 simultaneous requests for the same URL
seq 50 | xargs -P50 -I{} curl -s -o /dev/null -w "%{http_code}\n" \
  https://cdn.example.com/popular-page
# With coalescing: origin sees 1 request
# Without coalescing: origin sees 50 requests

Yes. Request Coalescing is also known as request collapsing, collapsed forwarding. A cache behaviour that serves many concurrent requests for one cache key from a single upstream fetch: the first miss goes to the origin, the rest wait on a queue and are answered from that one response. It exists to stop a cache stampede. Also called request collapsing or collapsed forwarding.