Cache Stampede
A self-inflicted flood of origin requests: a hot cached object expires or is evicted, every concurrent request for it misses at the same moment, and the origin takes that object's whole request rate in one burst. Request coalescing and stale-while-revalidate prevent it.
Also known as thundering herd, dog-piling.
Full Explanation
A cache stampede is a self-inflicted flood of origin requests. A popular cached object expires or is evicted. Every concurrent request for it becomes a miss at the same moment. The origin then receives that object’s whole request rate in one burst. It is also called a thundering herd and dog-piling. A thundering herd is the general name for many clients hitting one resource at once. Wikipedia classes it as a cascading failure. If the burst is large enough, no attempt to regenerate the object completes. So the object is never re-cached, and the herd keeps arriving.
It is not an attack, and it is not a traffic spike. Client demand has not changed. Only the cache’s ability to answer it has changed. The cache that normally absorbs load amplifies it instead. It is also not the same problem as many unrelated keys expiring on one boundary. That problem is cured by staggering TTLs. One hot key’s expiry is cured differently: by making its misses share a single origin fetch. The two standard defences are request coalescing and stale-while-revalidate. Request coalescing sends one request upstream and serves every waiting client from that response. Stale-while-revalidate keeps serving the stale copy while one background revalidation runs.
How it works
The trigger is expiry or eviction of a hot object while requests for it are in flight. A cache answers from storage for as long as the object is fresh, that is, for its TTL. The instant it is stale or gone, lookups stop answering. Every arriving request then becomes a miss that must be forwarded. CloudFront states the base behaviour plainly: when the object is not in the cache or the cached object is expired, it sends the request to the origin immediately.
The size of the burst is set by request rate multiplied by origin fetch time, not by how many clients exist. Here is the worked example on Wikipedia: a page that takes 3 seconds to render, receiving 10 requests per second, has 30 processes recomputing it simultaneously the moment the cached copy expires. Halve the render time and you halve the herd. Double the rate and you double it.
Three independent defences shrink or remove that burst:
- Request coalescing (request collapsing) combines multiple requests for the same object into a single request to origin. It then uses that response to satisfy the pending ones. Fastly calls the queue a “waiting list”. nginx implements the idea as proxy_cache_lock, which allows only one request at a time to populate a new cache element. Coalescing does not make the fetch faster. It makes the origin serve one request instead of thousands.
- Stale-while-revalidate is a Cache-Control extension defined in RFC 5861 section 3. Caches MAY serve a response after it becomes stale, up to the number of seconds given. They SHOULD revalidate it without blocking. Clients get a fast, slightly stale response, and the origin handles one refresh per window. Its sibling stale-if-error (section 4) covers the other half: serve stale rather than pass on a 500, 502, 503 or 504.
- TTL jitter randomises expiry times so keys do not all expire on the same instant. Redis gives the example of a random TTL between 55 and 65 minutes where a nominal one hour is wanted, so that cached items do not all vanish simultaneously. Jitter spreads different keys apart. It does nothing for a single hot key, whose expiry still produces one burst.
Why it matters for a CDN
Edge caching concentrates demand, so a handful of objects can be most of an edge location’s traffic. One expiry then converts a large share of the request rate into origin-bound misses. The same arithmetic bites before an object is ever cached. Fastly’s figure: for an object requested 50 times per second with a 500 ms fetch latency, the origin would be processing 25 concurrent requests for the same object before Fastly had the opportunity to store it in cache.
A CDN also multiplies the fan-in, because uncollapsed misses reach the origin from every server that accepted a request. Fastly notes that disabling clustering can significantly increase origin traffic. This can happen not only through a poorer cache hit ratio but through the inability to collapse concurrent requests that originate on different delivery servers. With shielding in place, requests from POPs across the network are focused into a single POP that collapses them before forwarding one request to origin. That makes an origin shield or a tiered cache a stampede defence, not merely a hit-ratio tool.
What CDNs do
- CloudFront: collapsing is built in and on by default, documented identically for custom origins and for Amazon S3 origins. Simultaneous requests sharing a cache key that arrive before the first response is received are paused. They are then all served from that response. In the logs the first request is a Miss, and the collapsed ones are a Hit in x-edge-result-type. See Simultaneous requests for the same object (request collapsing).
- Fastly: cache misses qualify for collapsing by default in both VCL and Compute services when using the readthrough or simple cache interfaces. The core cache interface supports collapsing only when explicitly configured within a cache transaction. Queued clients are then served from the one origin response. See Request collapsing.
- Cloudflare: the zone cache collapses by default. On simultaneous misses for the same asset in one data center, a cache lock forwards only the first request to origin and streams that response to the waiting ones. Workers Cache is opt-in (cache.enabled in the Wrangler configuration). It applies the same collapsing per cache key per data center to Worker responses, honouring
stale-while-revalidate. The Cache API does neither: it does not collapse concurrent requests, andstale-while-revalidateand stale-if-error are not supported by cache.put or cache.match. - Self-hosted caches: Varnish coalesces automatically. Several clients requesting the same page produce one backend request while the others are held. A positive grace serves the object after its TTL has expired while a new version is fetched (the default comes from the default_grace runtime parameter). nginx must be opted in: proxy_cache_lock is off by default.
stale-while-revalidatebehaviour needs proxy_cache_use_stale updating plus proxy_cache_background_update on, also off by default.
Watch out for
- Cache-key fragmentation silently disables collapsing. Only requests that share a cache key are collapsed. CloudFront forwards every request with a unique cache key to the origin. So caching on a header, cookie or query string that varies per request leaves you with no protection at all.
- An uncacheable response turns the queue into a serial trickle. Fastly warns that if the collapsed origin response cannot be cached, it satisfies only the request that triggered the fetch. The rest re-queue behind the next request. Requests then go to origin consecutively rather than concurrently. In some cases this can create extreme response times of several minutes. Mark single-user responses private. Pass known-uncacheable requests around the cache (or let a hit-for-pass object form) so no waiting list builds.
- Bound the wait. nginx’s proxy_cache_lock_timeout defaults to 5 s. When it expires, the request is passed to the proxied server. However, the response will not be cached. Fastly suggests failing long-queued requests fast in vcl_miss, for example erroring when time.elapsed > 1s. This way waiting lists dissolve quickly while an origin is unhealthy.
- Directives that suppress stale serving. Cloudflare serves stale content during revalidation only if the origin includes
stale-while-revalidate. With Origin Cache Control enabled, must-revalidate, proxy-revalidate, s-maxage or no-cache sent alongside it stop stale serving. Requests then return EXPIRED instead of UPDATING. See Revalidation. - The stale-while-revalidate window is finite. Per RFC 5861 section 3.1, if the window is too small or traffic too sparse, some requests fall outside it and block until the server can validate the cached response.
- Do not refresh on a timer. RFC 5861 section 5 suggests predicating background validation on an incoming request, to avoid the possibility of an amplification attack.
- Turning collapsing off restores the hazard. CloudFront’s documented ways to prevent it are the CachingDisabled managed cache policy, or a minimum TTL of 0 together with an origin Cache-Control of private, no-store, no-cache, max-age=0 or s-maxage=0. Both increase load on the origin and add latency for the simultaneous requests that were paused. A header alone, without the minimum TTL of 0, is not what AWS documents.
Best practice
- Verify collapsing is actually enabled on every cache in front of your origin. The defaults differ: on for CloudFront and for Cloudflare’s zone cache, on for Fastly’s readthrough and simple cache interfaces but explicit for core, opt-in for Cloudflare Workers Cache, and off in nginx until proxy_cache_lock on.
- Pair coalescing with
stale-while-revalidate, and add stale-if-error so a stampede that reaches a struggling origin does not turn into an error storm. RFC 5861 advises setting max-age plusstale-while-revalidateto the longest total potential freshness lifetime you can tolerate. - Jitter TTLs across keys so distinct hot objects do not share an expiry boundary.
- Refresh hot keys before they expire instead of waiting for the first miss. The standard formulation is probabilistic early recomputation. Each request may refresh the key with a probability that rises as expiry approaches. This way refreshes spread out rather than landing together on the boundary.
- Put a shield or tiered cache between the edge and the origin so per-location collapsing fans into a single origin fetch.
- Keep cache keys as narrow as the content allows: fewer distinct keys means more requests are eligible to collapse, and a hotter cache to begin with.
Examples
Varnish coalesces requests on its own:
# varnish default.vcl - coalescing is automatic
# But you can control grace period for stale serving
sub vcl_backend_response {
# Keep stale copy for 1 hour after TTL expires
set beresp.grace = 1h;
}
sub vcl_hit {
# Serve stale while revalidating in background
if (obj.ttl <= 0s && obj.grace > 0s) {
return (deliver);
}
}
Nginx locks the proxy cache the same way:
proxy_cache_lock on;
proxy_cache_lock_timeout 5s;
proxy_cache_lock_age 5s;
# Combined with stale-while-revalidate
proxy_cache_use_stale updating;
proxy_cache_background_update on;
Frequently Asked Questions
A self-inflicted flood of origin requests: a hot cached object expires or is evicted, every concurrent request for it misses at the same moment, and the origin takes that object's whole request rate in one burst. Request coalescing and stale-while-revalidate prevent it.
Varnish coalesces requests on its own:
# varnish default.vcl - coalescing is automatic
# But you can control grace period for stale serving
sub vcl_backend_response {
# Keep stale copy for 1 hour after TTL expires
set beresp.grace = 1h;
}
sub vcl_hit {
# Serve stale while revalidating in background
if (obj.ttl <= 0s && obj.grace > 0s) {
return (deliver);
}
}
Nginx locks the proxy cache the same way:
proxy_cache_lock on;
proxy_cache_lock_timeout 5s;
proxy_cache_lock_age 5s;
# Combined with stale-while-revalidate
proxy_cache_use_stale updating;
proxy_cache_background_update on;
Yes. Cache Stampede is also known as thundering herd, dog-piling. A self-inflicted flood of origin requests: a hot cached object expires or is evicted, every concurrent request for it misses at the same moment, and the origin takes that object's whole request rate in one burst. Request coalescing and stale-while-revalidate prevent it.
Related CDN concepts include:
- Origin Shield — A cache tier a CDN places between its edge servers and its origin. Cache misses …
- Request Coalescing — A cache behaviour that serves many concurrent requests for one cache key from a single …
- stale-while-revalidate — A Cache-Control response directive (RFC 5861) that lets a cache keep serving a stale stored …