Cold Cache
A cold cache holds no reusable stored responses, so requests miss and go to the origin until misses refill it. New edges and POPs, restarts of volatile storage, purges and evictions cause it; while cold, origin load and latency rise. A transient state of one store, not a failure.
Full Explanation
A cold cache is a cache store with no reusable stored responses. Requests find nothing to serve. They go to the origin instead. This continues until a run of misses refills the store. A cold cache is the transient opposite of a warm cache. It is not a failure, a separate component, or a cache miss type. A cold miss is one request. A cold cache is the state that produces a long run of them. Coldness is per store. Every edge server and every POP keeps its own cache. So one location can be stone cold while the rest of the network stays warm. The miss-then-store cycle that ends it is cache fill. Nothing else warms a cache: either traffic arrives, or you pre-warm deliberately.
How it works
HTTP defines a cache as "a local store of response messages and the subsystem that controls storage, retrieval, and deletion of messages in it" (RFC 9111 section 1). "When presented with a request, a cache MUST NOT reuse a stored response unless" the request matches a stored response that is fresh, is allowed to be served stale, or is successfully validated (RFC 9111 section 4). A cold store meets none of those conditions, so each request runs the full cycle:
- Miss. "If no stored response matches, the cache cannot satisfy the presented request. Typically, the request is forwarded to the origin server" (RFC 9111 section 4.1).
- Store, but only if permitted. "A cache MUST NOT store a response to a request unless" a list of conditions holds, among them that the response carries "an Expires header field", "a max-age response directive", or "a status code that is defined as heuristically cacheable" (RFC 9111 section 3). Some responses are not permitted to be stored. They never warm at all: they miss on every request, forever.
- Hit. "When a response is fresh, it can be used to satisfy subsequent requests without contacting the origin server, thereby improving efficiency" (RFC 9111 section 4.2). Warming is only this cycle repeating, once per object per cache. KeyCDN states the practical floor: "the resource needs to be requested at least twice from the same server in order to ensure the asset is cached" (KeyCDN troubleshooting guide).
- Collapse. Concurrent misses for one object need not each reach the origin. A cache may "combine multiple incoming requests into a single forward request upon a cache miss -- thereby reducing load on the origin server and network". But RFC 9111 adds that "if the cache cannot use the returned response for some or all of the collapsed requests, it will need to forward the requests in order to satisfy them, potentially introducing additional latency" (RFC 9111 section 4).
A store ends up cold for a handful of concrete reasons:
- A new location. A fresh edge or POP starts empty: "Every new file needs to be cached on each POP all around the world. This will result in one cache miss for each file and POP" (KeyCDN).
- A restart of volatile storage. Varnish stores objects in memory by default: "If the option is omitted, the malloc storage backend will be used, which stores objects in memory". Even its disk-backed file storage is wiped on restart: "Although disk storage is used for this kind of object storage, the file stevedore is not persistent. A restart will empty the entire cache" (Varnish Software, Configuring Varnish).
- A purge. "Each purge (either a particular file or the entire Zone) will delete the content on all POPs. The file(s) will then be cached again with the first request" (KeyCDN).
- Eviction under pressure. A full store makes room by dropping entries. It usually drops the least recently used ones first: Varnish "starts to remove the least recently used (LRU) objects" (Varnish Software). An evicted object is cold again even though nothing was purged and nothing restarted.
- A topology or cache key change. Content stays cached only where the lookup still finds it. Cloudflare warns that with Smart Tiered Cache, updating origin IPs or DNS records "may cause the existing assigned upper tiers to change, resulting in an increased MISS rate as cache is refilled in the new upper tiers" (Cloudflare, Tiered Cache).
- Long-tail demand. Some workloads never finish warming. With user-generated content, "new files need to be cached more often" because of the long tail, "which will lead to a lower CHR" (KeyCDN).
Why it matters for a CDN
"The goal of HTTP caching is significantly improving performance by reusing a prior response message to satisfy a current request" (RFC 9111 section 1). A CDN is that reuse spread across many locations. A cold cache suspends the whole benefit for as long as it lasts:
- Origin load. Misses are origin requests. Each location warms separately. So a network-wide cold start multiplies the same origin fetch by the number of caches that need the object.
- Latency. The first request for each object at each location pays a full trip to the origin, not an edge hit. Fastly describes exactly this in its video prefetch work: "if the cache is cold the request would go all the way back to origin for each segment" (Fastly blog).
- What you see. A cold cache shows up as a collapsed cache hit ratio, a run of misses, and a matching spike in origin requests. That is why hit ratio is the metric to watch after any deploy, restart or purge, not before.
What CDNs do
None of the three vendors below documents automatically pre-filling an empty cache. Warming still comes from traffic or an explicit pre-warm. What they do supply is a middle tier: a cold edge asks a warmer upstream instead of the origin. This is the origin shield pattern:
- Cloudflare Tiered Cache (opt-in toggle or API; listed as available on Free, Pro, Business and Enterprise). "Tiered Cache works by dividing Cloudflare’s data centers into a hierarchy of lower-tiers and upper-tiers". On a miss, "only the upper-tier can ask the origin for content", "which reduces origin load and makes websites more cost-effective to operate" (Cloudflare docs). Cloudflare now presents the Smart and Regional topologies as part of its Smart Shield product.
- Fastly Shielding (opt-in per origin: "shielding may be enabled when adding or editing an origin server, and may be selected per-origin"). Requests "funnel through a single, designated shield POP", which "can reduce upstream load" and "increases the probability of end user requests resulting in a cache HIT (albeit potentially not from the first POP which handles the request)" (Fastly docs). This is the precise value of a shield when the edge is cold.
- AWS CloudFront Origin Shield (opt-in per origin, chargeable: "Origin Shield is a property of the origin" and "You incur additional charges for using Origin Shield"). It is "an additional layer in the CloudFront caching infrastructure that helps to minimize your origin’s load". Also, "Requests for content that is not in Origin Shield’s cache are consolidated with other requests for the same object, resulting in as few as one request going to your origin" (CloudFront docs).
Watch out for
- A full purge is a planned cold start. Cloudflare’s Purge Everything "instantly clears all resources from your CDN cache in all Cloudflare data centers". The origin impact is conditional, not universal: "When a site with heavy traffic contains a lot of assets, requests to your origin server can increase substantially and result in slow site performance." It invalidates rather than deletes: "Purge Everything invalidates the resource, resulting in the CF-Cache-Status header indicating EXPIRED for subsequent requests" (Cloudflare docs). So the origin is still contacted. But a successful revalidation can avoid re-transferring the body.
- Persistent storage is the exception, not the rule. A restart empties an in-memory or file-backed Varnish cache. Only a persistent stevedore survives one. Varnish’s own answer is Enterprise-only: "The Massive Storage Engine (MSE) is a Varnish Enterprise stevedore that combines memory and disk storage to offer fast and persistent storage" (Varnish Software). Do not assume your cache tier keeps anything across a restart. Check its storage backend first.
- stale-while-revalidate cannot rescue an empty cache. It "allows a cache to immediately return a stale response while it revalidates it in the background, thereby hiding latency (both in the network and on the server) from clients" (RFC 5861 section 1). What it permits is this: "caches MAY serve the response in which it appears after it becomes stale" (RFC 5861 section 3). Both need a stored copy. A truly cold store has none. So it still blocks on the origin.
- Cold caches invite cache stampedes. Straight after a restart or purge, many clients ask for the same uncached object at once. Request collapsing limits the spike. It does not remove it. It degrades to separate forwarded requests when the response cannot be shared (RFC 9111 section 4).
- A shield does not fix everything a cold cache exposes. AWS notes Origin Shield "may not be a good fit in other cases, such as dynamic content that is proxied to the origin, content with low cacheability, or content that is infrequently requested" (CloudFront docs). Low-cacheability traffic reaches the origin cold or warm.
Best practice
- Pre-warm the objects that matter before traffic arrives: request your top URLs at each location after a deploy, restart or purge. Fastly documents the edge-side version of this for video. It concludes that "By using Compute to pre-warm the cache, you are not only using a powerful, globally distributed network to do the work, but you also solve the pitfalls associated with legacy prefetch" (Fastly blog).
- Run a tiered cache or origin shield. This way, a cold edge asks a warm middle tier. Concurrent misses then consolidate into as few as one origin request.
- Set stale-while-revalidate on cacheable responses so warm-but-expiring objects never block a client. Add stale-if-error too. Then a stale copy "MAY be used to satisfy the request, regardless of other freshness information" when the origin fails under the extra load (RFC 5861 section 4). Both protect the warm state. Neither substitutes for it.
- Purge narrowly. "To maintain optimal site performance, Cloudflare strongly recommends using single-file (by URL) purging instead of a complete cache purge" (Cloudflare docs).
- Give responses an explicit, generous freshness lifetime. Keep cache keys stable. Every avoidable key change or short TTL re-cools content you already paid to fetch.
- Instrument before you cut over. Alert on hit ratio and origin request rate. Stage a rollout so a cold network does not meet peak traffic.
Examples
Warm the cache after a deploy:
#!/bin/bash
# Crawl top 100 URLs to warm the cache
while IFS= read -r url; do
curl -s -o /dev/null -w "%{http_code} %{url}\n" "$url" &
done < top-urls.txt
wait
echo "Cache pre-warm complete"
Watch the cache recover in Varnish:
# Watch hit rate recover after restart
watch -n 1 'varnishstat -1 | grep -E "(MAIN.cache_hit|MAIN.cache_miss)"'
# Output right after restart (cold):
# MAIN.cache_hit 12 0.40/s
# MAIN.cache_miss 988 32.93/s
# After 10 minutes (warming up):
# MAIN.cache_hit 8523 284.10/s
# MAIN.cache_miss 1477 49.23/s
Frequently Asked Questions
A cold cache holds no reusable stored responses, so requests miss and go to the origin until misses refill it. New edges and POPs, restarts of volatile storage, purges and evictions cause it; while cold, origin load and latency rise. A transient state of one store, not a failure.
Warm the cache after a deploy:
#!/bin/bash
# Crawl top 100 URLs to warm the cache
while IFS= read -r url; do
curl -s -o /dev/null -w "%{http_code} %{url}\n" "$url" &
done < top-urls.txt
wait
echo "Cache pre-warm complete"
Watch the cache recover in Varnish:
# Watch hit rate recover after restart
watch -n 1 'varnishstat -1 | grep -E "(MAIN.cache_hit|MAIN.cache_miss)"'
# Output right after restart (cold):
# MAIN.cache_hit 12 0.40/s
# MAIN.cache_miss 988 32.93/s
# After 10 minutes (warming up):
# MAIN.cache_hit 8523 284.10/s
# MAIN.cache_miss 1477 49.23/s
Related CDN concepts include:
- Cache Fill — A cache fill is the fetch-and-store after a cache miss: the cache pulls the resource …
- Cache Miss Types — A cache miss is a request an edge node cannot answer from its own cache. …