Warm Cache

Caching

A cache whose store already holds the responses its current traffic asks for, so most requests are answered from the store instead of the origin. Warmth belongs to one store measured against demand, not to any single object, and it is not freshness. It cools on purge, eviction or restart.

10 min read Updated Aug 30, 2026

Full Explanation

A warm cache is a cache whose store already holds the responses its current traffic asks for. Most requests are then answered out of the store instead of being forwarded to the origin. Warmth is a property of one store measured against the demand reaching it. It is not a property of any single cached object. It is also not the same as HTTP freshness. A single entry can be perfectly fresh in a store that is cold overall. A warm store can also be full of entries that have gone stale. Warmth is not a feature you switch on either. It is the state a cache arrives at by serving traffic, and it is the opposite of a cold cache. Warmth is per store, so it is never global. Every edge and every tier warms on its own traffic. The same network can be warm in one location and cold in the next.

How it works

A cache warms through ordinary traffic, one miss at a time. There is no separate loading step. Each request is an opportunity for cache fill. HTTP states the reuse rule negatively: "When presented with a request, a cache MUST NOT reuse a stored response unless" the presented target URI and that of the stored response match; the stored response's method allows it to be used; the request header fields nominated by the stored response match those presented; no-cache does not apply; and the stored response is fresh, allowed to be served stale, or successfully validated (RFC 9111 section 4). The first conditions are the lookup: the cache key. A response's Vary header field extends that key to the request fields it nominates (RFC 9111 section 4.1). Matching is therefore not sufficient for a hit. The last condition decides whether the match can actually be served.

When nothing matches, "the cache cannot satisfy the presented request. Typically, the request is forwarded to the origin server" (RFC 9111 section 4.1). The returned response is then stored, so the next request for the same thing hits. Fastly describes exactly this loop: "The first time a cacheable resource is requested at a particular POP, the resource will be requested from your backend server and stored in cache automatically. Subsequent requests for that resource can then be satisfied from cache without having to be forwarded to your servers" (Fastly caching documentation). Repeat that per object and per POP, and the store becomes warm.

How long an entry stays usable without a trip upstream is its freshness lifetime. A fresh response is "one whose age has not yet exceeded its freshness lifetime". While it is fresh, "it can be used to satisfy subsequent requests without contacting the origin server" (RFC 9111 section 4.2). Origins normally set that lifetime explicitly. They use a Cache-Control max-age directive or an Expires field: the TTL. Where no explicit expiration time is available, a cache is allowed to calculate a heuristic one instead (RFC 9111 section 4.2.2). That is why an object served with no cache headers at all can still be held warm by a CDN's default.

Expiry does not empty the store. When a stored response goes stale, the cache can revalidate it with a conditional request. On a 304, the cache "MUST update its header fields with the header fields provided in the 304 (Not Modified) response" (RFC 9111 section 4.3.4). The entry is freshened in place with no body transferred, so warmth survives expiry. What actually removes content is a purge, a restart of volatile storage, or eviction. Fastly is explicit about the last one: "All data stored in the Fastly cache is ephemeral: it will expire, and may be evicted by the platform before it expires depending on how frequently it is used" (Fastly caching documentation). Warmth is the running balance of those fills, freshenings and losses. The figure that reports it is the cache hit ratio.

Why it matters for a CDN

  • Offload is the product. Every hit is a request the origin never sees. Warmth is what turns a network of caches into reduced origin bandwidth, reduced origin compute, and a response that never pays the trip upstream.
  • The cold start is the built-in cost of the design. Each POP fills from its own traffic, so a new deployment or a full flush means "every POP has to individually fetch the resources from your origin". The resulting "spike in traffic to your origin is temporary and will only last as long as it takes to populate the cache" (Fastly CDN tutorial).
  • Warmth is uneven inside one deployment. A cache is warm for its popular corpus and cold for its long tail at the same moment. This is the main argument for an origin shield. One shared tier stays warm, so cold edges pull from it instead of from the origin.
  • It is a measurement problem as much as a design one. A hit ratio only means something at a named unit: this POP, this shield, this zone, because that unit is the store whose warmth you are asking about.

What CDNs do

  • Fastly. The read-through HTTP cache is the only cache interface available to CDN services. It "works without any configuration or code required", so a service warms POP by POP as requests arrive. Request collapsing "allows us to identify multiple simultaneous requests for the same resource, and make just one backend fetch for it, using the resulting response to populate the cache and satisfy all waiting clients" (Fastly caching documentation). This is the standard's own "collapse requests" behaviour (RFC 9111 section 4). It caps how much origin traffic one cold object can generate. Shielding is a separate, opt-in setting configured on the origin. It "allows you to designate a single POP to handle all of the requests to your origin". Fastly warns in the same place that "Shielding can affect traffic, hit ratios, and performance" (Fastly CDN tutorial).
  • KeyCDN. The Cache Hit Ratio is surfaced in the dashboard. It is "measured among all files served from a Pull Zone regardless of the file size or file type". KeyCDN states that "A typical CHR for a normal website can easily be as high as 95%". It also warns on the same page that "The CHR can significantly vary depending on the environment and the setup". So treat 95% as an attainable figure for an ordinary static site, not a number every workload should be held to (KeyCDN cache hit ratio). Origin Shield is offered free to all customers as "an extra caching layer which will reduce the load on your origin server even further". It is enabled in the Zone settings. KeyCDN is candid about its price: it "adds an additional request from the edge server to the shield server if the content has not yet been cached" (KeyCDN CDN cache).

Watch out for

  • Warm does not mean up to date. A hit only tells you the store held a usable entry. Set the TTL too long, and the cache keeps serving an entry that is still fresh by the protocol clock while the origin has moved on. Serving an entry that has actually gone stale is a different and narrower case. A cache "MUST NOT generate a stale response unless it is disconnected or doing so is explicitly permitted by the client or origin server" (RFC 9111 section 4.2.4). Warmth and correctness are separate levers.
  • A high overall hit ratio hides a cold long tail. A handful of very popular objects can carry the average while most distinct URLs still miss to origin. Read the ratio next to absolute origin request volume rather than on its own.
  • Some content never warms. "Every new file needs to be initially fetched and cached from each POP all around the world. This will result in one cache miss for each file and POP". With user-generated content, "new files need to be cached all the time (which will lead to a lower CHR)" (KeyCDN cache hit ratio). For a large catalogue of rarely-requested objects, a cold edge is the normal state, not a fault to fix.
  • A purge is deliberate cooling. On KeyCDN, "Each purge (either the entire Zone or one particular file) will delete the content on all POPs globally. The file(s) will then be cached again" (KeyCDN cache hit ratio). For a hot object, the refill demand arrives all at once. That is how a routine flush becomes a cache stampede against the origin.
  • Fleet-wide numbers hide single cold stores. Warmth is per store, so a healthy aggregate can sit on top of one node that has just restarted empty. Alert per POP or per tier, not only on the global average.

Best practice

  • Set freshness explicitly. Give every cacheable response a Cache-Control max-age or an Expires value chosen for that content, rather than leaving the cache to compute a heuristic lifetime. Too short, and the store keeps re-fetching what it already had. Too long, and you cannot correct a mistake without a purge.
  • Cover the expiry gap with stale-while-revalidate. The extension tells caches they MAY serve a response after it becomes stale, "up to the indicated number of seconds". A cache serving stale for this reason "SHOULD attempt to revalidate it while still serving stale responses (i.e., without blocking)" (RFC 5861 section 3). Choose the window deliberately. It is a bounded permission to be out of date, not a free pass. Its companion stale-if-error "allows a cache to return a stale response when an error ... is encountered" (RFC 5861 section 1). So the same warm store can absorb an origin failure.
  • Purge surgically. "Purging can be expensive, so it's best to purge only the content that has changed and nothing else" (Fastly CDN tutorial). Prefer per-URL or key-scoped invalidation over a full flush. A full flush throws away warmth you paid origin traffic to build.
  • Put a warm tier behind the cold ones. A shield or mid-tier means a miss at a cold edge is usually answered by a warm upstream instead of the origin. This costs one extra internal hop when the shield misses too.
  • Monitor warmth per store, and read a drop as a signal. A sudden fall in one POP's hit ratio points at a leaked purge, a cache-key change that split the store, or a shift in traffic mix. All of these are far cheaper to catch as a hit-ratio anomaly than as origin saturation.

Try the interactive Cache Hit vs Miss animation in the course to compare response times between warm cache hits and cold cache misses.

Examples

Read the cache status in the response headers:

# Warm cache hit
curl -sI https://cdn.example.com/logo.png | grep -i x-cache
# X-Cache: HIT
# Age: 3542
# X-Cache-Hits: 847

# Cache miss (first request or expired)
curl -sI https://cdn.example.com/rare-page.html | grep -i x-cache
# X-Cache: MISS
# Age: 0

Watch the hit ratio in Varnish:

# Real-time hit rate
varnishstat -f MAIN.cache_hit -f MAIN.cache_miss

# Calculate hit ratio
varnishstat -1 -j | python3 -c "
import json, sys
d = json.load(sys.stdin)['counters']
hits = d['MAIN.cache_hit']['value']
miss = d['MAIN.cache_miss']['value']
print(f'Hit ratio: {hits/(hits+miss)*100:.1f}%')
"

Frequently Asked Questions

A cache whose store already holds the responses its current traffic asks for, so most requests are answered from the store instead of the origin. Warmth belongs to one store measured against demand, not to any single object, and it is not freshness. It cools on purge, eviction or restart.

Read the cache status in the response headers:

# Warm cache hit
curl -sI https://cdn.example.com/logo.png | grep -i x-cache
# X-Cache: HIT
# Age: 3542
# X-Cache-Hits: 847

# Cache miss (first request or expired)
curl -sI https://cdn.example.com/rare-page.html | grep -i x-cache
# X-Cache: MISS
# Age: 0

Watch the hit ratio in Varnish:

# Real-time hit rate
varnishstat -f MAIN.cache_hit -f MAIN.cache_miss

# Calculate hit ratio
varnishstat -1 -j | python3 -c "
import json, sys
d = json.load(sys.stdin)['counters']
hits = d['MAIN.cache_hit']['value']
miss = d['MAIN.cache_miss']['value']
print(f'Hit ratio: {hits/(hits+miss)*100:.1f}%')
"

Related CDN concepts include:

  • Cache Fill — A cache fill is the fetch-and-store after a cache miss: the cache pulls the resource …