Cache Fill
A cache fill is the fetch-and-store after a cache miss: the cache pulls the resource from upstream (origin or shield tier) and stores it so a later request is a hit. The miss is not the fill, and a 304 revalidation is not a fill. Fills are the traffic that still costs origin bandwidth and latency.
Full Explanation
A cache fill is the operation of fetching a resource from upstream. Upstream means the origin server, or a more central tier such as an origin shield. A fill happens after a request misses a cache. The cache then stores the response. A later request for the same thing can then be served without going upstream again. Google Cloud CDN states it at its shortest: data transfer to a cache is called cache fill.
A fill is not the miss that triggered it. The miss is the failed lookup. The fill is the fetch-and-store that follows. Nor is a fill every trip upstream. A conditional revalidation answered with 304 refreshes the stored copy without moving a body. A fetch the cache declines to store is a fetch, but not a fill. Fills are the slice of traffic that still costs origin bandwidth, origin capacity and user latency. That is why CDNs meter and price fills apart from hits. “Fill rate” and “fill bandwidth” are the metrics over that operation. This entry is about the operation itself.
How it works
The cache looks the request up by cache key. RFC 9111 defines this as the information a cache uses to choose a response, composed at a minimum of the request method and target URI. A request can fail to match for several reasons: nothing is held under that key, the stored copy is stale, or a Vary-nominated request header does not match. In that case, the cache cannot satisfy the presented request, and RFC 9111 says the request is typically forwarded toward the origin. The miss comes first. The fill is what happens next.
Upstream returns a response. The cache then does two separable things with it. It serves the response to the client. It also stores the response, if the storage rules allow it: no no-store directive, and nothing marked private in a shared cache. Only the store makes it a fill. That separation is real, not pedantic. RFC 9211 reports whether the cache stored the response as its own boolean, distinct from the reason the request was forwarded. Nginx can also be told to wait for a set number of requests before it stores anything at all, rather than the default of storing on the first. Once stored, the next request for that key is a hit with no upstream fetch.
A revalidation is not a fill. When a stale entry is revalidated and the next hop answers 304, the cache updates the stored response with the new information and keeps the stored body. Only a 200, or a 206 carrying ranges, brings bytes in. This is the cheapest lever on the whole subject. An origin that emits ETag or Last-Modified lets a cache confirm content is still current without downloading it again. That turns what would have been a full fill into a header exchange. One client request also need not mean one fill. With byte range requests a cache can hold part of an object and fetch the remainder. A single client request can then trigger several cache fill requests. A response served partly from cache and partly from the backend is a partial hit rather than a clean hit or miss.
In a multi-tier deployment the fill repeats per tier as the response travels outward. On Fastly with shielding enabled, the shield POP requests the resource from the origin, caches the response and returns it to the regional POP where it is also cached before reaching the user. That means one origin fetch and two stores.
Here is that path, drawn out:
Request flow during a cache fill:
User -> Edge (MISS) -> Shield (MISS) -> Origin
|
Response 200
|
Origin ----fill----> Shield ----fill----> Edge -> User
(cached) (cached)
Next request for same content:
User -> Edge (HIT) -> User # No fill needed
A filled object does not stay filled. It stops being usable for two independent reasons. First, freshness runs out: the primary mechanism for determining freshness is an explicit expiration time from the origin, set via Expires or the max-age directive. Second, and quite separately, the entry gets evicted. Every cache is finite. Google notes that CDN caches are usually full, and so are constantly evicting content, generally whatever has not been accessed recently, regardless of its expiration time. A long TTL is therefore not a promise of residency. Unpopular objects get refilled on almost every request, no matter how you configure them. Filling is also reactive rather than pushed. Google Cloud CDN documents the limit case plainly: an object stored in one cache does not automatically replicate into the others, cache fill happens only in response to a client-initiated request, and you cannot preload the caches except by causing each one to answer a request.
Why it matters for a CDN
The goal of HTTP caching is significantly improving performance by reusing a prior response message. A fresh stored response can satisfy later requests without contacting the origin at all. A fill is the case where that reuse was unavailable. The user waits out a full upstream round trip. The origin also spends bandwidth, plus a request slot it did not have to spend on a hit.
Fill rate is the share of requests that end in an upstream fetch. On traffic that is strictly hit-or-miss, fill rate and cache hit ratio are complements. A 5% fill rate and a 95% hit ratio describe the same traffic, because a hit is by definition a request that was satisfied by the cache and not forwarded. Cutting fills and raising the hit ratio are therefore one lever, not two. Every request moved from fill to hit removes an upstream fetch.
Fill bandwidth maps to money. On some platforms it is a priced line item under exactly this name. Google Cloud CDN bills a cacheable cache miss as cache lookup plus cache data transfer out plus cache fill, plus any load-balancing or storage operation charges. It prices cache fill at US$0.01 to US$0.04 per GiB depending on source and destination. The rate is cheapest within North America or Europe, and dearest inter-region. Google also records that for typical workloads serving popular content, cache fill is often less than 10% of total data transfer GiB. That last figure is a useful sanity check on the shape of a healthy cache. Fills should be a small minority of bytes, not of configuration effort.
What CDNs do
- Cloudflare: reports a cache ratio measuring how often Cloudflare serves content from cache instead of contacting your origin, and prescribes a remedy per cache status. For Miss the documented fixes are Tiered Cache and a custom cache key so several URLs share one entry. Short-TTL churn shows up as Revalidated or Expired instead, where the advice is to raise Edge Cache TTL or return revalidation headers. Tiered Cache is opt-in from the dashboard or API on every plan. Once on, only the upper tier may ask the origin. The Generic Global, Regional and Custom topologies are Enterprise-only, and a request served through the hierarchy is flagged CacheTieredFill in the logs.
- Fastly: each POP fills independently by default. Any time a POP lacks a cached copy, it will request directly from your origin, even if another POP is already making the same request. Shielding is opt-in per backend and funnels origin fills through a single POP. Its caching best practice is to aim for a 90%+ cache hit ratio, with no-cache headers, cache-key fragmentation, short TTLs and unexpected traffic named as the suspects when it is lower. Segmented Caching fills only the byte ranges actually requested. Without it, the maximum object size depends on when the account was created: 20 MB for accounts created on or after 17 June 2020, and 2 GB for older accounts, or 5 GB with Streaming Miss.
- AWS CloudFront: on a cache miss it sends a request to the origin to retrieve the object, called an origin request. A mid-tier is already the default: an edge miss is passed to a regional edge cache before the origin. Cache policies decide what enters the cache key, while origin request policies decide what is forwarded. That is how you deliver extra data to the origin without fragmenting the key. Origin Shield is a further opt-in layer that consolidates concurrent requests for the same object, resulting in as few as one request reaching the origin. Requests made in the same Region as the origin bypass it, and it carries its own per-request charge.
- Google Cloud CDN: uses the term in its billing, charging cache fill separately from cache egress on a miss. Its caches are demand-filled only, with no preload path. A range-capable origin can see several fill requests raised for one client request.
Watch out for
- A measured fill rate is not the same thing as origin fetches. Fastly excludes shielding from the global hit ratio, so a local MISS answered by a shield HIT gets reported as a miss and a hit in the statistics, even though there is no call to the backend. Segmented Caching skews the same way, because only outer requests enter the calculation. Judge origin protection by origin request volume, not by hit ratio alone.
- The complement with hit ratio holds only on cacheable, hit-or-miss traffic. Cloudflare's Dynamic status is its default for many file types, including HTML. It marks content that was never eligible for caching. No-store responses never fill either, so the two figures need not add to 100%.
- Do not read every non-HIT as a fill. Of Nginx's seven $upstream_cache_status values, MISS and EXPIRED are the ones that pull a body. REVALIDATED is a 304, while STALE and UPDATING serve the stored copy. BYPASS does go upstream, but was configured not to be served from cache.
- A cache stampede happens when many clients miss the same cold key at once, and each starts its own fill. Request coalescing collapses them into one fetch, and RFC 9211 defines a collapsed parameter to report it. Coalescing is not always on by default: on Nginx, proxy_cache_lock is off unless you enable it.
- Over-fragmenting the cache key or the Vary field shrinks every partition and raises fill rate. Vary: * is the pathological case: a stored response whose Vary value contains a member "*" always fails to match, so every request forwards.
- A shield is not absolute. If a Fastly shield POP is unreachable for a request, that request goes straight from the edge node to your origin. This bypasses the shield entirely.
- Fills can be billed at more than one tier. Fastly bills inbound traffic to a shield as regular traffic, including the requests that populate remote POPs. Fastly also notes those charges are likely to be offset by the origin bandwidth and origin load you save.
- Changing origin IPs or DNS records under Cloudflare's Smart Tiered Cache can reassign upper tiers. The MISS rate then rises while the cache refills in the new ones. Expect a fill spike after that kind of change.
Best practice
- Send explicit freshness on everything cacheable, using max-age or Expires, instead of leaning on heuristics or a platform fallback TTL. A short or absent TTL is the most common source of constant fills.
- Prefer a long TTL plus purge on change over a short TTL. Fastly's own guidance is to set a long cache lifetime and send a purge when content changes, rather than expiring content quickly in the hope of staying current.
- Always emit ETag or Last-Modified, so that an expiry costs a 304 rather than a refetch of the body.
- Take the fill off the user's critical path with stale-while-revalidate. RFC 9111 lets a cache serve stale only where explicitly permitted, and it names the RFC 5861 extension directives as one such permission. RFC 5861 has the cache return the stale response immediately and revalidate in the background without blocking. On Nginx that needs proxy_cache_use_stale updating together with proxy_cache_background_update, both off by default.
- Add a tier, such as an origin shield, an upper tier, or CloudFront Origin Shield, so that many edge fills collapse into one origin fetch.
- Enable request coalescing so a stampede reaches the origin as a single fetch.
- Keep cache keys lean: drop query parameters that do not change the response, and restrict Vary to headers that genuinely select a different representation.
- Watch fill rate, hit ratio and absolute origin request volume after every deploy. A sudden jump usually means the cache key or the Vary set changed, not that traffic did.
Interactive Animation
Examples
Nginx exposes the fill status in a header:
# Add upstream status header to reveal fills
add_header X-Cache-Status $upstream_cache_status;
# Values: HIT, MISS, EXPIRED, STALE, BYPASS, REVALIDATED
# MISS and EXPIRED both trigger a cache fill
# Check fill status on a request
curl -sI https://cdn.example.com/video.mp4 | grep -i x-cache
# X-Cache-Status: MISS (this request triggered a fill)
curl -sI https://cdn.example.com/video.mp4 | grep -i x-cache
# X-Cache-Status: HIT (served from cache, no fill)
Frequently Asked Questions
A cache fill is the fetch-and-store after a cache miss: the cache pulls the resource from upstream (origin or shield tier) and stores it so a later request is a hit. The miss is not the fill, and a 304 revalidation is not a fill. Fills are the traffic that still costs origin bandwidth and latency.
Nginx exposes the fill status in a header:
# Add upstream status header to reveal fills
add_header X-Cache-Status $upstream_cache_status;
# Values: HIT, MISS, EXPIRED, STALE, BYPASS, REVALIDATED
# MISS and EXPIRED both trigger a cache fill
# Check fill status on a request
curl -sI https://cdn.example.com/video.mp4 | grep -i x-cache
# X-Cache-Status: MISS (this request triggered a fill)
curl -sI https://cdn.example.com/video.mp4 | grep -i x-cache
# X-Cache-Status: HIT (served from cache, no fill)