Cache Key

Caching

The cache key is the identifier a cache derives from a request to decide which stored response, if any, may serve it. HTTP composes it from the request method and target URI, extended by the fields a response’s Vary header names; CDNs add or drop query parameters, headers and cookies on top.

Also known as Primary cache key, Cache ID.

11 min read Updated Aug 30, 2026

Full Explanation

The cache key is the identifier a cache derives from an incoming request. It decides which stored response, if any, may serve that request. The cache computes the key from the request alone, never from the response body. HTTP builds the key from at least the request method and the target URI. The key also grows to include whatever request fields the response’s Vary header names. There is no fixed formula for it: every CDN layers its own configurable rules on top. The cache key is also not a purge tag. Cache tags group objects for invalidation, and they play no part in lookup. Two requests that produce the same key are one object as far as the cache is concerned. Two requests that produce different keys are separate resources. That is exactly why cache busting works by changing the URL. Anything the origin varies on that the cache leaves out of the key is an “unkeyed” input. Unkeyed inputs are where both wasted duplication and cache poisoning come from. RFC 7234 split the idea into a “primary cache key” plus a Vary-derived “secondary key”. RFC 9111 collapsed both into one “cache key”. Akamai calls it a “cache ID”.

How it works

RFC 9111 section 2 defines it directly. The cache key “is the information a cache uses to choose a response and is composed from, at a minimum, the request method and target URI used to retrieve the stored response”. The same section adds a practical caveat: “many HTTP caches in common use today only cache GET responses and therefore only use the URI as the cache key”. The key component is the target URI, not just the path. So scheme, host, port, path and query are all in scope. RFC 9110 section 4.2.1 defines an http URI as a scheme, an authority, a path and an optional query. It states that “the hierarchical path component and optional query component identify the target resource within that origin server’s namespace”. That is why /page?id=1 and /page?id=2 are already different keys, before any CDN configuration exists.

When a stored response carries Vary, matching the URI is no longer enough. RFC 9111 section 4.1 says the cache “MUST NOT use that stored response without revalidation unless all the presented request header fields nominated by that Vary field value match those fields in the original request”. RFC 9110 section 12.5.5 puts it plainly: “Vary expands the cache key required to match a new request to the stored cache entry.” This is how content negotiation and caching coexist. With Vary: Accept-Encoding, a gzip request and a brotli request select different stored entries for one URL. Matching is not raw string comparison. Section 4.1 permits whitespace changes, folding repeated field lines and normalisation known to have identical semantics. An absent field can only match another absent field. A Vary value containing “*” always fails to match.

Above HTTP, each CDN and proxy decides which query parameters, headers and cookies join the key. No two defaults agree. Varnish’s built-in VCL calls hash_data() “on the host and URL of the request”. Fastly includes “req.url and req.http.host in the hash”. Cloudflare’s default is the full URL plus the Origin header and a set of proxy headers. CloudFront’s default is only the distribution domain name and the URL path. Take GET https://cdn.example.com/products?id=42&ref=homepage as an example. A typical CDN key is built from the scheme, the host cdn.example.com, the path /products and the query id=42&ref=homepage. Add Vary: Accept-Encoding, Accept-Language, and the negotiated values join it: gzip and en-US. Strip the tracking parameters (utm_*, fbclid, ref) and sort what is left. The key then collapses to scheme, host, /products, id=42. So every campaign variant of that page shares one stored copy.

Here is that key, built up step by step.

// Typical CDN cache key components
Request: GET https://cdn.example.com/products?id=42&ref=homepage
Cache key: https | cdn.example.com | /products | id=42&ref=homepage

// With Vary: Accept-Encoding, Accept-Language
Cache key: https | cdn.example.com | /products | id=42&ref=homepage | gzip | en-US

// After stripping tracking params (utm_*, ref, fbclid)
Cache key: https | cdn.example.com | /products | id=42
// Now ?ref=homepage and ?ref=twitter share the cached response

Why it matters for a CDN

Key design sets the cache hit ratio. It does so multiplicatively: each component you add splits one object into as many copies as that component has distinct values. CloudFront’s guidance is blunt: “if you include a value in the cache key that doesn’t affect the response that your origin returns, you might end up caching duplicate objects”. It also warns that “the User-Agent header can have thousands of unique variations, so it’s generally not a good candidate for including in the cache key”. The key is equally a correctness and security boundary. Anything the origin reads but the cache does not key on is an unkeyed input: “any web cache poisoning attack relies on manipulation of unkeyed inputs, such as headers”. A payload injected through one produces a response that, once cached, “will be served to all users whose requests have the matching cache key”. RFC 9111 section 7.1 makes the same point from the standards side: “storing malicious content in a cache can extend the reach of an attacker to affect multiple users”. It adds that this “is especially effective when shared caches are used to distribute malicious content to many clients”. There is one dial with two failure directions. Too much in the key wastes the cache and hammers the origin. Too little lets requests share responses that were never interchangeable.

What CDNs do

  • Cloudflare: the default key is the full URL (scheme, host, URI with query string), the Origin header sent by the client, and the x-http-method-override, x-forwarded-host and forwarded header families. All query parameters are included by default. Cache Rules can override Query String, Headers, Cookie, Host and User features. But those five settings are Enterprise-only. On Free, Pro and Business the cache-key controls are limited to Cache deception armor, Cache by device type, Ignore query string and Sort query string. In the default key $scheme is the origin scheme. So switching SSL mode from Off or Flexible to Full changes it and busts the cache. The Ignore Query String caching level “delivers the same resource to everyone independent of the query string”. But Cloudflare notes it “only disregards the query string for static file extensions”.
  • Fastly: the default hash is req.url plus req.http.host. It “takes no notice of HTTP method”, so unlike the HTTP definition the method is not keyed. TLS and plaintext requests “will find the same object”. The Host value is lowercased before hashing. But req.url “is not subject to normalization and will reflect the capitalization present on the request”. Fastly also mixes req.vcl.generation into the hash. That way a purge-all can drop the whole cache by making future requests miss. It recommends Vary over hash edits for “variations based on requested language, user location, or login state”. It also warns that variations are subject to resource limits. Beyond those limits, modifying the hash may be the better option. Ignoring the query string is an opt-in configuration change.
  • CloudFront: the default key is only the distribution domain name and the URL path: “other values from the viewer request are not included in the cache key, by default”. You opt query strings, headers and cookies in through a cache policy. The OPTIONS method is keyed for OPTIONS requests. So those responses cache separately from GET and HEAD. When caching of compressed objects is enabled for both Gzip and Brotli and the viewer advertises both, CloudFront “normalizes the header to Accept-Encoding: br,gzip and includes the normalized header in the cache key”. This normalisation is driven by the cache policy’s compression settings, not by the origin’s Vary header.
  • Akamai: “by default, cache keys are formed as URLs with full query strings”. The Cache Key Query Parameters behavior can keep the whole parameter set order-sensitive, alphabetize it so order stops mattering, ignore all parameters, or include or exclude a named list. Cache ID Modification adds headers, cookies, variables or the full URL. But “your changes to the cache key don’t apply to the two required cache key elements – the hostname, that is either the incoming Host header or the origin hostname, and the path to the object”. Only values the client actually sent can be keyed. A header injected or rewritten at the edge cannot define the key. If a client header is rewritten, “the original value sent from the client is used in the cache key”.

Watch out for

  • Unkeyed inputs are a poisoning primitive. If the origin reflects a request header the cache ignores, an attacker can elicit a poisoned response. That response is then served to everyone sharing the key. Where you set a custom key, Cloudflare recommends also enabling Normalize URLs to origin. This way “the URL in the cache key matches the URL sent to the origin, preventing cache poisoning”.
  • Header explosion. Each distinct value of a keyed header is another stored copy. CloudFront’s own example lists four spellings of English in Accept-Language: en-US,en, en,en-US, en-US, en, and en-US. These “all … indicate that the viewer’s language is English, but the variation can cause CloudFront to cache the same object multiple times”.
  • User-specific data in a shared key. RFC 9111 section 3.5 says a shared cache “MUST NOT use a cached response to a request with an Authorization header field … to satisfy any subsequent request”. An exception applies when a response directive such as public, must-revalidate or s-maxage allows it. RFC 9110 adds that there is “no need to send the Authorization field name in Vary because reuse of that response for a different user is prohibited by the field definition”. Session cookies are the same trap. CloudFront calls per-user cookie values “not good candidates for cache key inclusion”.
  • Cache tags are not the cache key. Surrogate keys “allow you to selectively purge related content”. They are a purge grouping. Lookup instead uses the hash, which Fastly calls “the primary object key”.
  • Changing the key throws the cache away. Akamai warns that “when you change the cache key, you invalidate cached content that uses existing cache key”. Akamai adds that edge servers then refetch “at a level that can create severe spikes in bandwidth”. Akamai’s reference docs repeat that “your origin server may experience traffic spikes before the new cache starts to serve out”. Cloudflare’s SSL-mode cache bust is the same failure with a different trigger.
  • Order and case fragment the key space. Under Akamai’s order-sensitive default, ?q=akamai&state=ma and ?state=ma&q=akamai “cache separately” until you alphabetize. Fastly lowercases only the host, so /Logo.PNG and /logo.png are distinct keys. Tracking parameters such as utm_source or fbclid sit inside the default key on Cloudflare and Akamai. This fragments one page into many.
  • Ignoring the wrong parameter serves the wrong content. Akamai’s own caution on the behavior: “be careful not to ignore any parameters that result in substantially different content, as it is not reflected in the cached object”.

Best practice

  • “Include only the minimum necessary values in the cache key.” The origin may want some values for analytics or telemetry even though they do not change the response. Those values belong in an origin request policy or forwarded header, not in the key.
  • Express representation differences with origin-side Vary rather than vendor hash edits: “an origin server SHOULD generate a Vary header field on a cacheable response when it wishes that response to be selectively reused for subsequent requests.” Keep the nominated value sets small or normalised. Also check your CDN’s limit on variations per object.
  • Normalise the query string: strip tracking parameters and sort the remainder so one page has one key. Akamai’s alphabetize option and Cloudflare’s Sort query string (available on every plan) both do this without custom code.
  • Keep Authorization, session cookies and User-Agent out of a shared key for content that is identical for everyone. Where content genuinely differs, encode the difference in the URL instead. CloudFront suggests language paths such as /en-US/… instead of keying Accept-Language.
  • Roll key changes out gradually and expect a cold cache. Akamai: “if you need to change the cache keys on an active configuration, limit the changes and send them out a few changes at a time over an expanded time frame.”

Interactive Animation

Loading animation...

Examples

This Varnish VCL example normalizes cache keys. It strips tracking parameters and sorts the rest.

sub vcl_recv {
    # Strip tracking query params from cache key
    set req.url = regsuball(req.url, "(\?|&)(utm_[a-z]+|fbclid|gclid|ref)=[^&]*", "");
    # Clean up leftover ? or &
    set req.url = regsub(req.url, "\?&", "?");
    set req.url = regsub(req.url, "\?$", "");
}
# These all resolve to the same cache key after normalization
curl https://cdn.example.com/page
curl https://cdn.example.com/page?utm_source=twitter
curl https://cdn.example.com/page?fbclid=abc123&utm_medium=social
# All three get the same cached response

Frequently Asked Questions

The cache key is the identifier a cache derives from a request to decide which stored response, if any, may serve it. HTTP composes it from the request method and target URI, extended by the fields a response’s Vary header names; CDNs add or drop query parameters, headers and cookies on top.

This Varnish VCL example normalizes cache keys. It strips tracking parameters and sorts the rest.

sub vcl_recv {
    # Strip tracking query params from cache key
    set req.url = regsuball(req.url, "(\?|&)(utm_[a-z]+|fbclid|gclid|ref)=[^&]*", "");
    # Clean up leftover ? or &
    set req.url = regsub(req.url, "\?&", "?");
    set req.url = regsub(req.url, "\?$", "");
}
# These all resolve to the same cache key after normalization
curl https://cdn.example.com/page
curl https://cdn.example.com/page?utm_source=twitter
curl https://cdn.example.com/page?fbclid=abc123&utm_medium=social
# All three get the same cached response

Yes. Cache Key is also known as Primary cache key, Cache ID. The cache key is the identifier a cache derives from a request to decide which stored response, if any, may serve it. HTTP composes it from the request method and target URI, extended by the fields a response’s Vary header names; CDNs add or drop query parameters, headers and cookies on top.