LRU Cache

Caching

LRU (least recently used) is a cache eviction policy: when the cache is full, it discards the object that has gone longest without a request. It decides which object leaves when space runs out, not whether a copy is still fresh, so an LRU eviction can drop an object whose TTL has hours left.

Also known as Least Recently Used, LRU eviction.

9 min read Updated Aug 30, 2026

Full Explanation

LRU (least recently used) is a cache eviction policy. When a cache has no room for a new object, it discards the object that has gone longest without being requested. It answers one question only: which stored copy leaves when space runs out. It is not a freshness rule and not an invalidation mechanism. A TTL decides when a stored copy is too old to serve. A purge removes a copy on command. LRU acts purely on capacity. The two limits are independent, so an LRU eviction can drop a copy that is still perfectly fresh. Cloudflare states that even if a TTL "specifies that content should be cached for a long time, we may still need to evict it earlier if it's less frequently requested than other resources". Akamai documents the same for its edge servers. This is the normal condition of a busy edge, not a fault: every cache in a CDN path evicts. What you actually control is the capacity LRU works against. That capacity also shapes the cache hit ratio, since each eviction of a still-wanted object turns a later request into a capacity cache miss.

How it works

  • Every cached object carries a recency marker. That marker is an age bit, a sequence number, or the timestamp of its last access.
  • Each read re-marks that object as the most recently used one. So anything re-read repeatedly keeps returning to the front of the order.
  • An insert that does not fit evicts from the other end. The object with the oldest last access goes first, and as many as needed go after it.
  • Nothing in the decision looks at how often an object was read. The decision also ignores the object's size and how long it is still allowed to be served.

Exact LRU is not free: the bookkeeping has to be updated on every hit. In hardware caches the cost rises with associativity, and "practical hardware usually employs an approximation to achieve similar performance at a lower hardware cost". Operating systems use the Clock approximation for the same reason. Software caches approximate too. Redis says its algorithm "uses an approximation of the least recently used keys rather than calculating them exactly. It samples a small number of keys at random and then evicts the ones with the longest time since last access". Redis has tracked a pool of good eviction candidates since Redis 3.0, because a true LRU "costs more memory".

Production caches normally pair capacity eviction with a second, distinct rule: an idle timer. nginx removes cached data not accessed within its inactive window "regardless of their freshness". Akamai's Cloud Wrapper applies a platform-level eviction once an asset has not been accessed for its idle TTL. Either rule can remove an object whose TTL has not expired.

LRU's closest relative is LFU (least frequently used). LFU counts accesses instead of ordering them: "how many times a block was accessed is stored instead of how recently". Recency is cheap and adapts quickly to change. Frequency remembers popularity that recency has already forgotten.

Why it matters for a CDN

An edge node caches a small slice of the object set it fronts. So on a busy PoP the store is effectively always full. Capacity, not expiry, is the usual reason a copy is missing. Request popularity is heavily skewed. That skew is what makes edge caching work at all. Breslau and colleagues measured five large proxy traces in 1998. They found requests follow a Zipf-like distribution "very well, but with alpha ranging from 0.57 to 0.67, instead of 1". LRU exploits that skew without measuring it. The hot set is re-read constantly, so it keeps re-marking itself as recent, while the long tail sinks to the eviction end. It also follows a hot set that moves. In the same traces, about two thirds of a day's 600 most popular URLs were still in the next day's list. The weekend hot sets changed by more than half against the working days.

Two consequences are specific to CDN work. Eviction is a per-node decision. So the same object can be resident in one PoP and already evicted in another, and hit ratio legitimately varies by location. Every eviction of a still-wanted object is paid for upstream too: the next request becomes a miss and a fetch from the origin. Cloudflare puts the cost plainly for large libraries that are not requested very frequently: in a traditional caching setup, "these assets might be evicted as they become less popular and, when requested again, fetched from the origin, resulting in egress fees".

What CDNs do

  • Cloudflare names LRU as its eviction algorithm: "our eviction strategy prioritizes content based on its popularity, employing an algorithm known as 'least recently used' or LRU". It may evict before TTL to make room for more frequently requested content (Cache Reserve goes GA, October 2023). Its Cache Reserve product is the documented backstop. Content admitted there is stored "for a much longer period of time — 30 days by default — without being subjected to LRU eviction".
  • Akamai documents that "Objects are removed from cache on a least recently used (LRU) basis. If an edge server's cache is full, it can remove objects that are infrequently accessed from the cache even if those objects have not yet reached their Time-to-Live (TTL)" (Learn about Akamai's caching). Its Cloud Wrapper mid-tier states both rules explicitly: LRU applies once cache space is full, and a platform eviction also applies for objects idle beyond their idle TTL.
  • Varnish keeps objects on an LRU list and exposes it in varnish-counters. n_lru_nuked counts objects "forcefully evicted from storage to make room for a new object". n_lru_moved counts "move operations done on the LRU list". Expiry has its own counter, n_expired, for "objects that expired from cache because of old age". So the two ways out of the cache are separately measurable.
  • nginx gives the job to a "cache manager" process that watches the max_size and min_free parameters of proxy_cache_path: "When the size is exceeded or there is not enough free space, it removes the least recently used data", in bounded iterations. Both parameters are optional. If neither is set, there is no capacity trigger at all, and only the inactive timer removes anything. The module documents no alternative eviction policy.

Watch out for

  • A miss on an object with a long TTL is not a bug. Both Cloudflare and Akamai warn that a full cache evicts unpopular content ahead of its TTL. Cloudflare notes this "can sometimes perplex users who wonder why a cache miss occurs unexpectedly". Raising the TTL does not stop it, because LRU never reads the TTL.
  • Idle timers, not just capacity. nginx's inactive default is 10 minutes. An object with a one-year TTL disappears after ten quiet minutes, even with the disk nearly empty. Check that value before blaming cache size.
  • Scans and one-hit objects. A crawler or a one-pass stream inserts objects that will never be read again. Each insert is marked most recently used. "Many cache algorithms (particularly LRU) allow streaming data to fill the cache, pushing out information which will soon be used again (cache pollution)." This is not a corner case for CDNs, where cache workloads "tend to show high one-hit-wonder ratios".
  • Frequency is invisible. LRU stores order, not counts. So an object read a thousand times an hour ago ranks below an object read once a minute ago. Under the independent-reference model that Zipf-like traffic approximates, the best online policy is in fact frequency-based. Breslau and colleagues found Perfect-LFU performed best on byte hit ratio "in most cases" in trace-driven simulation. They also noted that "LRU is the most widely-used Web cache replacement algorithm". LRU is a cost and adaptivity choice, not an optimum.
  • Thrashing. If the working set is larger than the cache, objects are evicted before their next request. Both hit ratio and origin load worsen as a result. Eviction rate, not miss rate, is the leading indicator.
  • Eviction has a ceiling. Varnish attempts at most nuke_limit objects (default 50, flagged experimental) to make space for one object body. It counts each time it hits that wall in n_lru_limited: "number of times more storage space were needed, but limit was reached in a nuke_limit". A large body arriving into a store full of small objects can hit the ceiling while eviction is working normally.

Best practice

  • Size storage so the hot working set fits. Treat a climbing eviction counter, not a bare miss ratio, as the signal to add capacity or cache less.
  • Graph evictions and expiries separately, since they have separate causes and separate fixes. In Varnish that is n_lru_nuked against n_expired, plus n_lru_limited for the large-object case.
  • Never lengthen a TTL to fight eviction. Capacity and freshness are independent limits. Only capacity, tiering or a reserve product changes an eviction outcome.
  • Put a mid-tier between edges and origin for long-tail libraries, so an edge eviction costs a shield hit rather than an origin fetch. This is what origin shield, Akamai's Cloud Wrapper and Cloudflare's Cache Reserve are for.
  • For scan-heavy or one-hit-heavy workloads, the fix is admission control as much as eviction order. TinyLFU decides "whether it is worth admitting the new item into the cache at the expense of the eviction candidate". Its W-TinyLFU variant "is demonstrated to obtain equal or better hit-ratios than other state of the art replacement policies on these traces". SLRU protects lines that "have been accessed at least twice". SIEVE and S3-FIFO are recent designs aimed squarely at web-cache one-hit-wonders. These are cache-engine algorithms you choose by picking an engine, not CDN product settings.
  • Validate any sizing or policy decision against a real request trace. A uniform synthetic load hides the skew that makes LRU a reasonable default in the first place.

Examples

Varnish cache size and LRU behavior:

# Start Varnish with 2GB cache (LRU is default eviction)
varnishd -a :80 -s malloc,2G

# Monitor evictions (nuke = LRU eviction)
varnishstat -f MAIN.n_lru_nuked

# If nuked count climbs fast, your cache is too small
watch -n 1 'varnishstat -1 | grep lru_nuked'

Nginx proxy cache with size limit:

# 10GB cache zone, LRU eviction when full
proxy_cache_path /var/cache/nginx
    levels=1:2
    keys_zone=cdn_cache:10m
    max_size=10g
    inactive=24h
    use_temp_path=off;

# inactive=24h removes items not accessed in 24h
# max_size=10g triggers LRU eviction at 10GB

Frequently Asked Questions

LRU (least recently used) is a cache eviction policy: when the cache is full, it discards the object that has gone longest without a request. It decides which object leaves when space runs out, not whether a copy is still fresh, so an LRU eviction can drop an object whose TTL has hours left.

Varnish cache size and LRU behavior:

# Start Varnish with 2GB cache (LRU is default eviction)
varnishd -a :80 -s malloc,2G

# Monitor evictions (nuke = LRU eviction)
varnishstat -f MAIN.n_lru_nuked

# If nuked count climbs fast, your cache is too small
watch -n 1 'varnishstat -1 | grep lru_nuked'

Nginx proxy cache with size limit:

# 10GB cache zone, LRU eviction when full
proxy_cache_path /var/cache/nginx
    levels=1:2
    keys_zone=cdn_cache:10m
    max_size=10g
    inactive=24h
    use_temp_path=off;

# inactive=24h removes items not accessed in 24h
# max_size=10g triggers LRU eviction at 10GB

Yes. LRU Cache is also known as Least Recently Used, LRU eviction. LRU (least recently used) is a cache eviction policy: when the cache is full, it discards the object that has gone longest without a request. It decides which object leaves when space runs out, not whether a copy is still fresh, so an LRU eviction can drop an object whose TTL has hours left.

Related CDN concepts include:

  • Cache Key — The cache key is the identifier a cache derives from a request to decide which …
  • Cache Miss Types — A cache miss is a request an edge node cannot answer from its own cache. …