Failover
Failover is the automatic switch to a backup server, origin, or CDN provider once a health check decides the primary has failed. It is the switch itself, not the standby that waits, and not load balancing, which spreads traffic over every healthy server all the time.
Also known as fail-over.
Full Explanation
Failover is the automatic switch to a backup server, origin, or whole CDN provider. It happens when the primary one fails, so the service keeps answering. NIST defines it as the capability to switch over automatically to a redundant or standby system, typically without human intervention or warning. This happens upon the failure of the previously active one. Three things must already be in place for it to happen. First, a probe, such as an origin health check, must decide the active component is down. Second, a standby must already be provisioned. Third, a routing layer must be able to move traffic onto it.
Failover is not redundancy. Redundancy is the spare capacity. Failover is the act of using it. Nor is failover load balancing. A load balancer spreads traffic across every healthy backend all the time, and merely stops using a sick one. A failover standby, by contrast, sits idle until the primary fails. Fastly's documentation draws exactly this line, calling the first pattern load balancing and the second fallback. The number that matters is failover time. It has two components: how long detection takes, plus how long the switch takes. On a CDN, those two add up very differently at each layer. At the edge it is routing convergence; for a DNS-based provider switch it can take minutes. The hyphenated spelling fail-over is also standard usage, including in RFC 4786.
How it works
Detection almost never turns on a single failed probe. That is deliberate: one lost packet must not evict a healthy server. Real systems therefore require agreement. The shape of that agreement sets the detection floor.
- Consecutive failures. Amazon Route 53 has each of its health checkers probe your endpoint every 30 seconds by default. In the chargeable Fast mode it probes every 10 seconds. Route 53 treats the endpoint as failed after a threshold of consecutive observations: three by default, configurable from 1 to 10. With the defaults, detection alone can take about 90 seconds.
- A sliding window. Fastly keeps the last window results for a backend. It calls the backend healthy while at least threshold of those results passed. This smooths both recovery and failure.
- A quorum of observers. Cloudflare sends health monitor requests from three data centres in each region you select. It needs a majority of data centres in a region to pass, and then a majority of regions to be healthy. Akamai's Global Traffic Management allocates seven liveness-test agents per data centre. It considers servers up while a majority of those agents succeed. The majority rule is there precisely so a local network fault cannot falsely declare a data centre down.
One checker's opinion is also not the verdict. Route 53 aggregates across its checkers. It considers an endpoint healthy while more than 18 per cent of them report it healthy. This way, an endpoint isolated from a few checking locations is not condemned by them.
The switch then happens at one of three layers. Each layer has a different bound on how fast traffic can actually move.
- Edge failover rides on anycast. A node's route advertisement can be coupled to the availability of the service it carries. Then availability triggers the advertisement and non-availability triggers a withdrawal, and requests converge on the next node still advertising the prefix. RFC 4786 recommends that coupling, but it describes the case where one advertisement corresponds to a single service address. Where one prefix covers several services, withdrawing it because one service failed may not be appropriate. The switch is bounded by routing convergence, not by a DNS TTL or a health-check interval. That is why this is the fastest layer.
- Origin failover retries the request against a second origin. On CloudFront you define an origin group with a primary and a secondary origin. The edge fails over when the primary returns one of the status codes you selected from 400, 403, 404, 416, 429, 500, 502, 503 and 504. It also fails over when it cannot connect, but only if you selected 503, or when the origin times out, but only if you selected 504. Failover applies only to viewer requests using GET, HEAD or OPTIONS, never to POST or PUT. By default CloudFront spends as long as 30 seconds on the primary before switching: three connection attempts of 10 seconds. You can shorten that to 1 attempt of 1 second.
- Provider failover moves traffic between CDNs, usually through DNS. It is the slowest layer, because a resolver keeps the answer it was given for the length of the DNS TTL. RFC 1035 defines the TTL as the interval a record may be cached before the source is consulted again.
Two things make a DNS-based switch slower than the TTL suggests. Both are worth knowing before you promise a number. First, RFC 8767 amended the TTL definition: a record whose data cannot be authoritatively refreshed may be used as though it were unexpired. That is the exact circumstance of provider failover, where the old authoritative servers have gone quiet. Second, resolvers do not all honour the TTL you publish. Moura et al. measured roughly 15,000 vantage points. They found about 10 per cent of resolvers answering with the parent zone's 172,800-second TTL rather than the child's 300 seconds. About 15 per cent of answers for one second-level domain were capped at 21,599 seconds. Their measurements also show the reassuring half of the picture: manipulation of TTLs shorter than an hour is rare.
Negative caching is the other DNS-side delay. RFC 2308 has resolvers cache "this name does not exist" and "this type does not exist" for a TTL. That TTL is taken from the minimum of the SOA MINIMUM field and the SOA's own TTL. RFC 9520 went further in 2023. It made caching of outright resolution failures mandatory rather than optional. Those failures include SERVFAIL, REFUSED, timeouts, and unreachable servers. The cache time is at least one second and never more than five minutes, with backoff recommended in between.
Why it matters for a CDN
A CDN's promise is that a failure at one point of a distributed system does not reach the user as an error. Failover is the layer that delivers on it. A CDN is unusual in having to engineer failover at three layers at once, each with a different time constant: an edge node lost, an origin lost, a whole provider lost. Getting the layer wrong is a common design error. Putting a DNS switch where routing convergence was available costs minutes for nothing.
The cache is what makes the delay survivable. While the switch is in progress, the edge can answer from what it already stored. RFC 9111 lets a cache that receives a 5xx while revalidating act as though the server had not responded, and send a previously stored response instead. The stale-if-error Cache-Control extension of RFC 5861 says explicitly that on an error, a stale stored response may be used to satisfy the request, up to the staleness limit you set. RFC 5861 scopes "error" to the situations that would produce a 500, 502, 503 or 504. That is why the resilient pattern is failover plus stale serving, rather than failover alone.
Provider failover is also what gives a multi-CDN deployment its point. Without a tested switch, a second provider is an invoice rather than a mitigation. Route 53's own documentation describes the minimal version. Keep a backup site, such as a static site in an S3 website bucket. Fail over to it when the primary becomes unreachable.
What CDNs do
- CloudFront: origin groups with a primary and a secondary origin. It fails over on selected status codes, or on connect and timeout failures when 503 and 504 are selected. This applies to GET, HEAD and OPTIONS only, with 30 seconds on the primary by default. CloudFront origin failover.
- Route 53: health-checked records. The failover routing policy gives active-passive routing: the primary record while it is healthy, the secondary once it is not. Any other routing policy gives active-active routing, where every healthy record stays in the answer set. If both the primary and the secondary are unhealthy, Route 53 returns the primary. Every routing algorithm also has a mode of last resort. When all records are unhealthy, it reverts to considering them all healthy, so a broken health check cannot turn a partial outage into a total one. Route 53 failover.
- Fastly: health-checked backends. No HTTP requests are sent to a backend Fastly considers sick. From there you choose automatic load balancing that steers around the sick backend, or a fallback director that uses a standby only while the primary is down. Note the platform limit: Compute does not expose backend health, so failover logic cannot be written in a Compute program. Fastly redundancy and failover.
- Cloudflare: Load Balancing with pools, monitors that issue health checks at an interval, and steering that diverts traffic once a pool drops below its health threshold. A fallback pool is the pool of last resort. It receives traffic when everything else is unhealthy, and its own health is not considered. If that fallback pool is disabled too, a proxied hostname returns error 1016, origin DNS failure. Cloudflare Load Balancing health.
- Akamai: Global Traffic Management liveness tests from seven agents per data centre, with majority rule. Agents are deliberately spread across ISPs. A data centre counts as up if any of its servers is up, tunable with Minimum Live Percentage. If every server for a property fails with no backup CNAME configured, GTM treats all data centres as up. Akamai GTM liveness tests.
- nginx (open source): the origin-side equivalent. Most CDN customers run it behind the edge. A backup server in an upstream block receives requests when the primary servers are unavailable. max_fails failed attempts within fail_timeout mark a server unavailable for that same fail_timeout (defaults: 1 attempt, 10 seconds). proxy_next_upstream decides which failures move a request to the next server. Its default is error timeout only: 500, 502, 503, 504 and 429 responses are retried, and counted as failures, only if you list them. nginx proxy_next_upstream.
Watch out for
- A DNS switch is not instant, and it is not exactly one TTL. Route 53 recommends a TTL of 60 seconds or less on any record you associate with a health check. But serve-stale behaviour and parent-centric resolvers both let old answers outlive the TTL you published. Treat the TTL as a floor, not a deadline.
- Flapping. A check that toggles produces repeated advertisements and withdrawals. BGP calls these flaps, and routers frequently dampen them. Dampening suppresses the path for a period that grows with the oscillation. Because some implementations penalise by AS_PATH, one unstable node's flapping can leave stable nodes unreachable. RFC 4786 says nodes should be configured so rapid oscillation is avoided. One way is to impose a minimum delay after a withdrawal before the service may be re-advertised: a failback delay.
- Advertising a route you cannot serve. RFC 4786 warns about a specific case. If some local services in a node are down, and the node is cut off from the other nodes, continuing to advertise the covering prefix can black-hole requests. The counter-measure is withdrawing the covering prefix as soon as any one covered service fails. That is a genuine design option in the RFC, but not a general recommendation: it takes every service on the node offline for one failure. The RFC says this may make it unsuitable for many applications. Choose deliberately.
- A standby nobody exercised. A cold backup has an empty cache and unproven capacity, so it can collapse under the load failover just handed it. Route 53's documentation is blunt that it does not consider endpoint load or capacity when it fails over. Ensuring the survivors can carry the traffic is your job.
- Health checks that probe the wrong thing. A check hitting a static path returns 200 while the database is down, so nothing is ever marked unhealthy and failover never fires. Probe a path that exercises the real dependency instead. Route 53 offers string matching in the response body for exactly this reason, and Akamai and Fastly let you assert on response content too.
- Cold-start sickness. On Fastly, a backend that has a health check is considered sick when the service initialises. It only becomes healthy once enough checks pass, so the first requests after a deploy can get a 503 unless initial is set to match threshold. Failover machinery can itself cause the outage it exists to prevent.
Best practice
- Fail over at the lowest layer that can fix the failure: routing convergence for a lost edge node, an origin group for a lost origin, DNS only for a lost provider.
- Exercise the standby on a schedule and at realistic load. An untested backup is an assumption. Route 53 will not check whether it can take the traffic.
- Probe the real dependency, and require agreement before switching: several consecutive failures, or a majority of checkers. Then set the interval and threshold knowingly. Route 53's 30-second interval with a threshold of 3 is roughly 90 seconds of detection, before any DNS change even begins.
- Keep health-checked DNS records at a TTL of 60 seconds or less. Still budget for resolvers that hold the old answer longer.
- Add a failback delay so a recovered primary cannot start an oscillation in routing or DNS.
- Pair the switch with graceful degradation. A Cache-Control policy carrying stale-if-error lets the edge keep answering from cache while the switch completes. That is what turns a failover into something the user never notices.
Examples
# Nginx: upstream failover
upstream origin {
server primary.example.com:443;
server backup.example.com:443 backup;
}
# CloudFront: origin failover group
resource "aws_cloudfront_distribution" "cdn" {
origin_group {
failover_criteria {
status_codes = [500, 502, 503, 504]
}
member { origin_id = "primary" }
member { origin_id = "backup" }
}
}
# Resilient Cache-Control for graceful degradation
Cache-Control: public, max-age=300, stale-if-error=86400
Frequently Asked Questions
Failover is the automatic switch to a backup server, origin, or CDN provider once a health check decides the primary has failed. It is the switch itself, not the standby that waits, and not load balancing, which spreads traffic over every healthy server all the time.
# Nginx: upstream failover
upstream origin {
server primary.example.com:443;
server backup.example.com:443 backup;
}
# CloudFront: origin failover group
resource "aws_cloudfront_distribution" "cdn" {
origin_group {
failover_criteria {
status_codes = [500, 502, 503, 504]
}
member { origin_id = "primary" }
member { origin_id = "backup" }
}
}
# Resilient Cache-Control for graceful degradation
Cache-Control: public, max-age=300, stale-if-error=86400
Yes. Failover is also known as fail-over. Failover is the automatic switch to a backup server, origin, or CDN provider once a health check decides the primary has failed. It is the switch itself, not the standby that waits, and not load balancing, which spreads traffic over every healthy server all the time.
Related CDN concepts include:
- Origin Health Check — An origin health check is a probe a CDN repeats at a set interval against …
- Origin Shield — A cache tier a CDN places between its edge servers and its origin. Cache misses …
- Request Coalescing — A cache behaviour that serves many concurrent requests for one cache key from a single …
- stale-if-error — stale-if-error is a Cache-Control extension directive (RFC 5861) that lets a cache serve an already-stored …