Origin Health Check

Architecture

An origin health check is a probe a CDN repeats at a set interval against an origin, usually an HTTP GET or HEAD to a path such as /health, to decide whether that origin is fit to receive traffic. It yields one bit of routing state, healthy or unhealthy, not a measure of user-visible performance.

Also known as Health Probe, Health Monitor.

12 min read Updated Aug 30, 2026

Full Explanation

An origin health check is a probe. A CDN repeats it at a configured interval against an origin. It is most often an HTTP GET or HEAD to a path such as /health. Its whole job is to decide whether that origin is fit to receive traffic. Cloudflare's description is a fair general one: "a service that runs on Cloudflare's edge network to monitor whether an origin server is online" (Cloudflare Health Checks). The output is one bit of routing state per origin: healthy or unhealthy. Fastly names the states healthy and sick. Everything downstream follows from that bit. The CDN can steer away from the sick origin, hand off to a backup through failover, serve stale content, or return an error page.

It is not performance monitoring. A passing probe says nothing about what real users experience. RUM measures that instead. The latency a probe records is edge-to-origin, not browser-to-pixel. It is not a status-page uptime service either. That is because the CDN acts on the verdict automatically instead of paging a human. And it is not free. Every edge location probes on its own, so the endpoint you nominate receives far more requests than the interval alone suggests.

How it works

You configure a probe plus its pass criteria. The CDN issues it from its own edge servers, then grades the answers.

  • Protocol and method. HTTP or HTTPS is the common case. The method is normally GET or HEAD. Azure Front Door supports both. "For new Front Door profiles, the probe method is set as HEAD by default" (Azure Front Door health probes). Azure gives an explicit tip: HEAD lowers load and cost at the origin. A health check need not be HTTP at all. Cloudflare also offers TCP, ICMP Ping, UDP-ICMP and SMTP monitors. The last three are on Enterprise only (Cloudflare monitors).
  • Status criteria. Every provider defines its own accepted set. They genuinely differ, so there is no universal "2xx passes" rule. Cloudflare's expected codes accept individual values or a range written as 2xx. Azure Front Door is far stricter: "A 200 OK status code indicates the origin is healthy. Any other status code is considered a failure." RFC 9110 defines that 2xx class as the one indicating "the client's request was successfully received, understood, and accepted" (RFC 9110, section 15.3).
  • Response and timeout criteria. No response, or a response slower than the timeout, is a failure. Cloudflare's monitor asks exactly this: "Does the endpoint respond to the health monitor request at all? If so, does it respond quickly enough (as specified in the monitor's Timeout field)?" (Cloudflare health details).
  • Body criteria, optionally. The probe can also require expected text in the body. This way, a server that returns 200 while broken still fails. Cloudflare matches a case-insensitive substring. Cloudflare warns that the value must be relatively static and inside the first 10 KB of the page (Cloudflare manage monitors).
  • Interval. Azure Front Door's default probe frequency is 30 seconds. Cloudflare's minimum interval is a plan limit, not a preference: 60 seconds on Pro, 15 on Business, 10 on Enterprise. Fastly's check interval runs from a 1-second minimum to a 1-hour maximum (Fastly health check API).

A single failed probe usually does not flip the state. Each provider damps differently. This is why detection time is never just the interval:

  • Cloudflare damps geographically. For each region you select, it sends requests from three separate data centres. A majority of data centres makes the region healthy, and a majority of regions makes the endpoint healthy. A timed-out check triggers retries immediately, rather than waiting for the next interval. The API-only consecutive_up and consecutive_down parameters can additionally require several consecutive results before a state change.
  • Azure Front Door damps over a sliding window. It looks at the last n probe responses and calls the origin healthy if at least x were healthy. n is set by SampleSize and x by SuccessfulSamplesRequired.
  • Fastly damps the same way under different names. window is how many recent results are kept. threshold is how many of them must pass. So a setting of 3/5 means three of the last five (Fastly health checks).

Once an origin is unhealthy, the CDN acts with no human in the loop. It stops sending the origin live traffic ("Fastly will not send HTTP requests to backends that are sick"). It fails over to a backup origin, serves a stored stale copy, or returns a custom error page. One caveat on the definition: not every CDN runs a periodic origin probe at all. Amazon CloudFront has none. It reaches the same decisions reactively, from real requests instead.

Names differ by provider. The shape does not:

# Example CDN health check configuration
health_check:
  path: /health
  method: GET
  interval: 15s
  timeout: 5s
  healthy_threshold: 2       # 2 consecutive passes to mark healthy
  unhealthy_threshold: 3     # 3 consecutive failures to mark unhealthy
  expected_status: 200
  expected_body: "ok"
  
failover:
  backup_origin: backup.example.com
  serve_stale: true
  stale_ttl: 3600

Why it matters for a CDN

A CDN can only satisfy a cache miss by fetching from the origin. So a dead origin turns every miss into an error, while hits keep succeeding. Your cache hit ratio is therefore the blast radius. An origin shield narrows it but never closes it. The health check converts that slow bleed into a switch. It can take the origin out of rotation, hand over to a healthy peer or backup, or answer from what is already stored.

The serve-stale route has rules attached. They are worth knowing before you rely on it. RFC 9111 permits a stale response only when the cache "is disconnected or doing so is explicitly permitted by the client or origin server". It prohibits one outright when an explicit in-protocol directive forbids it, such as no-cache, must-revalidate, s-maxage or proxy-revalidate (RFC 9111, section 4.2.4). stale-if-error is the explicit permission that closes the gap. It says a stale response MAY be used when an error is encountered, "regardless of other freshness information". Here, an error means a 500, 502, 503 or 504 (RFC 5861, section 4). Fastly states the end of that chain plainly. If all origins are marked unhealthy, Fastly "will attempt to serve stale". If no stale object is available, the client gets a 503.

What CDNs do

The mechanism is common. The defaults, and even the model, are not. So nothing here is portable in detail.

  • Cloudflare has two forms. Standalone Health Checks monitor an IP or hostname and notify you. That functionality "is now offered as part of Cloudflare's origin server safeguard, Smart Shield". It is not available on the Free plan, and the limit is 10 checks on Pro, 50 on Business and 1,000 on Enterprise. Health monitors attached to load balancing pools are the other form. They are what drives steering. Region choice sets the probe count. All Regions sends three probes in each of 13 regions, 39 in total. All Data Centers on Enterprise probes from every data centre. The probe's user agent is a fixed Cloudflare-Traffic-Manager string. It carries the pool ID and cannot be overridden.
  • Azure Front Door has every edge location send a synthetic HTTP or HTTPS request to all configured origins. It uses the same TCP ports as real traffic, and tags each request with a "User-Agent" header of "Edge Health Probe". Only 200 OK counts as healthy. Latency is wall-clock time from just before the request to the last byte of the response. Each check uses a fresh TCP connection, so warm connections do not bias it. If every origin in a group fails its probes, Front Door considers them all unhealthy. It round-robins across them anyway, rather than serving nothing. A single-origin group may disable probes altogether to reduce load.
  • AWS CloudFront has no periodic origin health probe. Origin failover is opt-in and reactive. You create an origin group with a primary and a secondary. Then you choose which of 400, 403, 404, 416, 429, 500, 502, 503 or 504 trigger the switch. A connection failure only counts when 503 is among your failover codes. A timeout only counts when 504 is. By default, CloudFront tries the primary "for as long as 30 seconds (3 connection attempts of 10 seconds each)" before failing over. It uses a separate 30-second default response timeout, and it fails over only for GET, HEAD and OPTIONS. There is no sticky unhealthy state: "CloudFront routes all incoming requests to the primary origin, even when a previous request failed over to the secondary origin." Custom error pages can be configured per origin (CloudFront origin failover).
  • Fastly marks each backend healthy or sick from a repeated predefined request. It sends no HTTP requests to a sick backend. One designated cache server per site runs the check and shares the result within that site. Fastly calls this process health check amortization. The same server also performs the backend DNS lookup. A persistently failing lookup eventually marks the backend sick. A backend with a health check starts sick when the service initialises. Only the simplest case, a single backend with nothing in cache, produces a blanket Fastly-generated 503.

Watch out for

  • Probe volume is a multiple of the interval, not the interval. Azure's own rough estimate for the 30-second default is the number of edge locations times two requests per minute. This is lower where an edge location sees no real user traffic. Fastly works a plausible worst case: one check per second across 150 sites, 50 services and 5 A records is 37,500 requests per second. Applying the same share key across those services drops the same example to 750. Size the endpoint for the multiple, not the setting.
  • Your own security layer can fail the probe. Cloudflare explicitly tells you to make sure your firewall or web server does not block or rate limit its health monitors. A WAF rule, a bot filter, or rate limiting triggered by the probe's own volume will mark a perfectly healthy origin unhealthy. A fixed, non-overridable probe user agent is easy for such a rule to catch.
  • Check depth cuts both ways. A static file that always returns 200 keeps passing while the database is down. So failover never fires. A check that exercises every dependency turns one transient dependency blip into a whole-origin eviction, even though the origin could still serve cached and static content. Assert exactly what a failover would actually fix, and no more.
  • There is no standard for the response body. The often-cited proposal, draft-inadarei-api-health-check-06, defines application/health+json with a status field of "pass", "fail" or "warn". It requires a 2xx-3xx code for pass. But it is an expired Internet-Draft. It lapsed on 19 April 2022 and never became an RFC (Health Check Response Format for HTTP APIs). Azure will only accept 200, whatever the body says. Cloudflare does plain substring matching. Match the endpoint to what your provider actually checks.
  • Deploys and cold starts move health state. On Fastly, a health-checked backend is sick at initialisation. So if initial is lower than threshold, the first request in each site after a deployment is likely to get a 503. Changing any backend property also resets that state. Cloudflare reports a newly attached monitor as unknown until the first result arrives. Plan the first minute after a release, not just steady state.

Best practice

  • Serve the probe path from a cheap, dependency-free handler. It should return 200 as soon as the process can serve requests. Put database and third-party checks on a separate readiness path that the CDN does not poll.
  • Use HEAD where the provider supports it. Azure recommends it specifically to lower origin load and cost.
  • Make the endpoint's status codes agree with the provider's accepted set. Under Azure, anything other than 200 is a failure. So never answer a health probe with a 204 or a redirect.
  • If you match on the body, match a short static substring near the top of the response. Never let that string appear in an error page.
  • Exempt the probe from WAF rules, bot rules and rate limits. Exclude it from analytics and log-based billing by its user agent: Edge Health Probe on Azure, and the Cloudflare-Traffic-Manager string on Cloudflare.
  • Set the damping deliberately, rather than accepting defaults. Use two or three consecutive results, or a window and threshold such as 3 of 5, so one timeout never evicts an origin. Then compute the real detection time from interval, timeout, retries and threshold together.
  • Cut duplicate probe traffic at the source. Share one check across identical backends on Fastly. Select only the regions you need on Cloudflare. Disable probes for a single-origin group on Azure.
  • On CloudFront, do not wait for a probe that does not exist. Include 503 and 504 in the failover codes. Shorten the connection timeout and attempt count if 30 seconds of failing is too long for your traffic.
  • Give the unhealthy verdict somewhere to land. Pair it with a backup origin, a fallback pool, or a serve-stale policy. Confirm the stale path is not blocked by a no-cache or must-revalidate directive on the objects you were counting on.

Examples

A minimal health check endpoint and how to test it:

# Django health check view
from django.http import JsonResponse

def health_check(request):
    return JsonResponse({"status": "ok"}, status=200)
# Test your health endpoint the way a CDN would
curl -o /dev/null -s -w "status: %{http_code}, time: %{time_total}s\n" \
  https://origin.example.com/health
# status: 200, time: 0.043s

Frequently Asked Questions

An origin health check is a probe a CDN repeats at a set interval against an origin, usually an HTTP GET or HEAD to a path such as /health, to decide whether that origin is fit to receive traffic. It yields one bit of routing state, healthy or unhealthy, not a measure of user-visible performance.

A minimal health check endpoint and how to test it:

# Django health check view
from django.http import JsonResponse

def health_check(request):
    return JsonResponse({"status": "ok"}, status=200)
# Test your health endpoint the way a CDN would
curl -o /dev/null -s -w "status: %{http_code}, time: %{time_total}s\n" \
  https://origin.example.com/health
# status: 200, time: 0.043s

Yes. Origin Health Check is also known as Health Probe, Health Monitor. An origin health check is a probe a CDN repeats at a set interval against an origin, usually an HTTP GET or HEAD to a path such as /health, to decide whether that origin is fit to receive traffic. It yields one bit of routing state, healthy or unhealthy, not a measure of user-visible performance.

Related CDN concepts include:

  • Failover — Failover is the automatic switch to a backup server, origin, or CDN provider once a …