Multi-CDN
Multi-CDN is delivering one property's content through two or more independent CDN providers, with a steering layer choosing which provider answers each request. It is not one CDN with more capacity, and not failover alone: two contracts without steering are two single-CDN setups.
Also known as Multi-CDN strategy, Multi-CDN architecture.
Full Explanation
Multi-CDN is the delivery of one property's content through two or more independent CDN providers. A steering layer decides which provider answers each request. Fastly draws the line plainly: in single-CDN architectures, the entirety of your site's content gets delivered over one provider's network. With multi-CDN, you distribute your traffic across multiple CDNs. Each of those CDNs then delivers your content over its own network (Fastly multi-CDN research survey).
It is not one CDN with more capacity. It is also not failover by itself. Provider failover is one mode of multi-CDN, not its definition. Nor is it something a second contract gives you. Signing contracts with multiple CDN providers is only the first step: the steering decision, the health checks, the TTLs and the security policy all still have to be built. Two providers with no steering layer are two single-CDN setups pointed at the same origin. What the complexity buys is reach and survivability. CDNs present different strengths and weaknesses relative to each other. They serve certain types of clients, or different geographic regions, better than each other. Any one of them can fail outright. Cloudflare stopped delivering the majority of its core network traffic from 11:20 UTC on 18 November 2025. It was largely back by 14:30 and fully normal at 17:06. That was its worst outage since 2019. What it costs is split caches, duplicated configuration, a switch bounded by DNS caching, and a measurement problem that spans vendors.
How it works
Two conditions have to hold before any steering decision means anything.
- Every provider can serve the whole object set. The second CDN has to accommodate the existing publishing workflow. Lumen's guidance is to audit it feature by feature: token authentication, content invalidation, geo-filtering and caching or origin-fill rules (How does a multi-CDN strategy work for my business?). The incumbent may have implemented one of those in a proprietary way. All of them have industry-standard equivalents that can be synchronised between vendors. That audit, not the routing, is where most of the migration work sits.
- Something decides, per request, which provider answers. Hydrolix calls this the decision engine or steering logic. It is a system that sits between the end user and the CDN vendors. It decides which path a request takes (How does a multi-CDN strategy work?). Without it, there is no multi-CDN. There are only two deployments running side by side.
The decision can be taken in four different places. The choice fixes both how precise the decision can be and how quickly you can change your mind.
- DNS. This is the most common method, because it is transparent to the client and needs no application code. You put a traffic manager, an authoritative DNS service, in the middle. It answers with a CNAME pointing at the provider it picked. Routing at the DNS layer lets you decide before requests even reach your CDNs. The steerer never joins the data path. Azure states it directly for Traffic Manager: clients connect to the selected endpoint directly. Traffic Manager is not a proxy or a gateway, and it does not see the traffic passing between client and service (How Traffic Manager works). The price is that the answer is cached, so a change only reaches users as the DNS TTL expires.
- An HTTP redirect. A primary gateway evaluates the request headers and answers with a 302 or 307 to a specific CDN URL. That allows decisions on file size, user tier or real-time server load. But every redirect adds a round trip before the download starts. That can negate the benefit for small objects and latency-sensitive APIs.
- The manifest. For chunked video, rewriting the URLs inside the manifest steers delivery down to individual chunks. One CDN still has to serve the manifest itself.
- The player. A client-side balancer uses its own telemetry to pick a provider. It can switch mid-stream when one network degrades. The cost is player integration and iterative testing. Lumen notes that not all adaptive-delivery protocols support CDN switching at the client level. So this can never be ubiquitous across devices.
On top of the steering point sits the distribution mode.
- Active-active. Traffic is split across providers, evenly or by weight. Every CDN therefore keeps a populated cache, and you get comparable performance data from all of them. This is what makes performance-based routing possible, down to a specific autonomous system number or geography.
- Active-passive. All traffic goes to a primary. A configured secondary sits idle until the traffic manager detects an outage and flips the DNS switch. This is simple, and it is the mode that hides the worst trap in this entry.
- Geographic and performance. Each region is answered by the provider that measures best there, usually through Geo DNS records or a latency table.
Weights are shares, not schedules. That is easy to misread from a configuration file. Route 53 sends traffic to a record as its weight divided by the sum of the weights in the group. Weights of 1 and 255 give one 256th and 255 256ths. A weight of 0 stops traffic to that record entirely. Azure Traffic Manager chooses an available endpoint at random for each DNS query, with a probability taken from weights between 1 and 1000. Either way, the split is a statistical outcome over many queries. That is exactly why it can miss.
Why it matters for a CDN
A single provider is a single point of failure. The failure mode is not only the total outage. Lumen's account is that even the best CDN suffers from failing hard drives, server quirks, bottlenecks and last-mile congestion. So a single CDN may provide satisfactory or even very good service most of the time, but it will never always deliver exceptional performance in all regions or markets. Micro outages can drop one provider's availability below 80 per cent for a few minutes: one request in five failing. These never move a median, and they are invisible to anyone measuring only monthly averages. The total outage happens too. Cloudflare's own post-mortem records that a change to a database system's permissions caused an oversized Bot Management feature file to propagate network-wide. The software routing traffic across the network failed to load it. In the preceding six years, it had had no other outage that stopped the majority of core traffic (Cloudflare outage on November 18, 2025).
The performance argument is separate. It is also better evidenced than the marketing usually admits. Singh, Dunna and Gill measured three years of Windows and iOS update delivery. They found that CDNs present different strengths and weaknesses relative to each other, serving certain types of clients or different geographic regions better than others. They gave a concrete case. Microsoft obtained lower latencies for clients in developing regions by directing them to Akamai's rich network of edge caches. Meanwhile, clients in developing regions fetching Windows updates from Level 3 got poor latencies, arising from the absence of Level 3's footprint in those regions. The same study measured the gap that regional reach has to close. It found a median client-side latency of 20 ms in developed regions against medians as high as 200 ms in developing ones (Characterizing the Deployment and Performance of Multi-CDNs, IMC '18).
Commercially, being able to move volume is the lever. Fastly's 2019 survey of more than 300 decision makers put cost control first among the benefits its multi-CDN users cited, ahead of resiliency, performance and scale. The mechanism is unglamorous. Organisations hold pre-negotiated commitments with several providers at different price points. Once the commitment with the expensive one is met, traffic shifts to a cheaper one. Hydrolix describes the same pattern as commit sparing, with the engine overflowing excess traffic to a lower-cost provider as a quota limit approaches. Capacity is the other lever. Spreading a very large live event across providers means no single vendor has to absorb the whole peak.
What CDNs do
No CDN sells multi-CDN as a feature of its own edge, because it is a property of the publisher's architecture rather than of any one network. The vendor behaviour that matters is in the steering layer, and it does differ, in what you can express as a policy and in how the layer guesses where the user is.
- Amazon Route 53 is the mechanism in this entry's code example. A weighted routing policy links several records of the same name and type and splits traffic by relative weight, with 0 draining a record. A failover routing policy gives active-passive. Geolocation, geoproximity and IP-based policies route on where the user is. Its latency routing policy is documented for the case where you have resources in multiple AWS Regions. That is not the same problem as choosing between third-party CDN hostnames. Choosing a routing policy.
- Azure Traffic Manager is DNS-based, with one routing method per profile out of Priority, Weighted, Performance, Geographic, Multivalue and Subnet. Priority is the active-passive pattern, and Weighted is the split. Every profile includes health monitoring and automatic failover of endpoints. Endpoints may be external and non-Azure, which is what lets it point at other providers. TTL in its answers is configurable from 0 to 2,147,483,647 seconds, which Microsoft describes as the maximum range compliant with RFC 1035. Microsoft also documents that Traffic Manager itself is resilient to failure, including the loss of an entire Azure region. That is a statement about the service, not about a region-failover feature it sells you. It is worth reading literally, because the steering layer has a blast radius of its own. Traffic Manager routing methods, Traffic Manager overview.
- Akamai Global Traffic Management is a DNS-based traffic manager. Its liveness tests decide whether each traffic target is up, with weighted random and performance-based load balancing as the two primary modes. Its load imbalance factor bounds how far demand to a target may exceed its configured share. With a 25 per cent allocation and a factor of 50 per cent, demand may grow to 37.5 per cent before load is shifted away. The default is 10 per cent, and Akamai warns that reducing it below 10 per cent is likely to produce oscillation. GTM load balancing.
- Specialist decision engines. Hydrolix's position is that most multi-CDN deployments need a purpose-built traffic manager. It names NS1, Cedexis, or a do-it-yourself system, because standard DNS providers generally lack the real-time data integration needed to route on performance. Treat that as a vendor's view of the market rather than a measurement. But the requirement it describes is real: a policy that reacts to performance needs a data feed, not just a record set.
Watch out for
- A cold standby converts a provider outage into an origin outage. In active-passive, if the secondary CDN is cold, its caches are empty. Shifting full production traffic to it makes the origin face a massive spike of cache misses. That can crash the backend just as you try to recover from the CDN outage. A cold cache plus a full cache fill at once is the failure the design was supposed to prevent.
- A weighted DNS split is statistical, not exact. Microsoft documents that client and recursive-resolver caching can significantly skew a weighted distribution when the number of clients or resolvers is small: development environments, app-to-app traffic, a user base behind one corporate proxy. Microsoft says explicitly that these DNS caching effects are common to all DNS-based traffic routing systems, not just Traffic Manager. Verify the realised split in logs. Do not assume the configured one.
- The steerer usually sees the resolver, not the user. Azure's Geographic method reads the source IP of the DNS query, typically the local resolver. Its Performance method looks up the recursive service's address in an internet latency table, rather than the client's. Route 53 narrows the gap with the edns-client-subnet extension of EDNS0, but only when the resolver supports it. Otherwise it uses the source IP address of the DNS resolver to guess where the user is. Lumen's version of the warning is that a resolver serving a wide geographical area causes inaccurate identification of those users' locations.
- The switch is not instant, and not exactly one TTL. If you change a routing decision because a CDN went down, users with cached records keep hitting the failed provider until the TTL expires. Hydrolix adds that even at a 60-second TTL, downstream ISPs may ignore the setting and cache longer. So multi-CDN mitigates outages, but it rarely delivers zero-downtime failover without some users seeing errors during the transition. Azure puts the same trade-off the other way round: longer TTLs mean directing traffic away from a failed endpoint takes longer.
- Splitting traffic splits the cache. When you split traffic between two CDNs, you split cache efficiency. If one user fetches an image through CDN 1 and another through CDN 2, the origin serves it twice. Hydrolix scopes the damage rather than asserting it universally. In an active-active setup with low traffic volume, this fragmentation can increase origin load and reduce overall performance. High volume dilutes it. A small site splitting three ways can visibly lose cache hit ratio.
- Load balance fights performance. Akamai's example is exact: two data centres, New York and Singapore, a 50/50 split, and users on both sides of the world. During New York business hours, almost all demand is American. Honouring the split sends half of it to Singapore, and many of those users get poor performance. That tension is why a fixed percentage split belongs underneath a regional policy, not above it.
- Latency steering does not know about load. Azure notes that Performance routing does not monitor load on an endpoint. It only removes endpoints that its health checks mark unavailable. Steering on measured latency alone can therefore keep pushing traffic at a provider that is saturating but not yet failing.
- Security policy drifts apart. WAF rules and rate-limiting policies must be synchronised across all providers: if the primary blocks a given SQL injection and the secondary does not, attackers eventually find the gap. The same applies to token authentication, geo-filtering and invalidation semantics, which are the four places Lumen found vendors implementing proprietary variants.
- Observability fragments before anything else does. Three vendors means three dashboards, three log formats, and possibly three definitions of latency. Akamai DataStream logs, CloudFront real-time logs and Fastly logs all differ, so comparing providers means normalising them first. Until that exists, you cannot tell whether an error rate rose globally or on one provider. That makes every steering decision a guess.
- Synthetic checks flatter the wrong provider. A CDN can look healthy from a data centre in Virginia and perform poorly for a user on a mobile network in Jakarta. Latency is a reasonable proxy for many performance elements, as Lumen puts it. But it may not tell the whole story. What shows what end users actually experience is real user monitoring and player-side telemetry.
Best practice
- Put the regional policy first and the percentage split underneath it. Lumen's recommendation is a hierarchy: regional splits for performance first, then a split by percentage of users for resilience and cache warming. Avoid steering by tagging specific content to specific CDNs. It slows down traffic changes, and it can induce large cache-fill volume when a CDN has to take over content it never cached.
- Give every provider enough live traffic to stay warm. Lumen's figure is that a CDN needs a minimum of more than 15 per cent of customer traffic in each market or region. That keeps caches warm for a traffic spike or a failover, and it is what makes the eventual ramp smooth. An idle standby is an invoice, not a mitigation.
- Move traffic in bounded, reversible steps. Lumen's rule is closed-loop control that avoids shifts of more than 50 per cent in a given region, because cache-fill traffic can overload origin infrastructure. Lumen adds that hands-on control and data analysis over several months should precede any attempt at full automation. Route 53 supports the same discipline directly: change the balance slowly by changing weights, and set a weight to 0 to stop sending traffic to a record.
- Write the policy down before buying the tooling, against the two documented criteria: in which regions each CDN performs best, and what commitment level makes sense with each vendor given that commitment drives unit cost.
- Decide on real-user data, not only synthetic probes, and keep both. Client-side telemetry is a direct reflection of end-user experience: viewing bit rates, rebuffering percentages, fatal errors. It is what turns a DNS balancer from a static policy into a control loop.
- Normalise logs from every provider into one place before you trust any steering decision. Then "which provider is better" is a query rather than an argument between dashboards.
- Deploy security rules to all edge networks simultaneously, rather than provider by provider. This is Hydrolix's explicit remedy for coverage gaps.
- Exercise the switch on a schedule. Hydrolix's minimum for an availability-only deployment is active-passive with regular game days to test the backup path. Test the health check, the TTL and a working cache together, because it is their combination that fails.
- Keep the steering layer out of the failure domain it protects against. If authoritative DNS and the primary CDN are the same vendor, a provider-wide event can take the steerer with it. Azure documents Traffic Manager's own resilience as extending to the loss of an entire Azure region, precisely because that question gets asked of a steering service.
Interactive Animation
Examples
# DNS-based multi-CDN with weighted routing (Route53)
resource "aws_route53_record" "cdn" {
zone_id = aws_route53_zone.main.id
name = "cdn.example.com"
type = "CNAME"
weighted_routing_policy { weight = 70 }
set_identifier = "cloudflare"
records = ["cdn.example.com.cdn.cloudflare.net"]
}
resource "aws_route53_record" "cdn_backup" {
zone_id = aws_route53_zone.main.id
name = "cdn.example.com"
type = "CNAME"
weighted_routing_policy { weight = 30 }
set_identifier = "fastly"
records = ["cdn.example.com.global.fastly.net"]
}
Frequently Asked Questions
Multi-CDN is delivering one property's content through two or more independent CDN providers, with a steering layer choosing which provider answers each request. It is not one CDN with more capacity, and not failover alone: two contracts without steering are two single-CDN setups.
# DNS-based multi-CDN with weighted routing (Route53)
resource "aws_route53_record" "cdn" {
zone_id = aws_route53_zone.main.id
name = "cdn.example.com"
type = "CNAME"
weighted_routing_policy { weight = 70 }
set_identifier = "cloudflare"
records = ["cdn.example.com.cdn.cloudflare.net"]
}
resource "aws_route53_record" "cdn_backup" {
zone_id = aws_route53_zone.main.id
name = "cdn.example.com"
type = "CNAME"
weighted_routing_policy { weight = 30 }
set_identifier = "fastly"
records = ["cdn.example.com.global.fastly.net"]
}
Yes. Multi-CDN is also known as Multi-CDN strategy, Multi-CDN architecture. Multi-CDN is delivering one property's content through two or more independent CDN providers, with a steering layer choosing which provider answers each request. It is not one CDN with more capacity, and not failover alone: two contracts without steering are two single-CDN setups.
Related CDN concepts include:
- Origin Health Check — An origin health check is a probe a CDN repeats at a set interval against …
- Anycast — Anycast announces one IP address from many locations at once; the routing system, usually BGP, …
- Failover — Failover is the automatic switch to a backup server, origin, or CDN provider once a …