DNS (Domain Name System)
The distributed, hierarchical naming system that resolves names like example.com to addresses. A query-response lookup service, not a routing protocol. For a CDN it is also the steering plane: the authoritative nameserver picks which edge address to return, and the TTL caps how fast that changes.
Full Explanation
The Domain Name System (DNS) is the distributed, hierarchical naming system. It turns a name such as example.com into the addresses a client can actually connect to. There is no single tidy definition of it. RFC 9499, section 1 describes the DNS as some combination of “a commonly used naming scheme for objects on the Internet; a distributed database representing the names and certain properties of these objects; an architecture providing distributed maintenance, resilience, and loose coherency for this database; and a simple query-response protocol”.
Two things it is not. It is not a routing or a transport protocol. An answer only tells the client which address to open a connection to. The packets then travel by ordinary IP routing. That separation is deliberate. The first design goal of the DNS is “a consistent name space which will be used for referring to resources”. The same section also says that “names should not be required to contain network identifiers, addresses, routes, or similar information as part of the name” (RFC 1034, section 2.2). It is also not a content path. Nothing a CDN delivers travels over DNS. For CDN work DNS matters twice. It is the first round trip of nearly every request made to a hostname. It is also the steering plane, because “Many Authoritative Nameservers today return different responses based on the perceived topological location of the user” (RFC 7871, section 1). That is how one request ends up at one edge server and the next at another.
How it works
- The name space is a tree, and authority is cut out of it. “The domain name space is a tree structure”, and “the labels that compose a domain name are printed or read left to right, from the most specific (lowest, farthest from the root) to the least specific (highest, closest to the root)”. So in www.example.com, www is the most specific label and com is nearest the root (RFC 1034, section 3.1). Cuts in that tree carve it into zones. Each zone “is said to be authoritative for all names in the connected region”. Every zone also has a highest node, and the “name of this node is often used to identify the zone” (RFC 1034, section 4.2). That node is the zone apex. The root servers know the .com servers. The .com servers know the example.com servers. Only the latter hold that zone's records.
- Resolution is a referral walk. A browser or operating system uses a stub resolver, “A resolver that cannot perform all resolution itself”. “Stub resolvers generally depend on a recursive resolver to undertake the actual resolution function”. The recursive resolver in turn “is expected to cache the answers it receives” (RFC 9499, section 6). That recursive resolver is normally your ISP's, or a public one configured in its place. For example, “Google Public DNS offers IPv4 addresses (8.8.8.8 and 8.8.4.4)”, and using it means “switching from your internet service provider's (ISP) domain name servers to Google's” (Google Public DNS, Get Started). The resolver starts at the root and works down. On each reply, “if the response contains a better delegation to other servers, cache the delegation information, and go to step 2” (RFC 1034, section 5.3.3).
- Answers are resource records. A is “a host address”, meaning IPv4 (see A record). NS is “an authoritative name server”. CNAME is “the canonical name for an alias”. SOA “marks the start of a zone of authority” (RFC 1035, section 3.2.2). AAAA “stores a single IPv6 address” (RFC 3596, section 2.1).
- Transport is UDP port 53 first, TCP when the answer will not fit. “Messages sent using UDP user server port 53 (decimal)” (RFC 1035, section 4.2.1). “Traditional DNS messages are limited to 512 octets in size when sent over UDP” (RFC 6891, section 4.3). A server whose reply would exceed the limit truncates it and sets the TC flag. The client “takes the TC flag as an indication that it should retry over TCP instead” (RFC 7766, section 4). EDNS(0) lets the requestor advertise a larger UDP payload, but the familiar 4096 is only advice, not a rule. “A good compromise may be the use of an EDNS maximum payload size of 4096 octets as a starting point” (RFC 6891, section 6.2.5). Treating DNS as UDP-only is wrong. Since RFC 7766, “All general-purpose DNS implementations MUST support both UDP and TCP transport” (RFC 7766, section 5).
- Every record carries a TTL. It is “the time interval that the resource record may be cached before the source of the information should again be consulted” (RFC 1035, section 3.2.1). Until it expires a resolver reuses what it holds. “If the data is in the cache, it is assumed to be good enough for normal use” (RFC 1034, section 5.3.3). The number belongs to the zone that owns the data by design. “Where there tradeoffs between the cost of acquiring data, the speed of updates, and the accuracy of caches, the source of the data should control the tradeoff” (RFC 1034, section 2.2).
- Failures are cached too. A name error (NXDOMAIN) answer “should be cached such that it can be retrieved and returned in response to another query for the same <QNAME, QCLASS>”. A no-data error (NODATA) is cached the same way, per <QNAME, QTYPE, QCLASS>. The clock on it is not the missing record's own TTL. That clock “is taken from the minimum of the SOA.MINIMUM field and SOA's TTL” (RFC 2308, section 5). See negative caching.
- EDNS Client Subnet (ECS) puts the client back into the query. It is an EDNS0 option “to allow Recursive Resolvers, if they are willing, to forward details about the origin network from which a query is coming when talking to other nameservers” (RFC 7871, section 5). An authoritative server can then tailor its answer to the client's network instead of the resolver's.
Why it matters for a CDN
- It is on the critical path, before any connection exists. A client cannot open TCP or TLS to an edge until the hostname resolves. An uncached lookup “may involve several network accesses and an arbitrary amount of time” (RFC 1035, section 2.2). The specification puts no number on that and neither should you. A warm resolver cache removes the lookup from the path entirely. A cold one adds a full walk down the tree.
- It is the steering plane. The location part is standard behaviour. “Many Authoritative Nameservers today return different responses based on the perceived topological location of the user” (RFC 7871, section 1). That is what geo DNS means. Health and load steering are vendor features rather than protocol behaviour. They make DNS a failover mechanism as well as a placement one. Akamai's Mirror Failover “monitors the primary data center, and as long as it is up, GTM sends users there. If the primary data center goes down, GTM sends users to the backup data center”. Akamai's Weighted Random Load Balancing with Load Feedback “begins to shift the traffic to other data centers that are below their target load” (Akamai, property type descriptions).
- The TTL is the clock on every steering change. TTL “controls how long each record is cached and — as a result — how long it takes for record updates to reach your end users” (Cloudflare, Time to Live). A new edge address, a failover, a migration: none of them reach a user before that user's cached answer expires.
- What the CDN sees is the resolver, not the user. “Since most queries come from Intermediate Recursive Resolvers, the source address is that of the Recursive Resolver rather than of the query originator”. Resolvers were traditionally close to their clients, but “a class of Recursive Resolvers has arisen that handles query sources that are often not topologically close”. Such cases “lead to less than desirable responses from topology-sensitive Authoritative Nameservers” (RFC 7871, section 1). ECS is the mechanism that narrows that gap.
- The authoritative DNS is a single point of failure for the whole property. If the zone stops answering, no client reaches any edge, however healthy the edges are. The long-standing operational rule is diversity. “Secondary servers must be placed at both topologically and geographically dispersed locations on the Internet, to minimise the likelihood of a single failure disabling all of them” (RFC 2182, section 3.1). CDNs sell their DNS on exactly that footing. Cloudflare pitches its authoritative service as protecting a domain “from DDoS attacks and route leaks and hijacking” (Cloudflare DNS). See anycast.
What CDNs do
- Cloudflare runs its own “fast, resilient, and easy-to-manage authoritative DNS service”, marked available on all plans (Cloudflare DNS). “By default, all proxied records have a TTL of Auto, which is set to 300 seconds. This value cannot be edited”. So “recursive resolvers will not cache them for longer than 300 seconds (five minutes)”. Cloudflare adds the caveat that “It may take longer than 5 minutes for you to actually experience record changes, as your local DNS cache may take longer to update”. For DNS-only records the floor depends on your plan. “you can choose a TTL between 30 seconds (Enterprise) or 60 seconds (non-Enterprise) and 1 day” (Cloudflare, Time to Live). CNAME flattening “allows you to use a CNAME record at your zone apex” by returning “the final IP address instead of a CNAME record” (Cloudflare, CNAME flattening).
- Fastly takes the opposite position: “Fastly does not provide a managed DNS service.” A CNAME is “The preferred method of connecting a domain to Fastly”. That is because it lets “a Fastly authoritative name server” answer the query and “route the user to the POP that offers the most consistently high performance for that specific end user”. Anycast A and AAAA records are the apex option, since “Anycast is required for apex domains (e.g., example.com) because they do not support CNAMEs in DNS”. But they cost you that steering. “When using anycast IPs, connections select a Fastly POP based on routing decisions made outside of Fastly and over which Fastly has less ability to direct traffic for best performance”. Fastly also recommends against ALIAS-style apex emulation, because “they often result in inferior performance or interfere with the ability of records to update automatically” (Fastly, Routing traffic to Fastly).
- Akamai steers with Global Traffic Management (GTM). Its property types are defined by the DNS answer they hand back. “Map by geographic location” means “GTM returns a CNAME based on the location (country, or country and state) of the requester”. Equivalents are keyed on AS number and on CIDR block. An IP Version Selector “lets you return different answers depending on whether the query is for an IPv4 (A record) or IPv6 (AAAA record) address” (Akamai, property type descriptions). Akamai describes the service as making “intelligent routing decisions” that “are based on real-time data center performance health and on global Internet conditions” (Akamai, Welcome to Global Traffic Management). Watch the vocabulary: on Akamai's own page, Performance-Based Load Balancing means “Traffic is sent to the target that is geographically closest to the user”, not to the lowest measured latency.
- AWS CloudFront hands the choice to DNS and says so plainly: “DNS routes the request to the CloudFront POP (edge location) that can best serve the request, typically the nearest CloudFront POP in terms of latency” (AWS, How CloudFront delivers content).
- Azure Front Door in its current Standard and Premium tiers uses “Unicast for both DNS (Domain Name System) and HTTP (Hypertext Transfer Protocol) traffic”. This is combined with “Azure Front Door's Traffic Manager based load management architecture” to pick the PoP. “If the preferred Front Door edge location is unhealthy, all traffic automatically moves to the next optimal edge location”. The older Classic tier used “Anycast for both DNS (Domain Name System) and HTTP (Hypertext Transfer Protocol) traffic”. Classic “retires on March 31, 2027” and already “no longer supports profile creation, new domain onboarding, or managed certificates”, so do not design against it (Azure, Traffic acceleration).
Watch out for
- A CNAME cannot share a name with anything else, which is what makes the apex awkward. “If a CNAME RR is present at a node, no other data should be present” (RFC 1034, section 3.6.2). RFC 2181 narrows the exception to signing records alone: an alias name “may, if DNSSEC is in use, have SIG, NXT, and KEY RRs, but may have no other data” (RFC 2181, section 10.1). The apex has to carry the zone's own NS records. So a CNAME there does not merely fail, it takes the zone with it. “Especially do not try to combine CNAMEs and NS records”, because then “the NS entries are ignored. Therefore all the hosts in the podunk.xx domain are ignored as well” (RFC 1912, section 2.4). Vendor flattening and anycast address records are workarounds. The standards-track answer is the HTTPS and SVCB record types, whose “primary purpose of AliasMode is to allow aliasing at the zone apex, where CNAME is not allowed”. Unlike CNAME, they “do not affect the resolution of other RR types” (RFC 9460, section 2.4.2; worked apex example in section 10.4.2). The catch is client support: clients that do not know the type “will just ignore the new record”, so the A and AAAA records still have to be correct.
- The TTL is a ceiling on reuse, not a purge. Lowering it does nothing to answers already handed out. Every resolver holding the old value keeps serving it until its own copy expires (RFC 1035, section 3.2.1). Plan a cutover backwards from the TTL that was in effect before you touched anything, not from the one you just set.
- A deleted or mistyped record keeps failing after you fix it. Negative answers are cached on a clock derived from the SOA rather than from the record you have just created. A zone with a large SOA minimum can keep returning NXDOMAIN long after the name exists (RFC 2308, section 5).
- DNS is neither authenticated nor private by default, and those are two different problems. DNSSEC extensions “provide origin authentication and integrity protection for DNS data, as well as a means of public key distribution”. They explicitly “do not provide confidentiality” (RFC 4033, section 1). More bluntly, “By intention, DNSSEC does not protect request and response privacy” (RFC 7858, section 1). DNS over HTTPS attacks the other half. It “encrypts DNS traffic and requires authentication of the server” and so “mitigates both passive surveillance” and “active attacks” (RFC 8484, section 8.1). Mind the scope, though: RFC 8484 “focuses on communication between DNS clients (such as operating system stub resolvers) and recursive resolvers” (RFC 8484, section 1), and DNS over TLS likewise “focuses on securing stub-to-recursive traffic” (RFC 7858, abstract). Encrypting a client's lookup hides it from the local network, not from the CDN's authoritative servers.
- ECS carries a privacy cost, and the RFC's default is off. “We recommend that the feature be turned off by default in all nameserver software, and that operators only enable it explicitly in those circumstances where it provides a clear benefit for their clients” (RFC 7871, section 2). It trades part of the client's network address for steering accuracy. It is not a free upgrade.
- Steering granularity is coarse without ECS. Every user behind one centralised resolver looks like one location. A whole ISP or a whole public-resolver population can land on the same edge (RFC 7871, section 1).
Best practice
- Set the TTL from how fast the answer has to be able to change. Keep it short for anything the CDN steers or fails over. Keep it longer for records that never move. Do a cutover in two steps: “lower the TTL on the existing record (e.g., to 60), wait for a period equal to the longer TTL, and then update the record to point to Fastly”. Keep the new records low “to enable rapid rollback”, and, “following a successful migration, consider increasing it to an hour or more” (Fastly, Routing traffic to Fastly).
- Point the hostname at the CDN by name, not by address. Edge addresses belong to the CDN and it will change them. Cloudflare pins proxied records at 300 seconds precisely so that “potential changes to the assigned anycast IP address will take effect quickly” (Cloudflare, Time to Live). Use a CNAME below the apex. At the apex, use whatever your vendor actually supports: flattening, anycast A and AAAA records, or an HTTPS record in AliasMode (RFC 9460, section 2.4.2). Do not use ALIAS emulation a vendor warns against.
- Get the actors right on ECS. You cannot switch it on for your users. The recursive resolver decides whether to forward the client's prefix at all (RFC 7871, section 5). Your authoritative provider decides whether to act on it. The RFC tells nameserver operators to leave it off unless there is “a clear benefit for their clients” (RFC 7871, section 2). Treat improved geographic accuracy as something to measure, not assume.
- Sign the zone, and be precise about what signing buys. DNSSEC “adds cryptographic signatures to your DNS records, preventing anyone else from redirecting traffic intended for your domain” (Cloudflare DNS). That protection only applies for resolvers that validate, and it buys integrity, not confidentiality (RFC 4033, section 1).
- Spread the authoritative servers, and treat DNS as part of the availability design. “Secondary servers must be placed at both topologically and geographically dispersed locations on the Internet, to minimise the likelihood of a single failure disabling all of them” (RFC 2182, section 3.1). A multi-CDN design that still hangs off one authoritative provider has not removed the single point of failure.
- Verify from the user's vantage point, not the zone's. dig +trace walks the referral chain from the root. dig @resolver name shows what one specific recursive resolver is handing out right now, which is the only thing your users experience. Fastly's own cutover check is the same tool: “To test that your DNS provider is correctly serving the new DNS records, you can use the dig tool in a terminal” (Fastly, Routing traffic to Fastly).
Examples
# Trace the full DNS resolution chain
$ dig +trace cdn.example.com
# Check which CDN edge you're hitting
$ dig +short cdn.example.com
104.16.132.229
# Check DNS response time
$ dig cdn.example.com | grep 'Query time'
;; Query time: 12 msec
# See EDNS Client Subnet in action
$ dig @8.8.8.8 +subnet=103.0.0.0/24 cdn.example.com
# Returns different IP than without subnet
Frequently Asked Questions
The distributed, hierarchical naming system that resolves names like example.com to addresses. A query-response lookup service, not a routing protocol. For a CDN it is also the steering plane: the authoritative nameserver picks which edge address to return, and the TTL caps how fast that changes.
# Trace the full DNS resolution chain
$ dig +trace cdn.example.com
# Check which CDN edge you're hitting
$ dig +short cdn.example.com
104.16.132.229
# Check DNS response time
$ dig cdn.example.com | grep 'Query time'
;; Query time: 12 msec
# See EDNS Client Subnet in action
$ dig @8.8.8.8 +subnet=103.0.0.0/24 cdn.example.com
# Returns different IP than without subnet