BGP (Border Gateway Protocol)
BGP (Border Gateway Protocol, currently BGP-4) is the internet's inter-autonomous-system routing protocol: networks announce the IP prefixes they can reach and choose the path traffic takes between them. It carries no latency or load data, but it is what makes CDN anycast work.
Also known as BGP-4.
Full Explanation
BGP stands for Border Gateway Protocol. The current version is BGP-4, specified in RFC 4271. It is an inter-autonomous-system routing protocol. Autonomous systems (ASes) use BGP to announce which IP prefixes they can reach. They also use it to choose which path traffic takes between networks. Every announcement carries an ordered list of the ASes that the reachability information has traversed. This list is the AS path. The AS path lets any network compare competing routes and prune loops. That is why BGP is called a path-vector protocol. BGP is not DNS. It is not an interior routing protocol such as OSPF. It does not forward packets itself. It also carries no measurement of latency, server load or server health. It only exchanges reachability information between network domains and selects one best route per prefix. For a CDN, one consequence dominates all the others: BGP is what makes anycast work. A single IP address can then be served from every location. Withdrawing an announcement is how traffic is moved off a failed site.
How it works
Two BGP speakers open a session over TCP port 179. They exchange UPDATE messages (RFC 4271 section 3). Each UPDATE carries IP prefixes plus path attributes. The most important attribute is AS_PATH. A session with a speaker in another AS is external BGP (eBGP). A session inside your own AS is internal BGP (iBGP). When advertising a route to an external peer, a speaker prepends its own AS number to AS_PATH. When advertising to an internal peer, it must leave AS_PATH unmodified (section 5.1.2).
Loop prevention happens at the receiver, not the sender. A speaker scans the full AS path of every route it learns. If its own AS number appears anywhere in that path, it excludes the route from its decision process (section 9.1.2). Cisco routers, for example, deny such paths on ingress. They never install them in the BGP RIB (Cisco, Select BGP Best-path Algorithm). Prepend on the way out, reject on the way in. That pair is how inter-domain routing loops are pruned.
A speaker keeps every route it learns. It then runs the decision process to pick exactly one best route per prefix. It first assigns each route a degree of preference. For routes learned from internal peers, this comes from LOCAL_PREF. For routes learned from external peers, it comes from local policy. The route with the higher degree of preference must be preferred (section 9.1.1, section 5.1.5). Among routes that tie on preference, the speaker applies the tie-breakers of section 9.1.2.2 in this order, until one route is left:
- Shortest AS path (an AS_SET counts as 1, no matter how many ASes are in the set).
- Lowest ORIGIN value.
- Lowest MULTI_EXIT_DISC (MED), which is comparable only between routes learned from the same neighbouring AS.
- Routes received via eBGP in preference to routes received via iBGP.
- Lowest interior cost to the route's NEXT_HOP.
- Lowest BGP Identifier, and finally lowest peer address.
The winner is installed in the Loc-RIB and the forwarding table. Subject to export policy, it is then re-advertised to neighbours. Do not expect a router's documentation to match that list step for step. RFC 4271 lets an implementation use any algorithm that produces the same results, and vendors add steps of their own. Cisco compares WEIGHT, a router-local Cisco-specific parameter, before LOCAL_PREF. Only then does it walk the familiar AS_PATH, origin, MED and eBGP-over-iBGP sequence (Cisco).
BGP-4 itself carries IPv4 prefixes. IPv6 and other address families ride on the multiprotocol extensions defined in RFC 4760. These extensions are backward compatible with speakers that do not support them.
Why it matters for a CDN
Anycast. The CDN announces the same prefix from every point of presence (PoP). Each user's own network routes to whichever announcing PoP its routing prefers. One address then serves the world with no per-user mapping. RFC 4786 section 3.1 explains why. For services distributed using anycast, there is no inherent requirement for referrals, and none for name-based distribution such as round-robin DNS. Instead, the routing system decides which node is used for each request. This decision is based on the topological design of the routing system, and on the point in the network at which the request originates.
Traffic engineering. Peering policy, LOCAL_PREF, MED and AS-path prepending all change which PoP a given network reaches. RFC 4786 section 3.1 states it plainly. The anycast node chosen to service a particular query can be influenced by the traffic engineering capabilities of the routing protocols. These protocols make up the routing system. The degree of influence available depends on the scale of that routing system.
Failover and drains. Withdraw the prefix at one PoP, and traffic converges on the next-best announcement. This is how a CDN drains a site for maintenance. It is also how a CDN survives the loss of one site.
Attack absorption. RFC 4786 section 3.2 lists common objectives of anycast. One is the mitigation of non-distributed denial-of-service attacks, by localising damage to single anycast nodes. Another is the constraint of distributed attacks or flash crowds to local regions. This gives the opportunity for traffic to be handled closer to its source, using high-performance peering links rather than oversubscribed paid transit.
What CDNs do
- Cloudflare answers for proxied hostnames from shared IP ranges. In its current documentation, these ranges form the backbone of Cloudflare's anycast network. Anycast is a routing method where the same IP address is announced from data centers worldwide. This way, each visitor's request is routed to a nearby data center (Cloudflare, Cloudflare IP addresses). Enterprise customers can instead have Cloudflare announce prefixes they lease or own (BYOIP). Magic Transit announces a customer's own prefixes by BGP, so that all traffic for that network is pulled through Cloudflare for filtering.
- Fastly offers anycast IP addresses to paid customers, chiefly for content that must sit on a zone apex where a CNAME is not allowed. Its documentation warns that anycast addressing methods do not offer Fastly as much flexibility in routing requests. As a result, its anycast options may not be as performant as its CNAME-based system. Fastly recommends the CNAME-based system for as much content as possible, particularly for large resources and streaming video (Fastly, Using Fastly with apex domains).
- Akamai pairs anycast with its own mapping, rather than relying on routing alone. Edge IP Binding uses anycast routing to provide a fixed set of IP addresses or CIDR blocks, up to 20 anycast addresses per Edge IP Binding-enabled edge hostname. That set is coupled with proprietary mapping schemes to identify the best edge server to respond to the client's request (Akamai TechDocs).
Watch out for
- Topologically near is not the same as fast. RFC 4786 section 3.2 is explicit. Topological nearness within the routing system does not, in general, correlate to round-trip performance across a network. In some cases response times may see no reduction, and may increase. Routing picks a path from topology and policy, not from measured performance.
- Anycast on its own is blind to your servers. Akamai explains why plain anycast was not enough. Too many server stacks sharing the same CIDR block can pollute BGP routing tables. Plain anycast also does not consider server load, capacity, or availability. RFC 4786 section 3.1 adds that load-balancing between anycast nodes is typically difficult to achieve, and distribution is generally unbalanced.
- Convergence is not instant. RFC 4271 rate-limits advertisements. The suggested default MinRouteAdvertisementIntervalTimer is 30 seconds on eBGP connections and 5 seconds on iBGP connections (section 10). A classic measurement study of the deployed internet found that the delay in interdomain path failovers averaged three minutes over two years. Some failovers triggered routing table fluctuations lasting up to fifteen minutes (Labovitz et al., Delayed Internet Routing Convergence, IEEE/ACM Transactions on Networking 9(3), 2001). Those figures are from 2001 and depend on the networks involved. Measure your own failover rather than trusting either number. But do not design for millisecond anycast failover.
- Route leaks. A route leak is the propagation of routing announcements beyond their intended scope. It happens in violation of the intended policies of the receiver, the sender, or an AS along the preceding AS path. The result can be redirection of traffic through an unintended path that may enable eavesdropping or traffic analysis. Leaks can be accidental or malicious, but most often arise from accidental misconfigurations (RFC 7908 section 2).
- Hijacks. Nothing inside BGP checks ownership. As Cloudflare puts it, any route can be originated and announced by any random network, independent of its rights to announce that route. That is why an out-of-band system is needed to manage which network may announce which route (Cloudflare, RPKI - the required cryptographic upgrade to BGP routing).
- RPKI covers the origin only. Prefix origin validation checks that the AS claiming to originate a prefix is authorised by the prefix holder. But complete path attestation against the AS_PATH attribute is outside its scope (RFC 6811 section 1). A route can be RPKI-valid and still have travelled a path nobody intended.
- Very specific prefixes may not travel. Acceptable specificity is agreed per peering. RFC 7454 section 6.1.3 records the RIPE community's documentation on this, from the time that RFC was written. It states that IPv4 prefixes longer than /24 and IPv6 prefixes longer than /48 are generally neither announced nor accepted in the internet. It notes these values may change. Size an anycast prefix with that in mind, instead of assuming a /26 will propagate.
- Long-running flows and per-packet load balancing. RFC 4786 section 4.1 warns that, especially for long running flows, there are potential failure modes using anycast that are more complex than a simple destination-unreachable failure using unicast. For services anycast across the global internet, equal-cost paths are normally not a consideration. This is because BGP's exit selection algorithm usually selects a single consistent exit per destination. But per-packet ECMP across links to different neighbour ASes can deliver packets of one transaction to different anycast nodes. This can effectively make the service unavailable (section 4.4.3).
Best practice
- Publish a Route Origin Authorization for every prefix you announce. A ROA is a digitally signed object. It provides a means of verifying that an IP address block holder has authorized an AS to originate routes to one or more prefixes within the address block (RFC 6482). On the receiving side, origin validation assigns each route a state of Valid, Invalid or NotFound. An implementation must be able to match and set that state in route policy (RFC 6811 sections 2 and 3).
- Configure an explicit import and export policy on every eBGP session. RFC 8212 updates RFC 4271: routes from an eBGP peer are not eligible in the decision process if no explicit import policy has been applied. They are also not added to that peer's Adj-RIB-Out if no explicit export policy has been applied. This turns the classic accident of leaking a full table into a no-op.
- Filter prefixes inbound and outbound on every peering (RFC 7454 section 6). Cap what you accept. It is recommended to configure a limit on the number of routes accepted from a peer. This limit should be lower than a full internet table for peers, and higher for upstreams that provide full routing. Review those limits regularly (section 8).
- Protect the session where the trade-off is worth it. TTL security (GTSM) is the cheap win: network administrators should implement TTL security on directly connected BGP peerings (RFC 7454 section 5.2). For authentication, TCP-AO should be preferred over the older TCP MD5 where it is implemented. The same RFC, though, calls TCP session protection not required, even across shared networks such as IXPs. It asks operators to weigh the configuration and key-management overhead (section 5.1).
- Treat anycast as resilience and coarse load distribution, rather than as a load balancer. Load-balancing between anycast nodes is typically difficult to achieve (RFC 4786 section 3.1). Steer with LOCAL_PREF, MED and AS-path prepending. If you need real latency or load control, pair BGP proximity with your own mapping or health data. This is precisely what Akamai's Edge IP Binding does.
- Budget for convergence in the failover design. Verify what the internet actually sees: check your announcements and their RPKI validity from outside your own network, rather than assuming propagation.
Examples
This shows what BGP announcements look like and how to inspect them:
# View BGP routes for a prefix using a looking glass
# Many networks provide public looking glass servers
# Check BGP route info via RIPE Stat
curl -s "https://stat.ripe.net/data/looking-glass/data.json?resource=104.16.0.0/13" | python3 -m json.tool
# Use bgp.he.net to look up AS paths
# Example: AS path to Cloudflare (AS13335)
# Your ISP (AS1234) -> Transit (AS5678) -> Cloudflare (AS13335)
# Check if RPKI validation is in place
curl -s "https://stat.ripe.net/data/rpki-validation/data.json?resource=AS13335&prefix=104.16.0.0/13"
In a CDN anycast setup, the same prefix comes from many PoPs. BGP picks the closest one from the AS path, not from physical distance. A user in Amsterdam might reach a PoP in Frankfurt when the path is shorter.
Frequently Asked Questions
BGP (Border Gateway Protocol, currently BGP-4) is the internet's inter-autonomous-system routing protocol: networks announce the IP prefixes they can reach and choose the path traffic takes between them. It carries no latency or load data, but it is what makes CDN anycast work.
This shows what BGP announcements look like and how to inspect them:
# View BGP routes for a prefix using a looking glass
# Many networks provide public looking glass servers
# Check BGP route info via RIPE Stat
curl -s "https://stat.ripe.net/data/looking-glass/data.json?resource=104.16.0.0/13" | python3 -m json.tool
# Use bgp.he.net to look up AS paths
# Example: AS path to Cloudflare (AS13335)
# Your ISP (AS1234) -> Transit (AS5678) -> Cloudflare (AS13335)
# Check if RPKI validation is in place
curl -s "https://stat.ripe.net/data/rpki-validation/data.json?resource=AS13335&prefix=104.16.0.0/13"
In a CDN anycast setup, the same prefix comes from many PoPs. BGP picks the closest one from the AS path, not from physical distance. A user in Amsterdam might reach a PoP in Frankfurt when the path is shorter.
Yes. BGP (Border Gateway Protocol) is also known as BGP-4. BGP (Border Gateway Protocol, currently BGP-4) is the internet's inter-autonomous-system routing protocol: networks announce the IP prefixes they can reach and choose the path traffic takes between them. It carries no latency or load data, but it is what makes CDN anycast work.