ECMP (Equal-Cost Multi-Path)
Equal-Cost Multi-Path (ECMP) is a forwarding technique: where a router holds several next hops of equal cost to one destination, it hashes a packet's flow fields to pick one, so every flow keeps one path. It turns existing path diversity into capacity and failover in the router, not a load balancer.
Also known as Equal-Cost Multi-Path routing, Equal-Cost Multipath, ECMP routing.
Full Explanation
Equal-Cost Multi-Path (ECMP) is a forwarding technique. It applies when a router holds more than one next hop of equal routing cost to the same destination. Instead of installing one best path and leaving the others idle, the router installs them all. It chooses between them per packet, by hashing the header fields that identify a flow. So every packet of a flow follows the same next hop. RFC 2992 describes it as “a routing technique for routing packets along multiple paths of equal cost” in which “the forwarding engine identifies paths by next-hop”.
ECMP is not a load balancer. People often get that distinction wrong. It holds no per-connection state, and it has no health signal beyond what the routing protocol tells it. It also does not look at how much traffic a path is already carrying: it spreads flows, not bytes. An L4 load balancer does all of those things. That is why, at CDN scale, one is usually placed behind ECMP rather than instead of it. Cloudflare put the gap plainly when it explained why it built exactly that on top: “ECMP does not allow us to do dynamic load balancing by adjusting the share of connections going to each server” (Cloudflare, Unimog). ECMP also creates no path diversity of its own. It can only use equal-cost next hops that the topology and the routing protocol have already produced. Under BGP, those next hops do not appear at all until multipath is explicitly enabled. The helicopter view: ECMP is how a router turns existing path diversity into aggregate capacity and automatic failover, at flow granularity, for free.
How it works
Three things must line up. The routing protocol has to offer several next hops for one prefix. The router has to install more than one of them. The forwarding engine needs a rule for choosing between them.
- Equal-cost paths must be installed, not merely exist. An IGP installs them readily. RFC 2991 notes that OSPF and IS-IS “explicitly allow” ECMP. RFC 4786 section 4.4.3 adds that “equal-cost paths are commonly supported in IGPs”. BGP is the opposite. Its decision process ends in one best path, so multipath is a separate feature. It must be turned on, and it applies its own equality test. In FRRouting it is maximum-paths (1 to 128, and “the maximum value listed, 128, can be limited by the ecmp cli for bgp or if the daemon was compiled with a lower ecmp value”). The multi-path check treats as equal “only routes received via iBGP with identical AS_PATHs or routes received from eBGP neighbours in the same AS”, unless bgp bestpath as-path multipath-relax is set (FRRouting BGP documentation).
- A hash of the flow picks the next hop. In the hash-threshold method analysed in RFC 2992 section 1, “the router first selects a key by performing a hash (e.g., CRC16) over the packet header fields that identify a flow”. It also notes that “the N next-hops have been assigned unique regions in the key space”. The region containing the key names the next hop. Which header fields go into the hash is an implementation choice, not part of ECMP. RFC 2992’s own example uses only “the source and destination fields of the packet”. Cloudflare’s Magic Transit documents that “the hash always uses the source and destination IP addresses” and that “for TCP and UDP packets, the hash includes the source and destination ports as well”.
- Encapsulation can hide the flow key. Once traffic rides a tunnel, the fields the hash wants are no longer in the outer header. The VL2 paper hit this in 2009: “some inexpensive switches cannot correctly retrieve the five-tuple values (e.g., the TCP ports) when a packet is encapsulated with multiple IP headers”. It worked around this by having the sending host compute “a hash of the five-tuple values” and write it “into the source IP address field, which all switches do use in making ECMP forwarding decisions”. Linux exposes the same problem as a policy choice: hash on the inner header rather than the outer one. For a CDN running GRE or IPsec overlays, this decides whether a tunnel is one flow or many.
- Every packet of a flow takes one path. “As long as the region boundaries remain unchanged the same next-hop will be chosen for a given flow” (RFC 2992 section 2.2). The only thing that moves those boundaries is adding or removing a next hop. That is what makes ECMP safe for TCP. Spreading per packet instead breaks three things at once, per RFC 2991 section 2. The path MTU can change packet by packet, “negating the usefulness of path MTU discovery”. Differing latencies make packets “always arrive out of order”, pushing TCP into fast-retransmit. And “ping and traceroute are much less reliable in the presence of multiple paths and may even present completely wrong results”.
- Changing the path set re-shuffles flows. RFC 2992 calls this disruption: “the measurement of how many flows have their paths changed due to some change in the router”. Section 2.2 works out that for hash-threshold “the range of possible disruption is (1/4, 1/2]” when a next hop is added or removed. Section 3 compares the alternatives. Modulo-N is “the most disruptive of the algorithms” at (N-1)/N. Highest random weight (HRW) has “minimal disruption (i.e., disruption due to adding or removing a next-hop is always 1/N.)”, at the cost of an O(N) lookup. That is the same minimal-disruption property that consistent hashing is chosen for elsewhere. RFC 2992 also recommends “adding new regions to the center rather than the ends”, because removing an edge region is the worst case.
- Even splitting is a property of the hash, not of ECMP. RFC 2992 section 2 assumes regions of equal size. It states that “if the output of the hash function is uniformly distributed the distribution of flows amongst paths will also be uniform, and so the algorithm will properly implement ECMP”. It explicitly declines to analyse balancing in depth, because that depends entirely on the chosen hash function. RFC 2991 section 5 adds the condition operators forget: “the commonly used hash functions only become uniformly distributed when the number of inputs is relatively large”, so these algorithms “are more applicable to routers used to route many flows”. ECMP is a large-numbers mechanism. On a handful of flows, it simply does not balance.
On Linux the choice is exposed as sysctls, documented in ip-sysctl. fib_multipath_hash_policy selects 0 for layer 3, 1 for layer 4, 2 for “layer 3 or inner layer 3 if present”, or 3 for a custom field set. The default is 0, layer 3 only, so ports are not hashed until you ask for them. With policy 3, fib_multipath_hash_fields is a bitmask over outer and inner addresses, protocol, flow label and ports, defaulting to 0x0007 (source IP, destination IP and IP protocol). fib_multipath_hash_seed defaults to 0, which means “the seed value used for multipath routing defaults to an internal random-generated one”. The same documentation warns that “the actual hashing algorithm is not specified” and that “there is no guarantee that a next hop distribution effected by a given seed will keep stable across kernel versions”. So never treat a measured split as a contract.
Why it matters for a CDN
A CDN’s capacity is built out of parallel links and parallel machines. ECMP is the cheapest way to use them. A point of presence with several equal-cost uplinks to transit providers, peers or the CDN backbone carries the sum of those links rather than the width of the best one. The same mechanism fans one address across a rack of servers inside the site. ECMP’s “original purpose is to allow traffic to be spread across multiple paths between two locations, but it is commonly repurposed to spread traffic across multiple servers within a data center” (Cloudflare). Both come from the existing routing table. No extra device sits in the data path.
Failure handling is the other half. When a path is withdrawn its flows are re-hashed onto the survivors as soon as routing converges, with no coordination and no state to hand over. The length of that gap is dominated by detection, not by the recomputation. RFC 5880 section 1 observes that “the time to detect failures (‘Detection Times’) available in the existing protocols are no better than a second” for the “Hello” mechanisms of routing protocols. That is why operators add BFD, whose goal is “low-overhead, short-duration detection of failures in the path between adjacent forwarding engines”. Note that the two directions are not symmetric. VL2 measured a Clos fabric losing links and getting them back. OSPF “re-converges quickly (sub-second) after each failure”, but “restoration, however, is delayed by the conservative defaults for OSPF timers that are slow to act on link restoration. Hence, VL2 fully uses a link roughly 50s after it is restored.” Capacity comes back an order of magnitude more slowly than it goes away.
Inside a data centre, ECMP is what makes a Clos (leaf-spine) fabric worth building. The VL2 paper measured the alternative. With ECMP turned on in the layer-3 part of a conventional tree, “the conventional topology offers at most two paths”. Above the top-of-rack switch, “the basic resilience model is 1:1”, which “forces each device and link to be run up to at most 50% of its maximum utilization”. In a folded Clos with n intermediate switches, “the failure of any one of them reduces the bisection bandwidth by only 1/n–a desirable graceful degradation of bandwidth”. VL2 reaches it with nothing more exotic than link-state routing, ECMP forwarding and IP anycast on commodity switches.
Where ECMP and anycast meet, read the small print: the popular version of this story is wrong. For a prefix anycast across the global Internet over BGP, RFC 4786 section 4.4.3 says “equal-cost paths are normally not a consideration: BGP’s exit selection algorithm usually selects a single, consistent exit for a single destination regardless of whether multiple candidate paths exist”. The exposure appears where BGP multipath is deliberately enabled, or where the anycast prefix is carried in an IGP. That second case is exactly the CDN’s own network. There, “where multiple, equal-cost paths exist and lead to different Anycast Nodes, there is a risk that different request packets associated with a single transaction might be delivered to more than one node”. TCP is the first casualty, because services provided over it “necessarily involve transactions with multiple request packets, due to the TCP setup handshake”. The same section names the remedy: careful IGP link metrics, or ECMP selection algorithms “which cause a single node to be selected for a single multi-packet transaction”. Short single-packet exchanges are the easy case. RFC 4786 section 4.1 offers “DNS transactions over UDP transport”, that is DNS over UDP, as its example of a transaction that may be “carried out using a single packet request and a single packet reply”. Be careful quoting that as a rule, though: the RFC “deliberately avoids prescribing rules as to which protocols or services are suitable for distribution by anycast”. The test it does give is that node selection “ought to be stable for substantially longer than the expected transaction time”.
What CDNs do
- Cloudflare: ECMP from the router, then a load balancer on top. Cloudflare deployed it early: “back in 2015 we deployed ECMP routing - Equal Cost Multi Path - within our datacenters”. This “allowed us to spread traffic heading to a single IP address across multiple physical servers”, described then as “a third layer of load balancing” after DNS and anycast (Cloudflare, 2018). It was configured through internal BGP: “it only required adjusting the BGP metrics to equal values, the ECMP-enabled router does the rest of the work”, hashing “(src ip, src port, dst ip, dst port)” for TCP (Cloudflare, 2015). That is no longer the whole picture. The update is the interesting part: “Cloudflare relied on ECMP alone to spread load across servers before we deployed Unimog”, and “these drawbacks mean that ECMP alone is not an effective approach”. Today the router still spreads packets with ECMP, but every server runs the Unimog load balancer. They all share one forwarding table, so “it doesn’t matter which packets are sent by the router to which servers, and so ECMP re-hashes are a non-issue”.
- Cloudflare Magic Transit: ECMP across customer tunnels. “When the priority values for prefix entries match, Cloudflare uses equal-cost multi-path (ECMP) packet forwarding to route traffic”, and “the ECMP algorithm divides the hash for each packet by the number of equal-cost next hops. The modulus (remainder) determines the route the packet takes”. That is the modulo-N variant, the one RFC 2992 section 3 ranks as the most disruptive when the next-hop set changes. Note the scope of the escape hatch: “you can apply an optional weight value to static routes to modify ECMP tunnel distribution”, maximum 256. That applies to static routes, not BGP-learned ones, and “because ECMP balances flows probabilistically, the use of weights is only approximate” (Magic Transit traffic steering).
- AWS Transit Gateway: opt-in, and conditional on BGP attributes. For site-to-site VPN it is a gateway setting: “for VPN ECMP support, select this option if you need Equal Cost Multipath (ECMP) routing support between VPN tunnels”. This requires that “the advertised BGP ASN, then the BGP attributes such as the AS-path, must be the same”, and “to use ECMP, you must create a VPN connection that uses dynamic routing. VPN connections that use static routing do not support ECMP” (AWS documentation). For Connect (GRE) attachments the appliance must “advertise the same prefixes to the transit gateway with the same BGP AS-PATH attribute” and “the AS-PATH and Autonomous System Number (ASN) must match”. The gateway “can use ECMP between Connect peers for the same Connect attachment or between Connect attachments on the same transit gateway” but “cannot use ECMP between both of the redundant BGP peerings a single peer establishes to it”. On that path, “Bidirectional Forwarding Detection (BFD) is not supported” (AWS documentation).
Watch out for
- Per-packet selection, and only some of it. “ECMP algorithms that select a route on a per-packet basis rather than per-flow are commonly referred to as performing ‘Per Packet Load Balancing’ (PPLB)”. RFC 4786 section 4.4.3 is precise about when this hurts. PPLB across parallel links between the same pair of routers, or across diverse paths that converge to a single exit from the AS, “should cause no node selection problems”. It is PPLB across links to different neighbour ASes that have chosen different anycast nodes that “will, in general, cause request packets to be distributed across multiple Anycast Nodes”. The RFC judges such networks “pathological”, because they also cause persistent reordering.
- Rehash on any membership change. The disruption figures above are not theoretical. Cloudflare found that “ECMP is vulnerable to changes in the set of active servers, such as when servers go in and out of service”. It also found that “these changes cause rehashing events, which break connections to all the servers in an ECMP group” (Unimog). Note what counts as a change: on Magic Transit, “routing changes in the number of equal-cost next hops can cause traffic to use different tunnels. For example, dynamic reprioritization triggered by health check events can cause traffic to use different tunnels.” A health check flapping a member is therefore worse than losing it once.
- ICMP does not carry the flow key. “ECMP will indeed forward TCP packets in a session to the appropriate server. Unfortunately, it has no special knowledge of ICMP and it hashes only (src ip, dst ip)”. That source is some router out on the Internet, so “this may sometimes end up delivering the ICMP packets to a different server than the one handling the TCP connection”. A Packet-Too-Big lands on the wrong machine and Path MTU Discovery silently fails. Cloudflare’s fix was to “broadcast ICMP MTU messages to all servers”, released as the open-source pmtud daemon. The interim hack was pinning the IPv6 MTU to 1,280 bytes (Cloudflare, 2015).
- ECMP equalises flow counts, not bytes. “Because ECMP is probabilistic, the algorithm routes roughly the same number of flows through each tunnel. However, it does not consider the amount of traffic already sent through a tunnel when deciding where to route the next packet” (Magic Transit). So one very high-bandwidth connection saturates its path while the others idle. VL2 named the same hazard: “if ‘elephant flows’ are present, then the random placement of flows could lead to persistent congestion on some links”. Do not quote that as a measured result, though, because the same paragraph adds “our evaluation did not find this to be a problem on data-center workloads”. It suggests “re-hashing to change the path of large flows when TCP detects a severe congestion event” as the fix if it does occur.
- The group has a maximum width. Cloudflare: “routers impose limits on the sizes of ECMP groups, which means that a single ECMP group cannot cover all the servers in our larger edge data centers”. VL2 hit the same ceiling from the other side. It “defines several anycast addresses, each associated with only as many Intermediate switches as ECMP can accommodate”. In software the ceiling is explicit too: FRRouting caps maximum-paths at 128, and lower still if the daemon was compiled with a smaller ecmp value. Scaling past the group width needs another layer, not more ECMP.
- Predictability, in both directions. This one is usually stated backwards. RFC 2991 section 7 warns that “when next-hop selection is predictable, an attacker can synthesize traffic that will all hash the same, making it possible to launch a denial-of-service attack that overloads a particular path”. But it then points out that “a special case of this is when the same (single) next-hop is always selected”, so “such an attack is easiest when multipath is not being used”. ECMP reduces this exposure rather than creating it: “the more unpredictable the hash is, the harder it becomes to conduct a denial-of-service attack against any single link”. That is the argument for a non-default hash seed, not the kernel documentation, which states only that seed 0 means an internal random-generated seed and offers no rationale.
- Hashing on ports is not free. RFC 2991 section 3 notes that “including transport-layer information in the next-hop selection process can actually be problematic”: “if packets are fragmented, the transport-layer information may not be available in every packet”, and “having the choice of path depend on transport-layer fields may negate the benefit of caching information such as MTU for use in subsequent connections between the same endpoints”. Ports buy spread between busy host pairs. They are not what keeps a flow in order.
Best practice
- Before blaming the hash, confirm the paths are installed. On an IGP they usually are. On BGP nothing is multipath until maximum-paths is configured. Candidates still have to pass the AS_PATH equality test unless you relax it.
- Choose the hash fields for the spread you need, not for ordering. Layer-3 hashing already pins a flow to one path. Adding the ports only raises entropy between endpoints that talk a lot. Weigh that against fragments losing their ports, and against per-endpoint MTU caching (RFC 2991 section 3).
- Check what the hash can actually see. Inside a GRE or IPsec overlay, the outer header may be one tuple for the whole tunnel. Hash on the inner header instead (on Linux, policy 2 or a custom field set including the inner fields), or the tunnel becomes a single flow on a single path.
- Keep selection per flow, never per packet, wherever a multi-packet transaction is in play. For anycast that means the algorithms RFC 4786 describes as causing “a single node to be selected for a single multi-packet transaction”. Where transactions are long, consider RFC 4786’s split: an initialisation phase handled by anycast servers, and a sustained phase on non-anycast servers chosen during it.
- Pair ECMP with a sub-second failure detector such as BFD, since routing-protocol Hellos are “no better than a second”. Check whether your platform supports it: AWS Transit Gateway Connect does not. Budget separately for restoration, which is slower than failure.
- Where the path set changes often, pick the least disruptive selection you can afford: HRW moves 1/N of flows, hash-threshold between 1/4 and 1/2, modulo-N (N-1)/N. Where connections must survive a membership change outright, put a stateless L4 load balancer with a shared forwarding table behind ECMP, as Cloudflare did with Unimog.
- Set an explicit hash seed rather than inheriting a predictable one. Read per-next-hop counters after every convergence: on Linux, ip -s route show per nexthop. Treat any observed split as an observation, not a guarantee. The kernel does not promise it across versions.
- Leave headroom on each path sized for your largest single flow, not for the average, because ECMP will not move it for you.
Examples
# View ECMP routes on Linux
ip route show
# default
# nexthop via 10.0.0.1 dev eth0 weight 1
# nexthop via 10.0.0.2 dev eth1 weight 1
# Add ECMP route (Linux)
sudo ip route add 10.10.0.0/24 \
nexthop via 10.0.0.1 dev eth0 weight 1 \
nexthop via 10.0.0.2 dev eth1 weight 1
# Configure ECMP hash fields (Linux)
sudo sysctl -w net.ipv4.fib_multipath_hash_policy=1
# 0 = layer 3 only (src/dst IP)
# 1 = layer 4 (src/dst IP + ports) - recommended
# 2 = layer 3 or inner for tunnels
# BGP ECMP with FRRouting
router bgp 65001
maximum-paths 4
maximum-paths ibgp 4
# Verify ECMP is working (check counters)
ip -s route show 10.10.0.0/24
# Shows per-nexthop packet/byte counters
# Traceroute to see multiple paths
traceroute -f 2 -q 1 example.com
# Run multiple times to see different paths
Frequently Asked Questions
Equal-Cost Multi-Path (ECMP) is a forwarding technique: where a router holds several next hops of equal cost to one destination, it hashes a packet's flow fields to pick one, so every flow keeps one path. It turns existing path diversity into capacity and failover in the router, not a load balancer.
# View ECMP routes on Linux
ip route show
# default
# nexthop via 10.0.0.1 dev eth0 weight 1
# nexthop via 10.0.0.2 dev eth1 weight 1
# Add ECMP route (Linux)
sudo ip route add 10.10.0.0/24 \
nexthop via 10.0.0.1 dev eth0 weight 1 \
nexthop via 10.0.0.2 dev eth1 weight 1
# Configure ECMP hash fields (Linux)
sudo sysctl -w net.ipv4.fib_multipath_hash_policy=1
# 0 = layer 3 only (src/dst IP)
# 1 = layer 4 (src/dst IP + ports) - recommended
# 2 = layer 3 or inner for tunnels
# BGP ECMP with FRRouting
router bgp 65001
maximum-paths 4
maximum-paths ibgp 4
# Verify ECMP is working (check counters)
ip -s route show 10.10.0.0/24
# Shows per-nexthop packet/byte counters
# Traceroute to see multiple paths
traceroute -f 2 -q 1 example.com
# Run multiple times to see different paths
Yes. ECMP (Equal-Cost Multi-Path) is also known as Equal-Cost Multi-Path routing, Equal-Cost Multipath, ECMP routing. Equal-Cost Multi-Path (ECMP) is a forwarding technique: where a router holds several next hops of equal cost to one destination, it hashes a packet's flow fields to pick one, so every flow keeps one path. It turns existing path diversity into capacity and failover in the router, not a load balancer.
Related CDN concepts include:
- Point of Presence (PoP) — A Point of Presence (PoP) is one location where a network keeps its own servers, …
- Anycast — Anycast announces one IP address from many locations at once; the routing system, usually BGP, …