L7 Load Balancing
Load balancing at layer 7 of the OSI model, the application layer: the balancer parses each HTTP request and picks a backend by content (path, host, headers, query, cookies) rather than by IP and port. It terminates the client connection and proxies to an origin, so it costs more CPU than L4.
Also known as Layer 7 load balancing, Application load balancing, HTTP load balancing.
Full Explanation
L7 load balancing works at layer 7 of the OSI model, the application layer. For web traffic, that means HTTP. The balancer parses each request and picks a backend by what the request says: URL path, host, headers, query parameters, method, and cookies. AWS describes its Application Load Balancer exactly this way: it functions at the application layer, the seventh layer of the Open Systems Interconnection (OSI) model.
Three things it is not. It is not L4 load balancing. L4 routes on IP addresses and TCP/UDP ports. In A10 Networks' words, it has no awareness of the application payload. It is not a pass-through either. An L7 balancer terminates the client connection. It decrypts where the traffic is encrypted, then opens a new connection to the upstream. That makes it a reverse proxy, not a forwarder. And it is not free. Terminating, decrypting and parsing costs more CPU per request than L4 does.
Here is the consequence to remember. L4 chooses a target once per connection and keeps it there. AWS says exactly that of its Network Load Balancer, an L4 device: each individual TCP connection is routed to a single target for the life of the connection. L7 chooses per request instead. So a single keep-alive connection carrying twenty requests can be sent to twenty different backends. That is why a CDN edge server is an L7 router. Every cache and origin decision it makes needs the request read first.
How it works
- Terminate the client connection. For HTTPS, the balancer uses the domain's certificate to end the TLS session and decrypt the traffic. That lets it read the HTTP request. AWS: the load balancer uses certificates to terminate connections and decrypt requests from clients. Decryption happens only where it is needed. Plaintext HTTP is already readable.
- Parse the request. Path, host, headers, query string, method and cookies all become routing inputs.
- Match a rule, first match wins. An ALB evaluates the listener rules in priority order to determine which rule to apply. CloudFront compares the path against cache behaviours in the order in which cache behaviors are listed, and the first match determines which cache behavior is applied to that request. The default catch-all behaviour is always processed last. Path conditions route /api/* and /static/* to different pools. Host conditions let one balancer serve many domains.
- Reduce to healthy candidates. Health gates the choice before any algorithm runs. Elastic Load Balancing routes traffic only to the healthy targets. Cloudflare treats pool and endpoint health as the first of the three factors that decide where a request goes.
- Pick a backend in the pool. On an ALB, the default routing algorithm is round robin; alternatively, you can specify the least outstanding requests routing algorithm. NGINX offers round-robin, least-connected and ip-hash, plus per-server weights. Weights make the split deliberate rather than even. An ALB target group also accepts a weighted_random algorithm. A Cloudflare load balancer can set weights on pools to determine the percentage of traffic sent to each pool. That is the mechanism a canary or A/B rollout rides on. It is only available because the request was read first.
- Optionally pin the session. Cookie affinity is the usual shape. Cloudflare sets a __cflb cookie on the first request. It sends later requests from that client to the same endpoint while the cookie lasts and the endpoint stays healthy. NGINX can instead hash the client IP. That directs a client to the same server except when this server is unavailable.
- Write the request upstream. The balancer initiates a new TCP connection to the appropriate upstream server, and writes the request to the server. It may instead reuse a pooled connection. Cloudflare keeps connections alive to the origin and reuses them for up to 15 minutes after the last request. The client and the origin never connect to each other.
- Rewrite, redirect or answer. The balancer has parsed the request, so it can act on it three ways. It can redirect the request: ALB supports redirecting requests from one URL to another, and CloudFront has a Redirect HTTP to HTTPS viewer protocol policy. It can return a canned response of its own. Or it can adjust headers. One header it must handle: because it opened its own connection, the origin sees the balancer's address. The client's address has to be carried explicitly instead. RFC 7239 standardises the Forwarded field for this. It notes that proxies otherwise make the requests appear as if they originated from the proxy's IP address.
Why it matters for a CDN
Every interesting decision an edge makes is an L7 decision. That is because each one depends on the content of the request. Which origin to fetch from. Whether this is a cache hit, and under which cache key and variant. Which TTL and signing policy apply to this path. Whether to answer the redirect at the edge instead of paying a round trip to the origin. Which slice of a weighted rollout this request belongs to. An L4 device cannot make any of these decisions. It forwards bytes it has not read.
The same parse pays for security. Once traffic is decrypted at the edge, you can, as A10 puts it, run advanced security services, such as a web application firewall (WAF) to block SQL injections, or application-layer rate limiting to stop malicious bots. That is why WAF and rate limiting are L7 features by construction, not add-ons. Per-request routing also matters for multiplexed protocols. With HTTP/2, many requests share one connection. Only an L7 balancer can steer them individually. A WebSocket upgrade or a gRPC call is routed the same way, by reading the request first.
What CDNs do
- Cloudflare. Its Load Balancing product distributes traffic among pools according to pool health and traffic steering policies. It has three affinity modes: Cloudflare cookie, cookie with client-IP fallback, and HTTP header. Cookie sessions default to 23 hours. The cookie is always HttpOnly, but it is only marked Secure when Always Use HTTPS is enabled. The proxy doing this work is no longer NGINX. Cloudflare replaced it in 2022 with Pingora, an in-house Rust proxy. Pingora now handles almost every HTTP request that needs to interact with an origin server.
- AWS. The Application Load Balancer routes by listener rules over path, host, headers, methods, query parameters and source IP. CloudFront does the CDN equivalent with cache behaviours: a path pattern (for example, images/*.jpg) specifies to which requests you want this cache behavior to apply. Each behaviour names the one origin to forward matching requests to.
- NGINX. NGINX is a widely deployed L7 reverse proxy in front of origins. It uses an upstream group of servers, proxy_pass per location, and round-robin, least-connected or ip-hash selection. Note the edition split. Open-source NGINX does only in-band (passive) health checks, marking a server failed after max_fails errors. Active application health checks are an NGINX Plus subscription feature.
Watch out for
- The edge reads your plaintext. Termination is a trust boundary. That is the real cost, not key custody as such. Keyless SSL lets an operator keep their own custom certificates ... without exposing their TLS private keys. Yet the edge still sees the decrypted request. That option is also narrow. Cloudflare offers it as an Enterprise paid add-on only, and not with TLS 1.3.
- Do not assume L4 never terminates TLS. A10 lists L4 TLS handling as pass-through typically handled by the backend. AWS Network Load Balancer target groups accept TLS and QUIC. The dependable distinction is different: L7 parses the application message and L4 does not. It is not that only L7 holds certificates.
- Rule order is access control. AWS warns bluntly: define path patterns and their sequence carefully or you may give users undesired access to your content. Say a broad behaviour that does not require signed URLs precedes a narrow one that does. Then the first match wins, and the content is served unsigned.
- Path normalisation can move a request to another rule. CloudFront normalises URI paths per RFC 3986 before matching. So /a/b/..?c=1 matches an /a* behaviour, not /a/b*. Query strings and cookies are never considered when the path pattern is evaluated.
- Stickiness by source IP is unsound behind NAT. RFC 6269 states the general case: simple address-based identification mechanisms that are used to populate access control lists will fail when an IP address is no longer sufficient to identify a particular subscriber. Under carrier-grade NAT, many subscribers share one address. So ip-hash both concentrates them on one backend and fails to separate them.
- Keep-alive breaks naive cookie handling. A CDN sends many HTTP requests over one origin connection. So a load balancer that sets a session cookie once per TCP connection, and ignores later ones, loses affinity. Cloudflare names F5 BIG-IP as an example of such a balancer. Configure affinity to parse every request's cookie header.
- Forwarded client IPs are attacker-controlled. RFC 7239 is explicit that the field cannot be relied upon to be correct, as it may be modified, whether mistakenly or for malicious reasons, by every node on the way to the server, including the client making the request. Trust only the value your own edge wrote.
- L7 needs a protocol it can parse. That is not the same as HTTP-only. NGINX load balances FastCGI, uwsgi, SCGI, memcached and gRPC as well as HTTP. But for opaque TCP or UDP, such as database clusters, SMTP, or anything that must stay encrypted end to end, L4 is the right tool.
Best practice
- Route on the cheapest signal that expresses the rule. Use path and host first, then headers and method. Reach for cookies or query parameters only when the decision genuinely depends on them: they also widen the cache key.
- Order rules most specific first, with the catch-all last. Review that order as a security control, not as a formatting choice.
- Prefer cookie-based affinity, or no affinity at all, over ip-hash. If you must fall back to client IP, treat it as a heuristic and pair it with a cookie.
- Run an origin health check against every pool member. Make routing depend on its result. That way, a request is never sent to a backend already known to be down.
- Terminate TLS once at the edge rather than at each backend. Decide the edge-to-origin leg deliberately. It is a separate connection the balancer opens, not a continuation of the client's. So its protection is a separate choice.
- Keep HTTP keep-alive enabled on the origin. Connection reuse is where the measurable win is. Rebuilding Cloudflare's proxy for better reuse cut median time to first byte by 5 ms, and the 95th percentile by 80 ms.
Examples
# Nginx L7 routing
upstream api_servers {
server 10.0.1.10:8080;
server 10.0.1.11:8080;
}
upstream static_servers {
server 10.0.2.10:80;
server 10.0.2.11:80;
}
server {
listen 443 ssl;
location /api/ {
proxy_pass http://api_servers;
}
location /static/ {
proxy_pass http://static_servers;
add_header Cache-Control "public, max-age=86400";
}
location / {
proxy_pass http://api_servers;
}
}
Frequently Asked Questions
Load balancing at layer 7 of the OSI model, the application layer: the balancer parses each HTTP request and picks a backend by content (path, host, headers, query, cookies) rather than by IP and port. It terminates the client connection and proxies to an origin, so it costs more CPU than L4.
# Nginx L7 routing
upstream api_servers {
server 10.0.1.10:8080;
server 10.0.1.11:8080;
}
upstream static_servers {
server 10.0.2.10:80;
server 10.0.2.11:80;
}
server {
listen 443 ssl;
location /api/ {
proxy_pass http://api_servers;
}
location /static/ {
proxy_pass http://static_servers;
add_header Cache-Control "public, max-age=86400";
}
location / {
proxy_pass http://api_servers;
}
}
Yes. L7 Load Balancing is also known as Layer 7 load balancing, Application load balancing, HTTP load balancing. Load balancing at layer 7 of the OSI model, the application layer: the balancer parses each HTTP request and picks a backend by content (path, host, headers, query, cookies) rather than by IP and port. It terminates the client connection and proxies to an origin, so it costs more CPU than L4.
Related CDN concepts include:
- Edge Server — An edge server is one of the caching reverse-proxy machines inside a CDN Point of …
- L4 Load Balancing — L4 load balancing spreads a service's connections across several servers, choosing each backend from the …
- Rate Limiting — Rate limiting caps how many requests one client may make in a given period, keyed …
- WAF (WAF) — Web application firewall: a reverse proxy that inspects HTTP(S) requests against rule sets and blocks …