Rate Limiting
Rate limiting caps how many requests one client may make in a given period, keyed to an IP address, API key, cookie or path, and refuses the excess, conventionally with HTTP 429. It counts requests without inspecting them, so it sits beside a WAF and bot management rather than replacing them.
Also known as Throttling, Request throttling, Rate limit.
Full Explanation
Rate limiting caps how many requests a single client may make in a given period. It refuses or delays the excess. Cloudflare's Learning Center states the idea plainly: rate limiting "puts a cap on how often someone can repeat an action within a certain timeframe". The operator picks the identity the cap is enforced against. It can be an IP address, an API key, a cookie or session, or a path. HTTP standardises only the refusal: the 429 Too Many Requests status code. RFC 6585, section 4 says outright that "this specification does not define how the origin server identifies the user, nor how it counts requests". So every counting rule you meet is a product or application design decision, not a protocol requirement.
Rate limiting is not authentication. MDN notes that restrictions are "typically" based on "a client's IP". They can be tied to a person only "if requests are authenticated or contain a cookie" (MDN, 429 Too Many Requests). It is not inspection either. A limiter counts requests without judging their content. Cloudflare says of its own product that "it doesn't distinguish between good bots and bad bots". So it belongs beside a WAF and bot management, not in place of them. It is not a complete answer to DDoS either, though it is a genuine layer of one. Cloudflare lists brute force attacks, "DoS and DDoS attacks" and web scraping among the attacks rate limiting "can help mitigate". The same article warns that "rate limiting is not a complete solution for managing bot activity". On a CDN the counter typically lives on the edge server, ahead of your origin. On both Cloudflare and Fastly it is scoped to a single data centre and deliberately approximate. That changes what a threshold can promise. Rate limiting is also called throttling, or request throttling.
How it works
A limiter makes three decisions: what to count by, how to count, and what to do when the count is exceeded. It runs wherever it can see the traffic first. That may be on a CDN edge, in an API gateway, in a reverse proxy such as nginx, or in the application itself.
What to count by is the key. It matters more than the threshold. Cloudflare calls these a rule's characteristics: "The set of parameters that define how Cloudflare tracks the rate for this rule." It keeps "separate counters for each unique combination of values in a rule's characteristics" (Request rate calculation). nginx keys a shared memory zone on any variable, conventionally the binary client address. Too coarse a key groups unrelated users into one bucket. Too fine a key lets an attacker vary the field you keyed on. The key must also be something the client cannot forge. That is a live problem when the limiter sits behind another proxy. MDN observes that then "the server only sees the final proxy's IP address". So the client address has to be read from X-Forwarded-For, and only from a hop you trust.
How to count. Five algorithms cover nearly every deployment. Redis's walkthrough of all five is the reference used here.
- Fixed window: "Divide time into fixed-length windows (e.g. 10-second blocks). Each window gets a counter." It is cheap, one key per client per window, but only approximate. A client "could send 10 requests at second 9 and another 10 at second 11 — 20 requests in 2 seconds while technically staying within a “10 per 10 seconds” limit, because the requests straddle two windows". Redis's comparison table rates its burst behaviour "Allows 2x burst at boundaries".
- Sliding window log: store every request timestamp and prune the ones that fall out of the window. This gives exact counts. "It eliminates the boundary-burst problem of fixed windows at the cost of higher memory usage". Memory grows with the number of requests in the window, not with the number of clients alone.
- Sliding window counter: two counters, the current window and the previous one, blended by how far into the window you are. This "smooths the boundary spike that plagues fixed windows" for the price of two keys. It is still not exact. Redis lists as a drawback that "The count is an approximation." Redis recommends it as the default if you are unsure.
- Token bucket: "Tokens refill at a steady rate (e.g. 1 per second). Each request removes one token. If the bucket is empty, the request is denied." Capacity is capped, so it "allows short bursts (up to the bucket capacity) while enforcing a steady average rate". AWS names the two knobs directly. The throttling rate is "the rate, in requests per second, that tokens are added to the token bucket". The throttling burst is "the capacity of the token bucket".
- Leaky bucket: a bucket that drains at a fixed rate, in one of two modes. Policing rejects overflow immediately. Shaping queues accepted requests and delays them into a smooth output rate. nginx's limit_req does both in sequence. "The limitation is done using the “leaky bucket” method", and "Excessive requests are delayed until their number exceeds the maximum burst size in which case the request is terminated with an error" (ngx_http_limit_req_module).
What to do at the limit. The conventional refusal is 429. It "indicates that the user has sent too many requests in a given amount of time" (RFC 6585, section 4). That response "SHOULD include details explaining the condition, and MAY include a Retry-After header indicating how long to wait before making a new request". Note: MAY, not MUST. Retry-After itself is defined in RFC 9110, section 10.2.3. Servers send it "to indicate how long the user agent ought to wait before making a follow-up request", as either an HTTP-date or a number of seconds. Two further rules from RFC 6585 are routinely missed. First, "Responses with the 429 status code MUST NOT be stored by a cache". Second, 429 is optional too. section 7.2 notes that under attack, "responding to each with a 429 status code will consume resources". So "servers are not required to use the 429 status code; when limiting resource usage, it may be more appropriate to just drop connections, or take other steps". Blocking is not the only action available either. A Cloudflare rate limiting rule can issue a challenge, or merely log, instead of blocking.
Telling clients their budget is still unstandardised. The familiar X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset headers are convention only. Redis, having suggested them, adds: "The headers above are only common patterns. There is no single standard." The IETF HTTPAPI working group is drafting RateLimit-Policy and RateLimit fields for precisely this purpose. That work is still an Internet-Draft. Its introduction says "Currently, there is no standard way for servers to communicate quotas so that clients can throttle their requests to prevent errors" (draft-ietf-httpapi-ratelimit-headers, work in progress, not a published RFC). Until it lands, assume clients understand the status code and Retry-After, and nothing else.
Why it matters for a CDN
The reason to rate limit on a CDN, rather than only in the application, is position. The edge sees the request first. Traffic refused there costs the origin nothing. Fastly names the two purposes: "to prevent abusive use of a website or service (e.g. by a scraping bot or a denial of service attack) or to apply a limit on use of an expensive or billable resource (e.g. to allow up to 1000 requests an hour to an API endpoint)". The second purpose is not security at all. It is a bill. Cloudflare makes the same point about APIs: "Every time an API responds to a request, the owner of that API has to pay for compute time".
A CDN's geometry changes what a limit means. This is where most misconfiguration lives. Counters are local to each point of presence. "Cloudflare does not support global rate limiting counters across the entire network. Each data center maintains its own counters". The exception is data centres associated with the same geographical location. Every rule silently carries the data centre id (cf.colo.id) as a mandatory characteristic: "This ensures counters remain scoped to each data center." Fastly's counts are per POP too. Its buckets hold "the estimated number of requests received up to and including that 10 second window of time across the entire Fastly POP". So a threshold of 100 requests per minute is 100 per data centre per minute. A client spread across twenty PoPs can spend twenty times what you configured. A botnet spread across continents sees a much higher effective limit than a single scraper does.
Enforcement also lags detection. Cloudflare is explicit that rate limiting rules "are not designed to allow a precise number of requests to reach your origin server. There may be a delay of up to a few seconds between detecting a request and updating rate counters. Due to this delay, excess requests could still reach the origin before Cloudflare enforces a mitigation action such as blocking or challenging." Fastly warns in the same spirit that "Rate counters are designed to quickly count high volumes of traffic, not to count precisely". Fastly explains the trade: its "rate counter is designed as an anti-abuse mechanism", where reacting quickly beats counting exactly. Resource limiting, by contrast, "often requires a globally synchronized count and must be precise". For that, Fastly points customers at real-time log streaming and post-processing rather than at edge counters. If a number has to be exact, a contractual API quota or a paid tier, the edge is the wrong place to enforce it.
One more CDN-specific wrinkle: a cache hit is still a request. On Cloudflare, rate limiting counts cached assets by default. Restricting the count to origin-bound traffic is a plan-dependent parameter. With it disabled, "only the requests going to the origin (that is, requests that are not cached) will be considered when determining the request rate". In the other direction, a counting expression that references response fields forces the request to the origin, "skipping any cached content". So a careless counting rule can cost you the cache it was meant to protect.
What CDNs do
- Cloudflare: a rate limiting rule combines an expression, counting characteristics, requests per period, a period, an action, and a duration. "Once the rate is reached, the rate limiting rule blocks further requests for the period of time defined in this field." The Block action's response code defaults to 429 and must be a value between 400 and 499. Capability is plan-scoped, not universal. The Free plan gets one rule, IP as its only counting characteristic, and a 10-second counting period. "IP with NAT support" and custom counting expressions start at Business. The wider set of characteristics (host, query, path, header, cookie, ASN, country, JA3/JA4, body) needs Enterprise with Advanced Rate Limiting. So does complexity-based limiting, which counts a cost score the origin returns in a response header, between 1 and 1,000,000, instead of counting requests. Only Enterprise customers can throttle requests above the rate with the Block action, instead of blocking for a fixed duration (rate limiting rules, parameters).
- Fastly: Edge Rate Limiting exposes two primitives to VCL and Compute, rate counters and penalty boxes. Accumulated counts "are converted to an estimated rate computed over one of three time windows: 1s, 10s or 60s", always in requests per second. check_rates() increments two counters in one call, so a sustained rate and a burst rate can be enforced together. Your own code returns the refusal. The documented VCL example answers with a 429 and puts the client in a penalty box for 15 minutes. The product "must be enabled on your account by a Fastly employee in order to use the primitives in VCL described on this page" (Fastly rate limiting).
- AWS: for REST APIs, API Gateway "throttles requests to your API using the token bucket algorithm, where a token counts for a request". Four settings apply in order: per-client or per-method limits from a usage plan, per-method limits on a stage, account limits per Region, then AWS Regional limits that customers cannot change. None of them is a hard ceiling. "Both throttles and quotas are applied on a best-effort basis and should be thought of as targets rather than guaranteed request ceilings". "When request submissions exceed the steady-state request rate and burst limits, API Gateway begins to throttle requests. Clients may receive 429 Too Many Requests error responses at this point" (Throttle requests to your REST APIs).
Watch out for
- The threshold is per data centre, not per network. Size it for one PoP. Never quote the configured number as a global guarantee. This bites hardest when IP is not one of the characteristics. Cloudflare flags that as "especially relevant for customers that do not add the IP address as one of the rate limiting characteristics".
- Boundary bursts. A fixed window admits about twice the limit when requests land either side of a window edge. Use a sliding variant, or a bucket, where accuracy at the boundary matters.
- A window is a rate, not a schedule. Fastly: "If a rate of 100rps is selected with a 60 second window, that will equally allow 100rps for sixty seconds or 6000rps for one second." A long window with a generous rate does not stop a short spike. That is what a second, shorter window is for.
- Shared and NAT addresses. One address can be thousands of people. Cloudflare offers "IP with NAT support" specifically "to handle situations such as requests under NAT sharing the same IP address". But it works by cookie (_cfuvid). Cloudflare warns: "Visitors who clear cookies, use private browsing, or do not accept cookies will not be individually identified. Requests from these visitors share a single counter bucket, which can cause false positives in high-traffic NAT environments." Cloudflare's own advice for login and payment endpoints is to combine it with another characteristic, such as path or a header.
- A key the client controls is not a key. MDN is emphatic about any security use of X-Forwarded-For, "such as for rate limiting or IP-based access control". Such use "must only use IP addresses added by a trusted proxy", because untrustworthy values "can result in rate-limiter avoidance, access-control bypass, memory exhaustion, or other negative security or availability consequences".
- Per-IP limits do not stop distributed abuse. Cloudflare: "If rate limiting is only applied by IP address, brute force attackers could bypass this by attempting logins from multiple IP addresses (perhaps by using a botnet)."
- It cannot tell a crowd from an attack. Rate limiting judges volume, not intent. RFC 4732, section 1 records the underlying limit: "in principle it is not possible to distinguish between a sufficiently subtle DoS attack and a flash crowd (where unexpected heavy but non-malicious traffic has the same effect as a DoS attack)". A launch, a sale or a legitimate batch client will trip a tight rule.
- nginx refuses with 503 by default. The limit_req_status directive defaults to 503, so returning 429 requires setting it explicitly. Until you do, your dashboards will read origin unavailability where the truth is a rate limit.
- Limiter state is finite. An nginx zone stores "64 bytes on 32-bit platforms and 128 bytes on 64-bit platforms" per key. That is roughly 16 thousand or 8 thousand states per megabyte. If the zone fills, "the least recently used state is removed". If a state still cannot be created, "the request is terminated with an error". An undersized zone starts refusing legitimate traffic.
- Never cache the refusal. RFC 6585 is unconditional: "Responses with the 429 status code MUST NOT be stored by a cache." A 429 stored under a shared cache key would be replayed to unrelated clients. That is the opposite of the intent. See negative caching for how error responses are normally handled.
Best practice
- Choose the algorithm from the shape of your traffic. Token bucket suits short, legitimate bursts. Leaky-bucket policing suits a downstream that cannot absorb bursts at all. Sliding window counter is a general default. Fixed window works only where boundary slop is acceptable. Sliding window log works where counts must be exact and you can pay the memory.
- Enforce two limits rather than one: a sustained rate and a short burst rate. Fastly provides check_rates() to increment two counters in a single call for exactly this. With a token bucket, the same shape comes from the refill rate plus the bucket capacity.
- Key on something forgery-resistant. Take the client address only from a proxy hop you control.
- Layer identities on authentication paths. On login pages, a rule can key on IP or on username. Cloudflare: "Ideally it would use a combination of the two". That is because IP-only rules lose to a botnet, while username-only rules let one address try common passwords across many accounts.
- Refuse with 429 plus a Retry-After value. Document the policy out of band. Do not depend on clients reading X-RateLimit-style headers. They are convention, not standard.
- Make your own clients back off instead of hammering. AWS recommends "an exponential backoff algorithm to calculate the sleep interval between API requests", with "a maximum delay interval, as well as a maximum number of retries", and "jitter (randomized delay) to prevent successive collisions" (Request throttling for the Amazon EC2 API). A tight retry loop against a rate limit deepens the incident it is reacting to.
- Deploy in observation mode first. Cloudflare rules support a log action. nginx has limit_req_dry_run, in which "requests processing rate is not limited, however, in the shared memory zone, the number of excessive requests is accounted as usual". Fastly exposes ratecounter_increment separately from penaltybox_add, so you can count traffic before you start penalising it. Tune the threshold against measured traffic, with headroom, before switching to block.
- Do not enforce a number that has to be exact at the edge. Put contractual quotas where they can be counted precisely. Fastly's own recommendation for resource rate limiting is real-time log streaming with post-processing in your logging provider.
- Keep it as one control among several. Rate limiting answers "how much", not "what" or "who". Pair it with a WAF, bot management and always-on DDoS mitigation.
Interactive Animation
Examples
# Nginx: rate limiting
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s;
limit_req_zone $binary_remote_addr zone=login:10m rate=1r/s;
server {
location /api/ {
limit_req zone=api burst=20 nodelay;
limit_req_status 429;
}
location /login {
limit_req zone=login burst=5;
limit_req_status 429;
}
}
# Cloudflare rate limiting rule (API)
curl -X POST "https://api.cloudflare.com/client/v4/zones/{zone}/ratelimits" \
-d '{
"match": {"request": {"url": "*.example.com/api/*"}},
"threshold": 100,
"period": 60,
"action": {"mode": "simulate"}
}'
Frequently Asked Questions
Rate limiting caps how many requests one client may make in a given period, keyed to an IP address, API key, cookie or path, and refuses the excess, conventionally with HTTP 429. It counts requests without inspecting them, so it sits beside a WAF and bot management rather than replacing them.
# Nginx: rate limiting
limit_req_zone $binary_remote_addr zone=api:10m rate=10r/s;
limit_req_zone $binary_remote_addr zone=login:10m rate=1r/s;
server {
location /api/ {
limit_req zone=api burst=20 nodelay;
limit_req_status 429;
}
location /login {
limit_req zone=login burst=5;
limit_req_status 429;
}
}
# Cloudflare rate limiting rule (API)
curl -X POST "https://api.cloudflare.com/client/v4/zones/{zone}/ratelimits" \
-d '{
"match": {"request": {"url": "*.example.com/api/*"}},
"threshold": 100,
"period": 60,
"action": {"mode": "simulate"}
}'
Yes. Rate Limiting is also known as Throttling, Request throttling, Rate limit. Rate limiting caps how many requests one client may make in a given period, keyed to an IP address, API key, cookie or path, and refuses the excess, conventionally with HTTP 429. It counts requests without inspecting them, so it sits beside a WAF and bot management rather than replacing them.
Related CDN concepts include:
- DDoS (Distributed Denial of Service) (DDoS) — A DDoS attack makes a site unavailable by aiming more traffic or work at it …