P95/P99 Percentiles
P95 and P99 are latency percentiles: the times 95% and 99% of requests beat, so the slowest 5% and 1% exceed them. They are order statistics read off the sorted list of response times, not an average, a rate, or a maximum. They describe the slow tail an average hides.
Full Explanation
P95 and P99 are latency percentiles. P95 is the response time that 95% of requests come in under. P99 is the response time that 99% of requests come in under. So the slowest 5% of requests exceed P95, and the slowest 1% exceed P99. A latency percentile marks a position in the sorted distribution of response times. That makes P95 and P99 order statistics: real observed values, not a computed summary. They are not an average, not a rate, and not throughput. They are also not a maximum and not a guarantee. By construction, something is always slower than P99. P50, the median, is the same kind of statistic at the halfway mark.
The reason they exist is that latency distributions are right-skewed. So the mean sits nowhere near what any real request experienced. ClickHouse's worked example makes the point. Of 1,000 requests, 988 complete in 52 ms and 12 stall for 2,400 ms behind a lock. The mean is 80 ms, a figure that "matches no request that actually happened". P50 is 52 ms and P99 is 2,400 ms. Quote a percentile and you are naming a time some identifiable share of your users actually waited. That is why the Google SRE book recommends percentiles over averages for latency indicators. Averaging request latencies "obscures an important detail: it's entirely possible for most of the requests to be fast, but for a long tail of requests to be much, much slower."
How it works
Sort the measured request times and read off a position. In 1,000 samples, P99 is roughly the 990th value and P95 is roughly the 950th. The mental model really is a sorted list. Each percentile reports one position, so a set of them (P50, P90, P95, P99) sketches the shape of the whole distribution. No single number can do that.
The same example shows why the pair is quoted rather than either one alone. With 12 slow requests in 1,000, the tail is 1.2% wide. So P95 lands at 52 ms, in the fast group, and reports nothing wrong at all. P99, at 2,400 ms, is the statistic that sees those 12 requests. P95 tells you where the bulk of traffic ends. P99 tells you how bad the tail gets. A tail narrower than 1% is invisible to both.
An exact percentile requires keeping every observation. So telemetry systems approximate. The two dominant approaches are bucketed histograms (Prometheus classic histograms, OpenTelemetry explicit-bucket histograms) and quantile sketches (HdrHistogram, t-digest, DDSketch). Both are compact and mergeable across hosts and time windows. Both trade a bounded approximation error at read time. Prometheus computes a quantile with histogram_quantile(). This walks the cumulative bucket counts to the requested quantile and interpolates inside the bucket the quantile falls into. Classic histograms use linear interpolation, but native histograms on a standard exponential schema use exponential interpolation. So the method is not uniform across histogram types. Accuracy is bounded by bucket width. Prometheus documents a specific edge: if a quantile lands in the highest bucket of a classic histogram, the function returns the upper bound of the second highest bucket.
Two further properties surprise people. First, there is no single accepted definition of a sample quantile. Hyndman and Fan catalogued that "there are a large number of different definitions used for sample quantiles in statistical computer packages". Even within one package, a different definition may be used for a boxplot than for an explicit quantile call. So two tools can disagree slightly on the same data. Second, percentiles cannot be averaged. Computing one needs the full sorted distribution. A finished percentile value cannot be recombined. Take two hosts. Host A serves 1,000 requests all at 100 ms, so its P99 is 100 ms. Host B serves 900 at 100 ms and 100 at 1,000 ms, so its P99 is 1,000 ms. Averaging the two per-host values gives 550 ms, but the true P99 across all 2,000 requests is 1,000 ms. Request-count weighting does not fix it. The information was destroyed when each host reduced its distribution to one number. The correct method is to merge the distributions, or mergeable sketches of them, and then take the percentile.
Why it matters for a CDN
CDN traffic is a mixture of structurally different request paths. That is exactly what produces a long right tail. A near-edge cache hit, a cache miss that must fetch from the origin, a request from a user with no nearby PoP, and a connection paying a fresh TLS handshake are not samples from one population. An average over that mixture describes none of them. P95 and P99 land squarely in the expensive paths. That is where a CDN's own failure modes live: cold caches, origin slowness, and geographic coverage gaps.
The cost of a miss is concrete and measurable. Cloudflare's origin response time metric starts its clock "when Cloudflare decides the request must go to origin (a cache miss)" and stops when it receives the response headers. It includes DNS resolution, TCP and TLS handshakes, request transmission, origin processing and response receipt. That whole round trip is what your TTFB tail is made of on a miss.
Percentiles are also the natural shape of a latency objective. The SRE book's canonical form is "99% of Get RPC calls will complete in less than 100 ms": a percentile and a threshold, not an average. It notes you can specify several targets at once when the shape of the curve matters. A high-order percentile "shows you a plausible worst-case value, while using the 50th percentile (also known as the median) emphasizes the typical case."
Finally, fan-out amplifies a rare slow component into a common slow experience. Dean and Barroso's "The Tail at Scale" shows that "variability in the latency distribution of individual components is magnified at the service level". Consider servers that typically respond in 10 ms but have a 99th-percentile latency of one second: "if a user request must collect responses from 100 such servers in parallel, then 63% of user requests will take more than one second." Anything that assembles a response from many fetches inherits this. The same paper delivers a warning aimed straight at CDN intuition. Caching layers, however useful, "do not directly address tail latency", except where the entire working set is guaranteed to fit in cache. A good cache hit ratio improves the typical case without necessarily touching P99.
What CDNs do
- Cloudflare: Origin Analytics reports origin response time "measured at the 50th, 95th, and 99th percentiles". It draws a reference line at the zone's configured origin timeout, so you can catch requests approaching it before they become 524 errors. It ranks request paths by P95 response time, error rate, request volume or TCP failure rate.
- AWS CloudFront: exposes per-request logs rather than ready-made percentiles. You compute them yourself in Athena with approx_percentile, which "returns the approximate percentile for all input values of x at the given percentage". Athena engine version 3 is built on Trino. AWS documents that approx_percentile "now uses tdigest instead of qdigest", warning that the function "returns different results than it did in previous engine versions".
- Prometheus and OpenTelemetry: the common way edge and origin fleets expose latency. Those fleets emit histograms, then derive percentiles at query time with histogram_quantile(). This keeps the buckets mergeable across instances.
- Tail-tolerant serving: P95 is used as a control input, not just a dashboard number. Dean and Barroso defer a hedged second request until the first has been outstanding longer than "the 95th-percentile expected latency for this class of requests". This "limits the additional load to approximately 5%". In one Google benchmark, hedging after 10 ms cut 99.9th-percentile latency from 1,800 ms to 74 ms for 2% more requests.
Watch out for
- Averaging percentiles. Rolling per-PoP or per-minute P99s into a mean produces a number that is the P99 of nothing. In the two-host example, that number is 550 ms where the truth is 1,000 ms. Merge distributions or sketches, then take the percentile.
- Trusting P95 alone. A tail narrower than 5% hides beneath it. In the worked example, P95 reads 52 ms while 12 users in 1,000 waited 2.4 seconds. Conversely, at low traffic P99 "jumps with every slow request". Below a few hundred requests per window, it is only a handful of samples. So prefer P95 or a longer evaluation window there.
- Coordinated omission. Named by Gil Tene, this is the pitfall no sketch or histogram fixes. Load generators "that wait for each response before sending the next stop sampling during stalls, silently deleting the worst observations before any percentile is computed." Your measurement tool can erase the tail before you ever compute P99.
- Where you measure changes the number. Cloudflare notes that its metric covers the full upstream round trip. Because of that, it "shows higher response times than your origin's own monitoring tools". Those tools measure only server-side processing. Neither is wrong. They measure different spans. And a CDN-side percentile is still not the user's experience. Client-side latency "is often the more user-relevant metric, but it might only be possible to measure latency at the server."
- Histogram resolution and tool disagreement. A percentile read from coarse buckets can sit anywhere inside the bucket. A quantile in the top classic-histogram bucket returns the bound below it. Add the differing sample-quantile definitions across packages, and small discrepancies between tools are expected, not bugs.
- A bare figure from a skewed metric. P99 over a global mixture is dominated by whichever segment is worst. A healthy global P99 can conceal one region or one ISP being badly served.
Best practice
- Report P50 alongside P95 and P99. The median gives the typical case and capacity trends. The high percentiles give the plausible worst case. One number alone hides the shape.
- Store mergeable histograms or sketches, never finished percentile values. That way, any later grouping by PoP, region or time window still yields a correct percentile.
- Express latency objectives as a percentile plus a threshold. Use several targets when the shape of the curve matters, rather than one blunt promise.
- Segment by region, PoP, ISP, device and cache status. The tail is usually a specific population. Aggregate percentiles are what hide it.
- Measure client-side with RUM as well as at the edge. Server-side timing cannot see what the browser waited for.
- Alert on the percentile that maps to your objective. Widen the window or drop to P95 on low-traffic endpoints, instead of paging on three slow requests.
- Keep average latency as a capacity and cost metric: it aggregates exactly, which percentiles do not. But never use it as the user-experience indicator.
Examples
# Calculate P95 and P99 from Nginx access log
# Assuming $request_time is the last field
awk '{print $NF}' /var/log/nginx/access.log | \
sort -n | \
awk 'BEGIN{n=0} {a[n++]=$1} END{
printf "P50: %.3f\n", a[int(n*0.50)];
printf "P95: %.3f\n", a[int(n*0.95)];
printf "P99: %.3f\n", a[int(n*0.99)];
printf "Max: %.3f\n", a[n-1];
}'
# Prometheus query for P99 latency
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)
# CloudFront real-time logs P95 analysis
aws athena start-query-execution --query-string "
SELECT
approx_percentile(time_to_first_byte, 0.50) as p50_ttfb,
approx_percentile(time_to_first_byte, 0.95) as p95_ttfb,
approx_percentile(time_to_first_byte, 0.99) as p99_ttfb
FROM cloudfront_logs
WHERE date = '2026-03-15'
"
# Python: calculate percentiles
import numpy as np
latencies = [12, 15, 18, 22, 25, 30, 45, 80, 150, 800]
print(f"P50: {np.percentile(latencies, 50):.0f}ms") # 27ms
print(f"P95: {np.percentile(latencies, 95):.0f}ms") # 507ms
print(f"P99: {np.percentile(latencies, 99):.0f}ms") # 735ms
print(f"Avg: {np.mean(latencies):.0f}ms") # 119ms
# Average says 119ms, but P95 shows the real story
Frequently Asked Questions
P95 and P99 are latency percentiles: the times 95% and 99% of requests beat, so the slowest 5% and 1% exceed them. They are order statistics read off the sorted list of response times, not an average, a rate, or a maximum. They describe the slow tail an average hides.
# Calculate P95 and P99 from Nginx access log
# Assuming $request_time is the last field
awk '{print $NF}' /var/log/nginx/access.log | \
sort -n | \
awk 'BEGIN{n=0} {a[n++]=$1} END{
printf "P50: %.3f\n", a[int(n*0.50)];
printf "P95: %.3f\n", a[int(n*0.95)];
printf "P99: %.3f\n", a[int(n*0.99)];
printf "Max: %.3f\n", a[n-1];
}'
# Prometheus query for P99 latency
histogram_quantile(0.99,
sum(rate(http_request_duration_seconds_bucket[5m])) by (le)
)
# CloudFront real-time logs P95 analysis
aws athena start-query-execution --query-string "
SELECT
approx_percentile(time_to_first_byte, 0.50) as p50_ttfb,
approx_percentile(time_to_first_byte, 0.95) as p95_ttfb,
approx_percentile(time_to_first_byte, 0.99) as p99_ttfb
FROM cloudfront_logs
WHERE date = '2026-03-15'
"
# Python: calculate percentiles
import numpy as np
latencies = [12, 15, 18, 22, 25, 30, 45, 80, 150, 800]
print(f"P50: {np.percentile(latencies, 50):.0f}ms") # 27ms
print(f"P95: {np.percentile(latencies, 95):.0f}ms") # 507ms
print(f"P99: {np.percentile(latencies, 99):.0f}ms") # 735ms
print(f"Avg: {np.mean(latencies):.0f}ms") # 119ms
# Average says 119ms, but P95 shows the real story
Related CDN concepts include:
- Latency — Latency is the time data takes to travel from one point on a network to …
- Throughput — Throughput is the rate at which data actually crosses a link in a measured interval …
- TTFB (Time To First Byte) (TTFB) — TTFB is the time from the start of a request until the first byte of …