HAR (HTTP Archive)
A HAR (HTTP Archive) file is a JSON record of every HTTP request and response a browser made during a page load, with status, headers, sizes and per-phase timings. It is the de-facto way to share a network trace: not a published standard, and only ever the client's view.
Also known as HTTP Archive.
Full Explanation
A HAR (HTTP Archive) file is a JSON record of every HTTP request and response a browser made while a page loaded. It records status, headers, body sizes, and a per-phase timing breakdown for each one. It is the de-facto way to hand a network trace to somebody else. That is why support engineers on both sides of a CDN ticket ask for it first ("send me the HAR"). It is not a standard. The HAR 1.2 specification is published as a Historical Draft dated 14 August 2012. It carries a DO NOT USE notice. It states that it "was never published by the W3C Web Performance Working Group and has been abandoned". It is also not a server-side record. A HAR holds what one browser on one network observed. So it can prove what a visitor received. It never proves what the origin or the CDN did internally.
How it works
The format is JSON. A HAR file must be saved in UTF-8; other encodings are forbidden. (The specification's own JSON reference is RFC 4627. The IETF has since replaced RFC 4627 twice, most recently with RFC 8259.) The root member must be present. It must be named log. It holds the format version, a required creator object naming the tool that wrote the file, an optional browser object, an optional pages array, and a required entries array. There is one page object for every exported web page. There is one entry object for every HTTP request.
Each entry carries four objects. The request object holds method, URL, HTTP version, cookies, headers, query string, any POST data, and header and body sizes. The response object holds status, HTTP version, headers, the redirectURL taken from the Location response header, and a content object. The other two objects are cache and timings. Three optional fields were added in 1.2, and they earn their keep. serverIPAddress is the IP address of the server that was connected, as resolved by DNS. connection is a unique ID for the parent TCP connection; it is how you spot connection reuse. The third field is comment. The httpVersion fields record the protocol version of each request and response. So a HAR also answers whether HTTP/2 or HTTP/3 was actually negotiated rather than assumed.
Time is split into phases. The timings object has seven numeric fields, all in milliseconds. blocked is time spent in a queue waiting for a network connection. dns is the time required to resolve a host name. connect is the time required to create the TCP connection. ssl is the time required for SSL/TLS negotiation. When this field is present, its time is also counted inside connect, for backward compatibility with HAR 1.1. send is the time spent sending the request. wait is the time spent waiting for a response from the server. receive is the time spent reading the entire response from the server or cache. send, wait and receive are not optional and must be non-negative. blocked, dns, connect and ssl may be omitted altogether by a tool that cannot supply them, or set to -1 where they do not apply. The specification's own example is that connect would be -1 for a request that re-uses an existing connection. An entry's time must equal the sum of the timings supplied, excluding any -1 values.
One phase is routinely misread. Chrome DevTools labels wait "Waiting (TTFB)". Chrome's own reference scopes that number precisely: it "includes 1 round trip of latency and the time the server took to prepare the response". So wait is not the whole time to first byte. TTFB is measured from the start of the request. So it also contains blocked, dns, connect (with ssl inside it) and send. Read the phases individually rather than the label. A small wait on a request that re-used a warm connection says nothing about what a first-time visitor pays for DNS and the handshake.
Sizes are recorded twice over. The two numbers mean different things. response.bodySize is the number of body bytes received on the wire. It is set to zero for responses coming from the cache (304) and to -1 when the information is not available. content.size is the length of the returned content. It "should be equal to response.bodySize if there is no compression and bigger when the content has been compressed". The optional content.compression field gives the number of bytes saved. The cache object is about the browser. The specification describes it as info about a request coming from browser cache, holding the cache entry state beforeRequest and afterRequest. It says nothing about any CDN cache. Finally, any field a tool invents outside the specification must start with an underscore, so parsers can ignore what they do not recognise.
Why it matters for a CDN
A HAR is the fastest way to judge whether a CDN did its job for one specific user, on one network, at one moment. Four readings do most of the work. Cache outcome lives in each response's headers. Look for a vendor marker such as CF-Cache-Status, a generic X-Cache, and the Age header. The Age header's presence indicates the object came from a cache, and its value says how long it had been there. Effect of the hit is the wait phase compared across a hit and a miss for the same URL. That comparison turns the payoff of a high cache hit ratio into per-request milliseconds instead of a dashboard percentage. Delivery hygiene comes from the rest of the entry. content.size against response.bodySize, together with Content-Encoding, shows which resources arrived without compression. redirectURL chains expose round trips spent before the real response. blocked exposes queueing. Chrome documents this as happening when six TCP connections are already open to an origin. That limit applies to HTTP/1.0 and HTTP/1.1 only, so the same page over HTTP/2 does not accumulate it. Which machine answered is serverIPAddress, the address of the edge server that served the request. You correlate it against the CDN's own PoP identifier when a fault looks regional. Because the capture happens in the browser, a HAR is evidence about delivery, not about CDN internals. It complements origin logs, CDN logs and CDN-side metrics, and it never replaces them.
What CDNs do
A HAR is captured by the browser. What differs between vendors is how their support process asks for it, and which diagnostic headers their responses put into the file.
- Cloudflare (support guide): the guide calls a HAR "often the first thing Cloudflare Support will request". It gives capture steps for Chrome, Firefox, Edge, Safari and mobile. It warns that "a HAR file can include sensitive details such as passwords, payment information, and private keys", and it points customers at a HAR sanitizer. On the response side (cache responses, on by default) CF-Cache-Status "indicates whether a resource is cached or not". Its statuses include HIT ("the resource was found in Cloudflare's cache"), MISS, EXPIRED, STALE, REVALIDATED, UPDATING, BYPASS and DYNAMIC. DYNAMIC means Cloudflare decided at request time that the asset was not eligible for cache at all. Age is "only present for responses served from the cache".
- Fastly (Fastly-Debug, opt-in): if an inbound request has the Fastly-Debug header set, "conventionally to 1 but actually any value is acceptable", Fastly cache servers add Fastly-Debug-Path, Fastly-Debug-TTL and Fastly-Debug-Digest. The digest is "a hash of the cache key created in the vcl_hash subroutine". The same header also stops Surrogate-Key and Surrogate-Control being removed from the response. A service can unset Fastly-Debug on the way in, so its absence from a HAR is not evidence of anything.
- Akamai (Pragma headers, opt-in): Akamai documents Pragma request headers that map to diagnostic response headers. akamai-x-cache-on returns X-Cache, which "returns information about how the edge server response was served", with example values including TCP_HIT, TCP_MISS, TCP_REFRESH_HIT and TCP_MEM_HIT. akamai-x-check-cacheable returns X-Check-Cacheable, "a flag to indicate whether the response is cacheable". akamai-x-get-true-cache-key returns X-True-Cache-Key, "the true cache key used for the MD5 hashing", excluding the scheme and the request method. Akamai documents these for its Edge Diagnostics tooling. So whether a property returns them to an ordinary browser request is a configuration matter. Do not read their absence from a HAR as a verdict.
- AWS (support case guide): AWS Support's own article walks customers through capturing a HAR for a case. In Chrome and Edge, use the "Export HAR (sanitized)..." icon. In Firefox, use "Save All As HAR". The article then walks through editing the file, because "HAR files and console logs can capture sensitive information, such as usernames, passwords, and keys". Cookies and authentication headers must be removed or masked before the file is attached.
Watch out for
- It is sensitive by construction. A HAR carries cookies, POST bodies, query strings and headers such as Authorization. The specification only says the user agent should find some way to notify the user before it transfers the file to anyone else. That is advice to the tool, not a safeguard in the file. Chrome's default export is the sanitized one, which "excludes sensitive information such as Cookie, Set-Cookie, and Authorization headers". A HAR exported with sensitive data, or written by some other tool, has no such protection.
- The cache object is the browser's cache. A CDN hit or miss is visible only in the response headers. Nothing in the HAR cache object refers to an edge cache.
- Missing phases are not zeros. connect is -1 on a re-used connection. An exporter that cannot measure blocked, dns, connect or ssl may leave them out entirely. A totals-only reading of such an entry is worthless. Only send, wait and receive are guaranteed present.
- Bodies and byte counts are not guaranteed. content.text is optional and is left out when unavailable, so many HARs hold headers and sizes but no response bodies. Cloudflare's mobile Chrome steps use "Save all as HAR with Content" precisely to get them. Neither size field is a plain wire count. bodySize is zero for a response coming from the cache (304). content.size is the decoded length, and it is larger than bodySize whenever the body was compressed.
- The capture is not the visitor's experience. A HAR taken in a private window with Disable cache ticked deliberately describes a cold client cache. That is the opposite of a returning visitor. Note which conditions you captured under. Pair the file with origin and CDN logs before concluding anything server-side.
- The format is frozen, so tools differ. The 1.2 draft was abandoned rather than finished. So almost everything outside log, version, creator and entries is optional. Field coverage and fidelity vary between exporters. Anything beginning with an underscore is one tool's private extension.
Best practice
- Capture in an Incognito or private window with Preserve log checked. Tick Disable cache when the question is about caching. That is exactly what Cloudflare asks for on a cache ticket.
- Reproduce the fault inside the recording. Each entry carries a startedDateTime. That is what lets a CDN line your capture up against its own logs, but only for requests that made it into the file.
- Add the vendor's debug headers before capturing when you need cache-key or TTL detail: Fastly-Debug on the request for Fastly, the documented Pragma headers for Akamai. Retrofitting them afterwards means capturing again.
- Export "with content" only when you genuinely need response bodies. It makes the file far larger and far more sensitive. Headers, sizes and timings are present either way.
- Sanitize before sharing: prefer the browser's sanitized export, then still search the file for credentials and tokens. Mask cookies and authentication headers as AWS Support instructs.
- Judge a CDN cache outcome from the response headers, never from the HAR cache object. Compare hit against miss for the same URL. Treat the result as one client's evidence, to be confirmed against origin and CDN logs.
Examples
# Export HAR from Chrome DevTools:
# 1. Open DevTools (F12) > Network tab
# 2. Load the page
# 3. Right-click > Save all as HAR with content
# Export HAR from curl (not standard, but useful)
$ curl -w '{"timings": {"dns": %{time_namelookup}, "connect": %{time_connect}, "tls": %{time_appconnect}, "ttfb": %{time_starttransfer}, "total": %{time_total}}}' -o /dev/null -s https://cdn.example.com/
# Analyze HAR with har-analyzer (npm)
$ npx har-analyzer report.har
# Key fields to check in HAR:
# response.headers: X-Cache, Age, CF-Cache-Status
# timings.wait: server processing time (TTFB)
# timings.blocked: time waiting in browser queue
Frequently Asked Questions
A HAR (HTTP Archive) file is a JSON record of every HTTP request and response a browser made during a page load, with status, headers, sizes and per-phase timings. It is the de-facto way to share a network trace: not a published standard, and only ever the client's view.
# Export HAR from Chrome DevTools:
# 1. Open DevTools (F12) > Network tab
# 2. Load the page
# 3. Right-click > Save all as HAR with content
# Export HAR from curl (not standard, but useful)
$ curl -w '{"timings": {"dns": %{time_namelookup}, "connect": %{time_connect}, "tls": %{time_appconnect}, "ttfb": %{time_starttransfer}, "total": %{time_total}}}' -o /dev/null -s https://cdn.example.com/
# Analyze HAR with har-analyzer (npm)
$ npx har-analyzer report.har
# Key fields to check in HAR:
# response.headers: X-Cache, Age, CF-Cache-Status
# timings.wait: server processing time (TTFB)
# timings.blocked: time waiting in browser queue
Yes. HAR (HTTP Archive) is also known as HTTP Archive. A HAR (HTTP Archive) file is a JSON record of every HTTP request and response a browser made during a page load, with status, headers, sizes and per-phase timings. It is the de-facto way to share a network trace: not a published standard, and only ever the client's view.
Related CDN concepts include:
- Cache Hit Ratio (CHR) — The share of requests a cache answers from its own stored copies instead of fetching …
- Latency — Latency is the time data takes to travel from one point on a network to …
- RTT (Round-Trip Time) (RTT) — RTT (round-trip time) is the delay from sending a packet until the response it triggers …
- TTFB (Time To First Byte) (TTFB) — TTFB is the time from the start of a request until the first byte of …