Gzip

Compression

Gzip is the HTTP content coding that wraps a DEFLATE stream (LZ77 plus Huffman coding) in a header and a CRC-32 trailer, defined by RFC 1952. RFC 1951 puts the gain at a factor of 2.5 to 3 for English text. It is HTTP's universal fallback coding; Brotli and Zstandard compress tighter.

Also known as x-gzip.

12 min read Updated Aug 30, 2026

Full Explanation

Gzip is the HTTP compression content coding. It carries a DEFLATE-compressed body. It is specified as a file format in RFC 1952. It is registered as an HTTP coding in RFC 9110 section 8.4.1.3. RFC 9110 calls it "an LZ77 coding with a 32-bit Cyclic Redundancy Check (CRC) that is commonly produced by the gzip file compression program". The compression itself is DEFLATE, from RFC 1951. DEFLATE uses LZ77 back-references plus Huffman coding. It is lossless, so the decoder reproduces the original bytes exactly. RFC 1951 puts the payoff at "a factor of 2.5 to 3" for English text.

Gzip is not the tightest coding available. Cloudflare measured average ratios across billions of production requests: 2.56:1 for gzip, 2.86:1 for Zstandard, and 3.08:1 for Brotli. Gzip loses to both on ratio. It is not "deflate" either. In HTTP those are two separate content codings. It is not a tool for media, either. Images, audio and video already carry their own compression. Re-coding them buys nothing. What gzip has instead is reach. It is the coding at the bottom of every negotiation ladder. Cloudflare documents falling back from Zstandard through Brotli to gzip, and only then to uncompressed. W3Techs counted it on 57.0% of all websites as of 12 September 2024. For a CDN, gzip is the floor. It is the coding you can always fall back to, rather than the one you would pick on merit.

How it works

A gzip response body is a gzip "member". RFC 1952 fixes its first ten bytes. They are: the identifying bytes ID1 = 31 (0x1f) and ID2 = 139 (0x8b); a compression method byte (CM = 8 means "deflate"); a flag byte; a four-byte modification time; and two bytes of extra flags and originating OS. If the flag byte says so, optional fields follow: an extra field, the original file name, a comment, and a header CRC. So the header is ten bytes only in the common case. Then comes the DEFLATE stream. Then comes an eight-byte trailer: CRC-32 over the uncompressed data, and ISIZE, the uncompressed length modulo 2^32. A gzip file may hold several members concatenated. An HTTP body normally holds one.

DEFLATE itself does two things. LZ77 replaces a repeated run of bytes with a back-reference to an earlier occurrence. It is written as a length and a distance. Gzip allows lengths of 3 to 258 bytes and distances up to 32,768. That distance is also its sliding-window size. Huffman coding then assigns short bit patterns to frequent symbols and long ones to rare symbols. That fixed 32 KB window is the structural reason newer codings win. Zstandard is not bounded by it. Brotli can use windows from 1 KB to 16 MB.

Negotiation over HTTP runs in this order:

  1. The client states what it will take in an Accept-Encoding request header. For example: Accept-Encoding: gzip, br. RFC 9110 section 12.5.3: "Accept-Encoding indicates the content codings acceptable in a response."
  2. The server compresses the representation. It labels the result with Content-Encoding: gzip. This is mandatory, not decorative: "the sender that applied the encodings MUST generate a Content-Encoding header field that lists the content codings in the order in which they were applied" (RFC 9110 section 8.4). Without it, the client has no way to know the body needs inflating.
  3. The bytes on the wire now depend on a request header. So the response carries Vary: Accept-Encoding (see Vary, RFC 9110 section 12.5.5). This stops caches from handing a compressed copy to a client that cannot inflate it. RFC 9110 makes sending Vary a SHOULD for origin servers, not a MUST.
  4. The client inflates the body. RFC 9110 notes that "typically, the representation is only decoded just prior to rendering or analogous usage". So the coding stays invisible to the media type in Content-Type.

One distinction worth holding onto: Content-Encoding is a property of the representation. Transfer-Encoding, by contrast, is a property of a single hop. RFC 9110 section 8.4 spells it out: "unlike Transfer-Encoding, the codings listed in Content-Encoding are a characteristic of the representation". That is why a gzip body survives being cached, revalidated and forwarded. A chunked transfer does not.

Compression level trades ratio against CPU. zlib is the reference implementation behind almost every gzip deployment. It offers levels 1 to 9, plus level 0 for no compression at all. Cloudflare compressed 10,655 HTML, CSS and JavaScript files at each level and published both curves. Output size fell from 31.5% of the original at level 1, to 28.9% at level 4, to 27.8% at level 6, and to 27.7% at level 9. Throughput fell from 125.5 MB/s to 90.7, 58.0 and 41.9 MB/s respectively. Ratio flattens after about level 4, while cost keeps climbing. Levels 8 and 9 produced identical output size in that test. Cloudflare's own production baseline at the time was zlib quality 8.

Why it matters for a CDN

Compression is the cheapest large win a CDN has on text. It is also the most expensive thing a CDN does per byte. Cloudflare has described it as "one of the most compute expensive operations our servers perform". The bill is CPU at the edge. The return is fewer bytes on the wire: "compressed content takes less time to transfer, and consequently reduces load times".

The saving lands on both legs, not just the last one. Toward the client, it shortens transfer on slow and metered links. That is why Akamai brands its gzip behaviour Last Mile Acceleration. Toward the origin, it cuts egress. Cloudflare requests content with accept-encoding: br, gzip on every plan. It frames end-to-end compression as a way for customers to reduce egress costs. An edge can therefore compress on the way out, cache a compressed copy it made on the way in, or simply forward what the origin already compressed.

The catch is that a gzip response is a variant. That complicates the cache key. RFC 9111 section 4.1 turns the header names in Vary into a secondary cache key. A cache "MUST NOT use that stored response without revalidation unless all the presented request header fields nominated by that Vary field value match those fields in the original request". An absent header only matches another absent header. Accept-Encoding has many semantically equivalent spellings. Taken literally, this shatters a shared cache into near-duplicate objects. Fastly therefore normalises the header value. A copy stored for Accept-Encoding: gzip, br can then also serve Accept-Encoding: gzip, br, deflate.

Compressing at the edge also rewrites response metadata. Cloudflare documents that it may omit Content-Length when it compresses a response. A cache-control: no-transform header from the origin stops it re-coding and preserves the length. Cloudflare honours that header only from the origin, never from the client. Validators are affected too. A content coding is part of the representation. RFC 9110 defines a strong validator as one that "changes value whenever a change occurs to the representation data" and is "unique across all versions of all representations". Because of that, the gzipped and identity forms of a resource must not share a strong ETag (RFC 9110 section 8.8.1).

What CDNs do

  • Cloudflare compresses toward the visitor by default, with gzip, Brotli or Zstandard. It chooses based on the visitor's accept-encoding value, the zone plan, and any Compression Rule. The plan default is documented: Zstandard on Free, Brotli on Pro and Business, gzip on Enterprise. It compresses only status 200, 403 and 404. It compresses only above a minimum size: 48 bytes for gzip, 50 for Brotli and Zstandard. Toward the origin it always sends accept-encoding: br, gzip. Enabling features that rewrite the body (Polish, Rocket Loader, Email Address Obfuscation and others) forces it to decompress and recompress. Cloudflare content compression docs.
  • Fastly offers gzip and Brotli at the edge in two shapes. Static compression runs pre-cache, when the response arrives from origin. So the compressed object is what gets stored and reused. It is CDN-services only, incompatible with Edge Side Includes, and billed on the compressed size. Dynamic compression runs post-cache, on the way to the client. It is switched on per response with an X-Compress-Hint header. It is billed on the size before compression. Fastly also normalises Accept-Encoding to keep the number of cached variants down. Fastly compression docs.
  • Akamai compresses at the edge with Last Mile Acceleration. This is gzip only. The property looks for Accept-Encoding: gzip and then applies Compress Response as configured (Always, Never, or Same as origin response). 860 bytes is the smallest object Akamai edge servers will gzip. The docs advise applying LMA to objects larger than 4.2 KB. Brotli is a separate behaviour that does not compress at the edge at all. The origin must produce the Brotli copy. If a gzip copy is already cached, Akamai serves that to Brotli-capable clients until its TTL expires. Akamai LMA docs, Akamai Brotli Support docs.

Watch out for

  • "deflate" is not gzip. The HTTP "deflate" coding is "a 'zlib' data format containing a 'deflate' compressed data stream". RFC 9110 warns that "some non-conformant implementations send the 'deflate' compressed data without the zlib wrapper". That is two incompatible things behind one token. In practice nobody offers it: none of the Cloudflare, Fastly or Akamai compression docs list deflate as a coding they will produce.
  • A missing Accept-Encoding does not forbid compression. This is the rule most often stated backwards. RFC 9110 section 12.5.3 says that "if no Accept-Encoding header field is in the request, any content coding is considered acceptable by the user agent". It also says that * matches any coding not explicitly listed. What actually forbids gzip is an explicit signal. An empty Accept-Encoding field value means the client wants no coding. A coding carrying q=0 is not acceptable. CDNs are usually stricter than the spec. Akamai's LMA compresses only when it sees Accept-Encoding: gzip. But do not write logic that assumes an absent header bans gzip. And do not read the spec's permission as a licence to compress for clients that said q=0.
  • Small and already-compressed responses are pure loss. Cloudflare skips anything under 48 bytes. Akamai will not gzip below 860 bytes, because the compress and decompress overhead cancels the gain. nginx defaults gzip_min_length to 20 bytes. But nginx derives that length only from Content-Length. So a response of unknown length escapes the floor entirely. For data that is already compressed there is nothing to find: DEFLATE's worst case actually expands the input by 5 bytes per 32 KB block. Akamai additionally warns off PDFs, which throw an error. It also warns off chunked responses.
  • The top levels buy almost nothing. On Cloudflare's corpus, moving from zlib level 6 to level 9 took output from 27.8% to 27.7% of the original. Throughput fell from 58.0 to 41.9 MB/s in that move. Reserve the top levels for build time.
  • Compression over TLS leaks. BREACH recovers secrets from HTTPS responses. It does this by injecting chosen plaintext and watching how the compressed size responds. Its authors list three preconditions: the application is served from a server using HTTP-level compression, it reflects user input in response bodies, and it reflects a secret such as a CSRF token in the same response. nginx repeats the warning on its own gzip page. Compressing responses that mix attacker-controlled input with secrets is the hazard, not compression as such.

Best practice

  • Treat gzip as the floor and negotiate upward. Offer Brotli or Zstandard to clients whose Accept-Encoding advertises them. Fall back to gzip for everyone else. This is exactly the ladder Cloudflare documents.
  • Pre-compress static assets at build time rather than at every request. Serve the stored file untouched. nginx does this with gzip_static. It sends a matching .gz file in place of the original. Note that the module is not built by default; it needs --with-http_gzip_static_module. nginx also asks that the original and the .gz carry the same modification time. Build-time CPU is free, so this is where the expensive settings belong. zopfli emits gzip-compatible output smaller than zlib can manage, at a cost that would be unacceptable on the fly.
  • For responses compressed per request, pick a mid level rather than the maximum. The measured curve flattens after about zlib level 4. Cloudflare's own production baseline was quality 8. So treat any specific number as something to measure on your own content, not a constant.
  • Set a minimum size. Never compress media, archive formats, or a body that already carries a Content-Encoding.
  • Send Vary: Accept-Encoding from the origin whenever the coding varies by request. Keep the header values you send narrow. Every extra distinct Accept-Encoding spelling is another cache variant unless your CDN normalises it.
  • Use cache-control: no-transform from the origin on any response the edge must not re-code. Keep Content-Encoding out of the picture entirely for endpoints that reflect user input alongside secrets.

Interactive Animation

Loading animation...

Examples

# Nginx gzip configuration
gzip on;
gzip_comp_level 6;
gzip_min_length 256;
gzip_types
    text/html
    text/css
    text/javascript
    application/javascript
    application/json
    application/xml
    image/svg+xml;
gzip_vary on;

# Pre-compress static files (build step)
$ gzip -k -9 bundle.js  # Creates bundle.js.gz

# Nginx: serve pre-compressed files
gzip_static on;  # Serves .gz file if it exists

# Test gzip support
$ curl -H 'Accept-Encoding: gzip' -sI https://example.com/ | grep content-encoding
Content-Encoding: gzip

Frequently Asked Questions

Gzip is the HTTP content coding that wraps a DEFLATE stream (LZ77 plus Huffman coding) in a header and a CRC-32 trailer, defined by RFC 1952. RFC 1951 puts the gain at a factor of 2.5 to 3 for English text. It is HTTP's universal fallback coding; Brotli and Zstandard compress tighter.

# Nginx gzip configuration
gzip on;
gzip_comp_level 6;
gzip_min_length 256;
gzip_types
    text/html
    text/css
    text/javascript
    application/javascript
    application/json
    application/xml
    image/svg+xml;
gzip_vary on;

# Pre-compress static files (build step)
$ gzip -k -9 bundle.js  # Creates bundle.js.gz

# Nginx: serve pre-compressed files
gzip_static on;  # Serves .gz file if it exists

# Test gzip support
$ curl -H 'Accept-Encoding: gzip' -sI https://example.com/ | grep content-encoding
Content-Encoding: gzip

Yes. Gzip is also known as x-gzip. Gzip is the HTTP content coding that wraps a DEFLATE stream (LZ77 plus Huffman coding) in a header and a CRC-32 trailer, defined by RFC 1952. RFC 1951 puts the gain at a factor of 2.5 to 3 for English text. It is HTTP's universal fallback coding; Brotli and Zstandard compress tighter.

Related CDN concepts include:

  • Content-Encoding — The HTTP header naming the content codings applied to a body (gzip, br, zstd), which …
  • Vary Header — Vary is the response header that names the request headers the origin used to select …
  • Brotli — Brotli is a lossless compression format from Google, specified in IETF RFC 7932 and carried …
  • Compression — An HTTP content coding that shrinks a message body, almost always a response body, so …