LL-HLS (Low-Latency HLS)

Streaming

Apple's low-latency mode for HLS. The packager publishes sub-second partial segments (EXT-X-PART) that players render before the parent segment is finished, and preload hints plus blocking playlist reloads replace polling, cutting live delay from 18-30 seconds to roughly 3-5.

Also known as LL-HLS, Low-Latency HLS, Low-Latency HTTP Live Streaming.

13 min read Updated Aug 30, 2026

Full Explanation

LL-HLS is the low-latency mode of HLS. Apple introduced it at WWDC 2019. It is specified as the low-latency extensions of HTTP Live Streaming 2nd Edition, revision 7 and later. It is not a new protocol, a new transport or a new codec. The stream is still an M3U8 playlist pointing at ordinary segments fetched over HTTP. A player that does not recognise the new tags ignores them and watches the same stream at regular latency. Two things change. Media is published in sub-second partial segments rather than whole segments. The client also stops polling for updates. It requests the next piece before that piece exists, and the server holds the request open until it can answer. Apple set a design target of one to two seconds from live. Implementers quote roughly 3 to 5 seconds in practice, against 18 to 30 for standard HLS. LL-HLS is the HLS counterpart of LL-DASH. The two converge when parts are addressed as byte ranges of CMAF resources.

How it works

Regular HLS is slow for two independent reasons. The first is publishing latency. The specification says “a Segment cannot be distributed until it has been completely encoded and packaged. A long Segment encoded in real-time introduces a delay equal to its duration” (specification, section 3.2). With the recommended six-second Target Duration, six seconds pass before anything can even reach the CDN. The second is discovery. The client learns that a new segment exists only by polling the playlist. In the bad case that takes “almost another six seconds before the client even finds out that there’s a new segment there”, followed by a further round trip to fetch it (WWDC 2019, session 502). LL-HLS attacks both problems. It uses three mechanisms that only work together.

  1. Partial segments, the EXT-X-PART tag. The live edge of the playlist gains a parallel channel. In it, “the media is divided into a larger number of smaller pieces, such as CMAF Chunks”. Each part “can be packaged, published, and added to the Media Playlist much earlier than its Parent Segment” (section 3.2). Each part is published as soon as it is packaged, so the player fetches and renders it long before its parent segment exists. Apple’s documentation offers 200 milliseconds as an example part length against six-second segments. The value Apple’s authoring requirements actually recommend is a one-second Part Target Duration. Apple’s own reference playlist uses 0.33334. A part that begins with an independent frame carries INDEPENDENT=YES. That is where a player can start playback or switch rendition.
  2. Preload hints, the EXT-X-PRELOAD-HINT tag. The playlist advertises the URL of the next part before that part exists. The tag “allows a Client loading media from a live stream to reduce the time to obtain a resource from the Server by issuing its request before the resource is available to be delivered. The server will hold onto the request (‘block’) until it can respond” (section 4.4.5.3). The round trip is spent before the media exists instead of after.
  3. Blocking playlist reload. The client appends delivery directives to the playlist URL. _HLS_msn=M names the media sequence number it wants next. _HLS_part=N names the part within it. The server then “MUST defer responding to the request until the Playlist contains the Partial Segment with Part Index N and with a Media Sequence Number of M or later”. It advertises the capability with CAN-BLOCK-RELOAD=YES on EXT-X-SERVER-CONTROL (section 6.2.5.2). Polling disappears. Every playlist update arrives on a request the client has already made.

Two supporting features pay for the extra playlist traffic. Playlist delta updates let the client add _HLS_skip=YES and receive only what changed. The older portion of the playlist collapses into an EXT-X-SKIP tag. Apple’s authoring requirements say services with playlist windows longer than two minutes should offer them. Rendition reports (EXT-X-RENDITION-REPORT) publish the last segment and part of every other rendition inside each playlist. This lets a player switch bitrate without first fetching the other playlist. Rendition reports are a hard requirement of the low-latency server profile. Without them a client would have to run the tune-in algorithm on every switch.

The live edge of a low-latency playlist

Apple’s reference playlist shows the shape. Reading down from the header:

  • #EXT-X-TARGETDURATION:4: parent segments are four seconds.
  • #EXT-X-PART-INF:PART-TARGET=0.33334: the Part Target Duration. This tag is required whenever the playlist contains an EXT-X-PART tag.
  • #EXT-X-SERVER-CONTROL:CAN-BLOCK-RELOAD=YES,PART-HOLD-BACK=1.0,CAN-SKIP-UNTIL=24.0: blocking reload is offered. The server recommends playing 1.0 second back from the end of the playlist. Delta updates may skip everything older than 24 seconds.
  • #EXT-X-PART:DURATION=0.33334,URI="filePart273.0.mp4",INDEPENDENT=YES: the parts of the segment currently being produced, listed ahead of the EXTINF line of their parent.
  • #EXT-X-PRELOAD-HINT:TYPE=PART,URI="filePart273.3.mp4": the next part, which does not exist yet. The client requests this URL immediately.
  • #EXT-X-RENDITION-REPORT:URI="../1M/waitForMSN.php",LAST-MSN=273,LAST-PART=2: how far every other rendition has got.

Here is the live edge of such a playlist, with round numbers: six-second segments cut into half-second parts.

#EXTM3U
#EXT-X-VERSION:6
#EXT-X-TARGETDURATION:6
#EXT-X-MEDIA-SEQUENCE:98
#EXT-X-SERVER-CONTROL:CAN-BLOCK-RELOAD=YES,PART-HOLD-BACK=1.5
#EXT-X-PART-INF:PART-TARGET=0.5
#EXT-X-MAP:URI="init.mp4"
#EXT-X-PROGRAM-DATE-TIME:2025-01-01T10:00:00.000Z

#EXTINF:6.0,
segment098.m4s
#EXTINF:6.0,
segment099.m4s
#EXTINF:6.0,
segment100.m4s
# segment101 is still being made, so its finished parts are listed instead
#EXT-X-PART:DURATION=0.5,INDEPENDENT=YES,URI="segment101.0.m4s"
#EXT-X-PART:DURATION=0.5,URI="segment101.1.m4s"
#EXT-X-PART:DURATION=0.5,URI="segment101.2.m4s"
# this part does not exist yet: the player asks for it, the server holds the request
#EXT-X-PRELOAD-HINT:TYPE=PART,URI="segment101.3.m4s"

PART-HOLD-BACK is the latency dial. The specification requires it to be at least twice the Part Target Duration. It says PART-HOLD-BACK should be at least three times the Part Target Duration. Apple’s authoring requirements for Apple devices make three times a MUST. Once a parent segment is complete, its parts are redundant. The server “removes Partial Segments from the Media Playlist once they’re greater (older) than three target durations from the live edge”. So the playlist does not grow without bound.

One point is widely mis-stated. HTTP/2 push is not part of LL-HLS. The 2019 design pushed parts over HTTP/2 in response to a blocking playlist request. Apple removed that requirement in 2020. It replaced push with the preload hint, which is an ordinary blocking GET. The Low-Latency Server Configuration Profile does require that “HTTP-delivered Playlists and Segments are served via HTTP/2 [RFC9113] or HTTP/3 [RFC9114]” (appendix B.1). That is a requirement on the delivery path, not merely a recommendation. The requirement now includes HTTP/3.

Why it matters for a CDN

  • The CDN is inside the latency budget. At 18 to 30 seconds of standard-HLS delay, cache fill and network distance were noise. At three seconds they are a measurable share of what the viewer sees. One avoidable edge-to-origin round trip on a 333-millisecond part can consume the whole error budget.
  • Blocking requests are the normal case. A low-latency edge holds large numbers of playlist and part requests open. Each one waits on media that does not yet exist. The profile makes this the CDN’s job explicitly: “CDNs and other proxy caches recognize blocking requests for Playlists and Media Segments whose cache fill is already pending, and hold the duplicate requests until they can be delivered from that cache fill. This minimizes the load on the active origin.” An edge that treats a long-lived request as a stall, or that forwards every waiter to the origin, breaks LL-HLS.
  • The cache key now carries semantics. _HLS_msn, _HLS_part and _HLS_skip determine which response is correct, not merely which variant is convenient. Ignore them and the edge serves one stale playlist to everybody.
  • Request rate multiplies. Sub-second parts plus a playlist fetch per part mean far more requests per viewer than standard HLS. LL-HLS “results in more requests per second. This is mainly due to video segments being smaller (sub-second as opposed to second duration)” (Fastly). That means many small objects at high rate, each with a short useful life.
  • Tune-in depends on the Age header. A client that starts from a cached playlist cannot honour PART-HOLD-BACK correctly. So the CDN tune-in algorithm in appendix C uses the Age response header to work out how stale its first playlist was. The profile therefore requires that “HTTP caches used to deliver Playlists or Segments will set the Age HTTP Response header”. An edge that strips Age makes correct low-latency tune-in impossible.

What CDNs do

  • Akamai Adaptive Media Delivery. Configure your own origin. Set Segmented Media Delivery Mode to Live with Enable ULL Streaming on. Enable HLS in Content Characteristics. Add the HTTP/2 behaviour over HTTPS. Add _HLS to the cache-key parameter list with exact match off, because “an LL-HLS player asks for manifests explicitly using special _HLS query parameters”. Parts can be addressed two ways. DISCRETE mode gives each part a unique URL, and “is the most common mode and all players support it”. BYTERANGE mode requests segments at a byte-range offset. Akamai says BYTERANGE “can reduce the number of requests from the client to the edge by 40%”. Because “the origin holds back requests for the duration of an individual part”, the ULL defaults of a 2-second read timeout at the edge and 1 second at the parent sit uncomfortably close to a one-second part. They need extra margin. Akamai also suggests emitting Timing-Allow-Origin. That lets a web player use the Resource Timing API to separate hold-back time from download time and estimate bandwidth correctly.
  • Cloudflare Stream. LL-HLS is an open beta. It is enabled per Live Input by a “Low-Latency HLS Support” toggle, with latency “to as little as three seconds” (announcement). The built-in player uses low latency when the input is enabled. It falls back to regular HLS otherwise. A custom player can append ?protocol=llhls to the HLS manifest URL. Cloudflare documents this as a way to test the low-latency manifest. It warns the flag “may change in the future and is not yet ready for production usage”.
  • Fastly. Fastly publishes a sample Compute application, compute-ll-hls, that sits in front of an LL-HLS origin. It “aims to absorb some of this increased load by handling most of those requests at the edge” through request collapsing and by “responding to Delta playlist requests at the edge”. It is demonstration code rather than a product feature. But it is a clean statement of the two edge behaviours that matter for LL-HLS: collapse the duplicate waiters, and synthesise the delta playlist at the edge from a cached full one.

Watch out for

  • Do not build on HTTP/2 push. Many LL-HLS write-ups still describe the 2019 design in which the server pushed parts. That requirement was withdrawn in 2020. The current mechanism is a blocking GET against a preload hint. HTTP/2 or HTTP/3 is still required for delivery, but push is not.
  • A part goes out at line rate, not as it is produced. This is the rule most often misread. The server “MUST refrain from transmitting any bytes belonging to a Partial Segment until all bytes of that Partial Segment can be transmitted at the full speed of the link to the client” (section 6.2.6). The trigger is being able to send the entire part at line speed, not merely having finished producing it. The spec states the reason: dribbling a part out would corrupt the client’s throughput measurement and therefore its bitrate adaptation. Chunked transfer of a part is out.
  • A hinted part may never arrive. “A server MAY choose not to publish previously-hinted resources if the planned segmentation changes, such as the case of early return from an ad”, answering 404 instead. So a player must not stake playback on a hint. A hint is, however, a promise about availability, not about the future: “a hinted resource MUST be available for request when its EXT-X-PRELOAD-HINT tag is added to the Playlist”.
  • Shorter parts are not simply better. The specification gives its own warning: “A shorter Target Duration reduces latency but also reduces available buffer, handicaps adaption and increases delivery overhead, increasing the likelihood of playback stall.”
  • Blocking has defined failure codes. A server that cannot answer after blocking for more than three target durations should return 503. An _HLS_msn more than two segments beyond the playlist, or an _HLS_part beyond the Advance Part Limit, should get a 400. Monitor these apart from real origin errors. At the edge they look like origin faults, but they are the protocol working.
  • Cache lifetimes are not one value. The profile recommends caching a successful blocking playlist response for six target durations. It recommends caching a non-blocking one for only half a target duration. Same path, very different freshness. That is because a blocking response is fresh by construction at the moment it is produced.
  • The bill moves. A large multiple of the request count at the same bitrate shifts cost onto per-request charges, TLS handshakes and edge CPU rather than egress.

Best practice

  • Size the Part Target Duration to the audience rather than to a latency target. Apple requires it to be at least the round-trip time that 95% of clients see (P95 RTT). Apple says it should be at least three times the P95 RTT, and recommends one second. PART-HOLD-BACK must then be at least three times the Part Target Duration (Apple authoring requirements, section 14).
  • Include _HLS_msn, _HLS_part and _HLS_skip in the cache key. Enable request collapsing so that a thousand players waiting on the same part produce one origin fill.
  • Serve the low-latency path over HTTP/2 or HTTP/3 with TLS 1.3. Gzip the playlists and preserve the Age header. Apply the profile’s differentiated cache lifetimes for blocking and non-blocking responses.
  • Emit EXT-X-PROGRAM-DATE-TIME on every media playlist and a rendition report for every other rendition. Both are profile requirements. Program date-time is optional in standard HLS but mandatory here (AWS Elemental MediaPackage).
  • Check that the edge read timeout exceeds a part duration with margin. Alert separately on 503 and 400 responses to blocking requests.
  • Prefer byte-range part addressing where the player set allows it. It means fewer edge requests, and it is the mode that interoperates with LL-DASH.
  • Verify with a genuine LL-HLS player, and verify that a legacy HLS client still plays the same stream at regular latency. That fallback, “clients fall back to regular-latency HLS playback if they discover that the server doesn’t support an aspect of the required configuration”, is what makes LL-HLS safe to switch on for an existing stream.

Examples

A live auction platform uses LL-HLS to keep latency under 3 seconds. The packager makes 0.5-second partial segments. It publishes preload hints for each part. The CDN edge supports blocking playlist reloads over HTTP/2. The player gets each new part at once, with no polling delay.

Frequently Asked Questions

Apple's low-latency mode for HLS. The packager publishes sub-second partial segments (EXT-X-PART) that players render before the parent segment is finished, and preload hints plus blocking playlist reloads replace polling, cutting live delay from 18-30 seconds to roughly 3-5.

A live auction platform uses LL-HLS to keep latency under 3 seconds. The packager makes 0.5-second partial segments. It publishes preload hints for each part. The CDN edge supports blocking playlist reloads over HTTP/2. The player gets each new part at once, with no polling delay.

Yes. LL-HLS (Low-Latency HLS) is also known as LL-HLS, Low-Latency HLS, Low-Latency HTTP Live Streaming. Apple's low-latency mode for HLS. The packager publishes sub-second partial segments (EXT-X-PART) that players render before the parent segment is finished, and preload hints plus blocking playlist reloads replace polling, cutting live delay from 18-30 seconds to roughly 3-5.