LL-DASH (Low-Latency DASH)
Low-latency mode of DASH, not a separate protocol: each CMAF segment is cut into small chunks and streamed in one HTTP/1.1 chunked response while it is still being written, so the player decodes before the segment is complete. Takes glass-to-glass latency from tens of seconds to a few.
Also known as CMAF-CTE (Chunked Transfer Encoding), Low-Latency Modes for DASH.
Full Explanation
LL-DASH (Low-Latency DASH) is the low-latency mode of DASH, MPEG's adaptive HTTP streaming standard (ISO/IEC 23009-1). It is not a separate protocol and not a new file format. It is ordinary DASH live content. Each segment is cut into small CMAF chunks and published while it is still being written, and the player fetches those chunks before the segment is finished. The codecs, the segment URLs and the shape of the manifest do not change. What changes is when the first byte of a segment becomes available. Also new: every hop must forward that segment while it is still growing. LL-DASH is the DASH counterpart of LL-HLS. One CMAF encode can feed both.
Latency in plain live DASH is dominated by waiting for whole objects. First, the encoder finishes a segment. Then the origin publishes it. The player downloads all of it. Then it buffers several more. The DASH-IF/DVB report addresses this problem for segmented HTTP streaming, naming HLS. It says: "Latencies known from proprietary system such as HLS up to 30-seconds or more can make or break a viewing experience" (DASH-IF/DVB Report on Low-Latency Live Service with DASH). LL-DASH removes the whole-object wait at every hop. It lands in the range of a few seconds. The rules come from DASH-IF's Low-Latency Modes for DASH, part 4 of DASH-IF IOP v5. The transport is normally HTTP/1.1 chunked transfer coding. That is why the industry also calls the technique CMAF-CTE. That name is only half the story. Dolby OptiView notes it is "not entirely accurate as CMAF-CTE is only a method used in LL-DASH".
How it works
The chunk. A CMAF chunk is a fragment-of-a-fragment. It is the smallest unit the player can parse and decode without waiting for its parent segment. A low-latency segment is still a normal CMAF fragment. But it contains more than one moof/mdat pair. Each pair "may contain any number of ISO BMFF samples between 1 and the full segment duration inclusive" (Low-Latency Modes for DASH, clause 9.X.5). The dash.js documentation puts the consequence plainly. Multiple moof and mdat boxes allow "the client to access the media data before the segment is completely finished" (dash.js low-latency guide).
The transport. Chunked transfer coding lets the origin start the response before it knows how long the body will be. This is stated directly: "Chunked enables content streams of unknown size to be transferred as a sequence of length-delimited buffers" (RFC 9112 section 7.1). One HTTP response carries one segment. Its body grows as the encoder produces chunks. DASH-IF is careful to call this an example rather than the mechanism. It states: "A low delay protocol, e.g. HTTP Chunked Transfer Encoding, of partially available Segments is used such that clients can access the Segments before they are completed." Mapping each CMAF chunk to one HTTP chunk "is a recommendation for low-latency operation, but not a requirement". Also, "By no means, the client should assume that this 1-to-1 mapping is preserved to the client". Any edge or proxy on the path may re-frame the HTTP chunking.
The manifest signals. A conforming service advertises the profile identifier http://www.dashif.org/guidelines/low-latency-live-v5. It also carries the signals a low-latency client needs. @availabilityTimeOffset says how much earlier a segment can be fetched than its computed availability start time. @availabilityTimeComplete set to FALSE declares that the segment is produced progressively. So "it may be inferred by the client that the segment is available at its announced location prior to completion". A ServiceDescription element carries a Latency target. This tells the client what live edge to aim for. At least one UTCTiming element must be present. It must be millisecond accurate and use one of urn:mpeg:dash:utc:http-xsdate:2014, http-iso:2014 or http-ntp:2014. A ProducerReferenceTime element is also required. Together these let the client tie media time to the producer's wall clock. In the current v5 text, the offset must be greater than zero. It must be smaller than the maximum segment duration for the Representation. It must be the same for every Representation. The difference between it and the maximum segment duration must be smaller than the target latency.
Here is how @availabilityTimeOffset and @availabilityTimeComplete look in a segment template:
<!-- LL-DASH MPD with availability time offset -->
<SegmentTemplate
timescale="1000"
duration="6000"
availabilityTimeOffset="5.5"
availabilityTimeComplete="false"
media="chunk_$Number$.m4s"
initialization="init.mp4" />
<!-- duration is counted in timescale units,
so 6000/1000 gives 6 second segments -->
<!-- availabilityTimeOffset tells the player it can
request the segment 5.5s before it is fully available -->
<!-- availabilityTimeComplete="false" says the segment is
still being written when it is published; the default
of true would claim it is already complete -->
The player. It computes the early availability time from the manifest and its synchronised clock. It issues the GET. It decodes each chunk as it lands rather than at end of segment. A DASH-IF low-latency client "should play the content within 500ms tolerance of the target latency". ABR has to be rebuilt around this. The reason: "the download time of a segment is often times similar to the segment duration" (dash.js). The link is idle between chunks. So segment-size divided by elapsed time no longer measures the network. Low-latency clients time the arriving chunk bytes instead.
Why chunks rather than just shorter segments. This is the point most summaries miss. "One reason for using chunks instead of shorter Segments is that Segments must start with an SAP type 1 or 2, i.e., for video a closed GOP. This decreases the coding efficiency of video. CMAF Chunks do not have this restriction since they are generated for neither bitrate switching nor randomly accessing the Representation." Chunking therefore buys latency without paying in keyframe frequency or bitrate. The trade is that an ordinary chunk is not a switch point. The current v5 text closes that gap with the Resync element (added by ISO/IEC 23009-1:2020/Amd.1). This element signals chunk properties. It lets an encoder mark chunks starting with an SAP of type 1, 2 or 3, "permitting downswitching, resynchronization and random access" inside a segment.
Chunk sizing. Not a dial to turn to zero. One sample per chunk "is not desirable" with efficient B-frame encoding. For video, the boundary should fall so that every B-frame displayed before a P-frame is in the same chunk. For the worked example in the spec, this is a multiple of 4-5 frames, "160-200ms for 25Hz video". For audio, a 200 ms chunk at 64 Kbps is only 1.6 kB. It "may possibly be too small to propagate through network or receiver buffers". So "it may make sense to have longer chunks (e.g. 0.5s for audio) or even not applying chunking for audio but run at shorter Segment duration". Subtitles are better left unchunked and delivered as short segments.
Why it matters for a CDN
LL-DASH changes what a CDN is asked to be. A cache of finished objects becomes a relay of an object that does not exist yet. DASH-IF draws the line explicitly: "HTTP chunked transfer encoding must at least be supported up from the ingest into the packager up to the CDN edge, whereas the last mile delivery is expected happen using HTTP chunked transfer encoding or HTTP in regular mode." So the hop that must not buffer is origin-to-edge. A legacy player on the last mile can still pull whole segments. That is why the same stream stays playable by ordinary DASH clients at ordinary latency.
DASH-IF is deliberately undramatic about the CDN itself. The specification states it "makes full use of the ubiquitously available HTTP/1.1, and modern CDN architectures are therefore fully usable, possibly with some updates on handling the relay of partially complete Segments through the network." That clause is the whole engineering problem in one line. Everything else about the CDN is normal.
The pay-off is that LL-DASH is unusually cache-friendly for a low-latency format. The objects the CDN keeps are still whole segments. The design "enables CDN-cache-friendly and efficient operation as the Segments that are stored on the CDN for later consumption are of typical Segment duration ranges of several seconds". Also, "the low-latency content is fully cacheable for consumption in time-shift and network-PVR mode without modification". The same bytes serve the DVR window and catch-up. LL-HLS, by contrast, addresses each part as its own resource. Gcore runs both. It describes the split: "LL-DASH requires the CDN to intelligently cache and serve partially complete files. LL-HLS requires the CDN to handle a massive volume of short, bursty requests and hold connections open for manifest updates. A traditional CDN is optimized for neither" (Gcore: how we engineered a single pipeline for LL-HLS and LL-DASH).
One protocol detail catches people out at the edge. Chunked transfer coding is an HTTP/1.1 mechanism only. HTTP/2 has no such thing: "The chunked transfer encoding defined in Section 7.1 of [HTTP/1.1] cannot be used in HTTP/2" (RFC 9113 section 8.1). An edge terminating HTTP/2 or HTTP/3 for the player streams the same partial segment as DATA frames without a Content-Length instead. This works fine. DASH-IF says the design "may benefit from improved network protocols, such as HTTP/2 or HTTP/3, but does not require any of those". The requirement is incremental forwarding, not a particular framing.
Finally, here is an honest caveat about the last mile. Broadcast-parity latency assumes a good connection. The DASH-IF/DVB report warns that "the low latency requirement implied by this is likely to require a high performance internet connection and that not all viewers may have suitable connections to achieve this today" (DASH-IF/DVB report).
What CDNs do
Support is real but opt-in, and it differs per vendor. Check the vendor's own current documentation rather than assuming.
- Akamai exposes low-latency DASH as an explicit workflow in Adaptive Media Delivery. Use a Media Services Live origin. Set Enable ULL Streaming to On under Segmented Media Delivery Mode. Have players request a path containing /dash/live-ull/ (or /cmaf/live-ull/ for DASH media in CMAF containers). Akamai applies an "aggregating response" that breaks the content into smaller chunks. Its best-practice list states the requirement squarely: "The CDN needs to propagate this content all the way to the client using HTTP chunked encoding transfer at each step in the distribution chain." The page notes this method is DASH-only. HLS needs a third-party origin. See Akamai: DASH and a Media Services Live origin.
- Gcore built a purpose-made edge module. In its own words: "For LL-DASH, we developed a custom caching module we call chunked-proxy. When the first request for a new .m4s segment arrives, our edge server requests it from the origin. As bytes flow in from the origin, the chunked-proxy immediately forwards them to the client." A second viewer of the same segment gets the cached prefix and then rides the same live tail. Gcore reports "a glass-to-glass latency of approximately 2.0 seconds for LL-DASH and 3.0 seconds for LL-HLS" from one packaging pipeline. See Gcore.
- Fastly streams misses through by design. But two relevant features pull in opposite directions, and you have to pick one. Streaming Miss "ensures the response is streamed back to the client immediately and is written to cache only after the whole object has been fetched". That is the behaviour LL-DASH needs. Enable it by setting beresp.do_stream in VCL. Segmented Caching is the feature for caching very large objects as byte ranges. It cannot be used on the same path. Among its documented limitations: "HTTP chunked transfer encoding between Fastly and origin isn't supported. Your origin server must frame responses to Range: requests with the Content-Length header." An LL-DASH origin emits an unknown-length chunked response. That does not satisfy the requirement. So route low-latency live away from Segmented Caching.
Watch out for
- One buffering hop cancels the whole thing. Consider the chain: encoder, packager, origin, shield, edge, player. If any of them replies only once the object is whole, the stream behaves exactly like plain DASH. Akamai's conditions for "stable latency reduction" list every actor separately for this reason. This runs from "The encoder needs to push content to the origin using HTTP 1.1 chunked encoding transfer" through to "The player needs to decode the bitstream as it's received and not wait for the end-of-segment."
- A compliant manifest over a non-compliant path. The classic LL-DASH failure is an MPD that correctly advertises @availabilityTimeOffset and @availabilityTimeComplete FALSE. But some server or intermediary still waits for segment completion before it answers. Conformance checks on the manifest pass. Measured latency does not move. Test the wire, not the MPD.
- Throughput estimation silently misreads the network. A chunked download lasts roughly the segment duration. Because of this, classical ABR reads a fraction of the true capacity and downswitches for no reason. The player must measure the incoming chunk bursts, as dash.js does.
- Clock drift. Fetching at an announced early availability time only works if the client agrees with the producer about now. Get UTCTiming wrong and the client asks too early. That means a 404 or a stall. Or it asks too late. Then it has quietly given up its latency. Dolby OptiView is blunt about the dependency. It says: "This does require an accurate synchronization between the clocks of the client and server. For this, MPEG-DASH allows you to configure a time server." DASH-IF requires a millisecond-accurate UTCTiming element. It also requires a ProducerReferenceTime the client can anchor to.
- Segment duration is bounded by the latency target, not by taste. For a chunked Adaptation Set with more than one Representation, the maximum segment duration "shall be smaller than the signaled target latency and should be smaller than half of the signaled target latency". Chasing latency by shrinking segments instead of chunking costs coding efficiency. It also multiplies request and cache overhead for little gain.
- DRM and SSAI spend the budget you no longer have. With small buffers, a license round-trip "could cause a few hundreds of milliseconds or even seconds delay". So Dolby OptiView recommends putting the DRM initialisation data (PSSH) in the manifest rather than the init segment. This lets the client start license negotiation before any media is fetched. Server-side ad insertion via DASH Periods forces frequent manifest refreshes, because the client can no longer predict where the next segment appears. This "makes doing SSAI with Low Latency DASH far from trivial, and should be evaluated carefully".
- Reach. DASH needs a Media Source API in the browser. On iPhone that arrived late: "Safari 17.1 now brings the new Managed Media Source API to iPhone" (WebKit). Older iOS still needs LL-HLS or plain HLS. So LL-DASH rarely ships alone.
Best practice
- Signal the contract properly. Include the http://www.dashif.org/guidelines/low-latency-live-v5 profile identifier. Include a ServiceDescription with a Latency target. Include a millisecond-accurate UTCTiming element using one of the three permitted schemes. Include a ProducerReferenceTime. Set @availabilityTimeComplete to FALSE. Set @availabilityTimeOffset greater than zero, smaller than the maximum segment duration, identical across Representations, and within the target latency of that maximum. Add a Resync element per Representation so clients can downswitch and resynchronise inside a segment.
- Prove the partial relay end-to-end before trusting the numbers. Request a segment at its early availability time and confirm first bytes arrive well before end-of-segment at the edge, not just at the origin. If time-to-first-byte tracks segment completion, the low latency is nominal.
- Pick the cache path deliberately per vendor: incremental forwarding of a growing object on the live path, and normal segment caching for the catch-up and DVR reads of the same URLs.
- Tune chunk duration per media type rather than globally. For video, set chunk boundaries at B-frame group edges. For audio, use around 0.5 s chunks, or leave audio unchunked with shorter segments. Leave subtitles unchunked.
- Use a player that actually implements the low-latency mode. The DASH-IF reference client dash.js documents its requirements and its chunk-aware throughput estimator. Verify the player holds within 500 ms of the signalled target latency.
- Measure real glass-to-glass latency continuously. Keep a graceful fallback for viewers whose connections cannot hold the live edge. LL-DASH content remains ordinary cacheable DASH for them.
Examples
A live football match uses LL-DASH. The stream has 500ms CMAF chunks inside 6-second segments. The player asks for segment 500 while the encoder still makes it. The CDN streams the chunks as they arrive. The viewer sees the goal about 3 seconds later, not 25 seconds.
Frequently Asked Questions
Low-latency mode of DASH, not a separate protocol: each CMAF segment is cut into small chunks and streamed in one HTTP/1.1 chunked response while it is still being written, so the player decodes before the segment is complete. Takes glass-to-glass latency from tens of seconds to a few.
A live football match uses LL-DASH. The stream has 500ms CMAF chunks inside 6-second segments. The player asks for segment 500 while the encoder still makes it. The CDN streams the chunks as they arrive. The viewer sees the goal about 3 seconds later, not 25 seconds.
Yes. LL-DASH (Low-Latency DASH) is also known as CMAF-CTE (Chunked Transfer Encoding), Low-Latency Modes for DASH. Low-latency mode of DASH, not a separate protocol: each CMAF segment is cut into small chunks and streamed in one HTTP/1.1 chunked response while it is still being written, so the player decodes before the segment is complete. Takes glass-to-glass latency from tens of seconds to a few.