Skip to content
Skip article header Engineering

Secure HLS Delivery

RFC 8216 spends exactly one sentence on protecting an HLS decryption key, and says nothing about who may ask for one. This guide works through the decisions that actually secure a stream: the key request as an authorization decision, signed manifests against signed segment URLs, CDN caching rules that can silently leak a key response and the boundary past which AES-128 stops helping and a licensed DRM system takes over.

Updated 20 min read 58 views
An engineer checking a key rotation log against a printed authorization flow sheet beside a media server rack.
Skip key takeaways

Secure HLS delivery is usually treated as a DRM problem, and RFC 8216, the specification describing HTTP Live Streaming, treats it as an afterthought. The document gives exactly one sentence to protecting a decryption key in transit, and nothing to who is allowed to ask for one. The specification text puts it plainly, saying only that "Encryption keys are specified by URI. The delivery of these keys SHOULD be secured by a mechanism such as HTTP Over TLS [RFC2818] (formerly SSL) in conjunction with a secure realm or a session token." That sentence is the entire standards basis for entitlement, session binding and revocation on an encrypted HLS stream. Everything past it, the key server, the token, the CDN check, is engineering a platform builds rather than a requirement the protocol hands over.

In short: the control table below is the centerpiece of this guide, ten controls a streaming platform actually runs against the two things RFC 8216 never specifies, entitlement and revocation. The key request is an authorization decision the protocol leaves entirely to the implementer. The DRM boundary marks the point where AES-128 stops helping and a licensed system such as FairPlay takes over.

What RFC 8216 actually says, and what it does not

RFC 8216 is not a mandate. Its own Status of This Memo says so on the rfc-editor.org text: "This document is not an Internet Standards Track specification; it is published for informational purposes." It is an Independent Submission by R. Pantos of Apple and W. May of MLB Advanced Media, describing HTTP Live Streaming as Apple publishes it rather than a requirement the IETF imposed on the industry, and the same text is also published as an HTML render on the IETF Datatracker. Writing that the HLS standard requires something misstates what the document claims to be. It also describes an old protocol version: dated August 2017 and scoped to version 7 of HLS, while Apple's current developer documentation points implementers at a newer draft, HTTP Live Streaming 2nd Edition, under the name draft-pantos-hls-rfc8216bis, not consulted for this guide.

One control at a time, and what each one misses

Each row below stands on its own. The stops column is sourced wherever the sections after the table give a citation for it; the does-not-stop and failure-mode columns are Pharos Production practice rather than requirements of either specification.

Control What it stops What it does not stop Failure mode seen in practice
EXT-X-KEY with METHOD=AES-128 An unauthenticated fetch, since the segment is completely encrypted under AES-128 CBC Anyone who obtains the Key file, which decrypts every segment in scope encryption runs but the key endpoint has no entitlement check, so it costs CPU and buys nothing
Key rotation, a new EXT-X-KEY per playlist The window a leaked key unlocks, since a tag applies only up to the next with the same KEYFORMAT Re-fetch of an old key from an endpoint that still serves it rotation is configured but the key server has no expiry, so the leak window is the whole archive
Authorization on the key request An unentitled viewer, if and only if the key endpoint checks entitlement A legitimate viewer copying the key or its URI out of their own session the key URI is signed but never bound to the viewer, so one session's URI plays for everyone
Signed manifest Direct fetches of the manifest by an unauthorized client Anything about the segment URLs it lists, once those are readable only the manifest carries a token, and the segment URLs inside it are shared directly
Signed segment URLs at the edge Direct segment fetches without a valid per-request token Nothing about the key; a signed segment beside an open key endpoint is the same leak one hop later the CDN checks the segment token but never the key request
Authorization header on the key response A shared cache reusing that response for a second request, per RFC 9111 An endpoint that authenticates by query string, cookie or custom header instead the token moved into the query string, and the cache protection silently stopped applying
Cache-Control: private or no-store on the key response A shared cache storing and serving one viewer's key response to another A malicious or compromised cache, which the spec says may ignore the directive the key response inherits the CDN's default HTML TTL and is served to the next viewer
Vary on the key response A shared cache reusing a response across the header dimension Vary names Everything not named; Vary narrows the cache key, it authorizes nothing Vary is set on a header that never carries viewer identity
EXT-X-PLAYLIST-TYPE and reload A client treating a live playlist as fixed, or a VOD playlist as still changing Nothing about who may fetch the manifest; a correctness rule, not a security one a per-user manifest is cached at the edge as if it were the static VOD file everyone else requests
Platform DRM (FairPlay by name; Widevine and PlayReady the other two) Key extraction from the player, on platforms that implement it Anything on a platform that does not; it is a build-and-run commitment, not a checkbox DRM is adopted to solve an entitlement problem the team could have solved at the key endpoint

EXT-X-KEY, key rotation and the CMAF case

The tag that carries the whole cryptographic argument is EXT-X-KEY, and its rotation rule is built into the definition rather than added afterward. The specification text states it in one sentence: "Media Segments MAY be encrypted. The EXT-X-KEY tag specifies how to decrypt them. It applies to every Media Segment and to every Media Initialization Section declared by an EXT-X-MAP tag that appears between it and the next EXT-X-KEY tag in the Playlist file with the same KEYFORMAT attribute (or the end of the Playlist file)." A new tag with the same KEYFORMAT ends the previous key's scope, the entire mechanism of key rotation in HLS: a playlist construct rather than a revocation channel, and nothing in the spec says how long a rotated-out key must keep working elsewhere. The specification does not let a server erase that history either, stating plainly that "The server MUST NOT remove an EXT-X-KEY tag from the Playlist file if it applies to any Media Segment in the Playlist file, or clients who subsequently load that Playlist will be unable to decrypt those Media Segments." Rotation retires a key from segments going forward; it does not, and cannot, revoke a key that is still in scope for segments already published under it.

Three methods are defined, and encryption is off by default: a Media Playlist with no EXT-X-KEY tag carries unencrypted segments, so cleartext is the starting point rather than an exception. Under AES-128, the method behind most production deployments, the whole segment is encrypted. The same text states it precisely: "An encryption method of AES-128 signals that Media Segments are completely encrypted using the Advanced Encryption Standard (AES) [AES_128] with a 128-bit key, Cipher Block Chaining (CBC), and Public-Key Cryptography Standards #7 (PKCS7) padding [RFC5652]. CBC is restarted on each segment boundary, using either the Initialization Vector (IV) attribute value or the Media Sequence Number as the IV; see Section 5.2."

Under SAMPLE-AES, only the media samples inside a segment are encrypted, not the container around them. The spec puts it this way: "An encryption method of SAMPLE-AES means that the Media Segments contain media samples, such as audio or video, that are encrypted using the Advanced Encryption Standard [AES_128]." For fragmented MP4 segments this is where CMAF enters, in one sentence: "fMP4 Media Segments are encrypted using the 'cbcs' scheme of Common Encryption [COMMON_ENC]." That one reference is the entire sourced basis this guide has for CMAF; a scheme-by-scheme comparison is outside what RFC 8216 supports.

The IV attribute matters more than its short definition suggests: "The value is a hexadecimal-sequence that specifies a 128-bit unsigned integer Initialization Vector to be used with the key." Reuse weakens the cipher, per the same specification: "[AES_128] REQUIRES the same 16-octet IV to be supplied when encrypting and decrypting. Varying this IV increases the strength of the cipher." Under the identity key format with no IV attribute present, the Media Sequence Number itself becomes the IV, predictable across the whole stream rather than random per segment, exactly what the specification describes happening by default rather than a defect in some implementation.

The key request is an authorization decision

The URI attribute is where the standards basis for authorization begins and ends. The specification states it: "The value is a quoted-string containing a URI that specifies how to obtain the key. This attribute is REQUIRED unless the METHOD is NONE." That sentence says how to obtain a key, nothing about who may. A Key file is defined just as plainly: "An EXT-X-KEY tag with a URI attribute identifies a Key file. A Key file contains a cipher key that can decrypt Media Segments in the Playlist." Possession of that file is total, for everything the tag's scope covers: "If a Media Playlist file contains an EXT-X-KEY tag that specifies a Key file URI, the client can obtain that Key file and use the key inside it to decrypt all Media Segments to which that EXT-X-KEY tag applies." Nothing in the syntax distinguishes a request from an entitled subscriber and a request from a script that just parsed the manifest.

Apple's own player documentation puts the burden exactly where RFC 8216 leaves it, on the client and on whatever the client talks to. Apple's HTTP Live Streaming documentation states that the client is responsible for fetching decryption keys, authenticating or presenting a UI for authentication and decrypting media files as needed. Read that as an architecture statement: the key request is a separate, authenticatable HTTP transaction, and it can carry a session token, an entitlement check or a per-user key endpoint the way any authenticated API call can, because RFC 8216 never says it cannot.

Building that endpoint is Pharos Production practice: issue a short-lived, single-purpose token when the player requests the manifest, bind the key request to that token and the viewer's session and check entitlement on every key fetch rather than once at manifest time, since the manifest and the key request are two different requests a client can replay independently. That stateless-API pattern, validating a short-lived token against a session store on every call, is the same one our backend architecture guide covers for any high-throughput authenticated endpoint.

Signing, caching and the gaps between them

Nothing in RFC 8216 or RFC 9111 describes a signed-URL format, a token-validation scheme or a CDN edge rule, so this section is Pharos Production delivery practice.

RFC 8216 itself invites the practice the rest of this section argues against. The specification text states plainly that "The server MAY set the HTTP Expires header in the key response to indicate the duration for which the key can be cached." That is a permission to cache, not a promise of protection: the sentence says nothing about which cache, shared or private, may hold the response, or whose subsequent request that cached copy may then answer. A platform that sets Expires on a key response and stops there has taken the specification up on its invitation without also setting Cache-Control: private and Vary on the same response, which is exactly the failure mode the rest of this section describes.

The security considerations section does note that a playlist is itself a URI-fetch surface a client will exploit however it can: "Playlist files contain URIs, which clients will use to make network requests of arbitrary entities. Clients SHOULD range-check responses to prevent buffer overflows."

Signing only the manifest is the common mistake in HLS delivery: a signed or bearer-token manifest stops an unauthorized client from ever seeing the playlist, but the segment URLs written inside that playlist are usually plain, unsigned CDN paths. Once a client has fetched one manifest, those segment URLs work for anyone who copies them, regardless of whether the manifest itself was protected.

Signing segment URLs individually closes that leak: each request carries its own short-lived token, checked at the edge before the segment is served. It closes nothing about the key, though, since a signed segment sitting beside an open key endpoint is the same leak one hop later. Fetch access and decrypt access are two different questions, and a platform that answers only one has answered the easier half.

RFC 9111 is Internet Standards Track, STD 98, and it governs whether a CDN edge node may reuse a key response across two viewers: a shared cache, the kind a CDN edge runs, serves more than one user, while a private cache, the kind a player runs, is dedicated to a single one. The gap worth naming is that the rule governing an authenticated key request is narrower than it looks. The RFC 9111 text states: "A shared cache MUST NOT use a cached response to a request with an Authorization header field (Section 11.6.2 of [HTTP]) to satisfy any subsequent request unless the response contains a Cache-Control field with a response directive (Section 5.2.2) that allows it to be stored by a shared cache, and the cache conforms to the requirements of that directive for that response." Read plainly, that protection is keyed on the Authorization header field specifically: a key endpoint that authenticates by a query-string token, a cookie or a custom header gets none of it, because the condition that triggers the rule never fires.

Two directives close the gap for an implementation that does not rely on the Authorization header. The same document narrows the cache key with Vary: "When a cache receives a request that can be satisfied by a stored response and that stored response contains a Vary header field (Section 12.5.5 of [HTTP]), the cache MUST NOT use that stored response without revalidation unless all the presented request header fields nominated by that Vary field value match those fields in the original request (i.e., the request that caused the cached response to be stored)." Private keeps a shared cache out entirely while still letting the player's own cache hold the response: "The unqualified private response directive indicates that a shared cache MUST NOT store the response (i.e., the response is intended for a single user). It also indicates that a private cache MAY store the response, subject to the constraints defined in Section 3, even if the response would not otherwise be heuristically cacheable by a private cache."

The weakest directive, and the one most often mistaken for a guarantee, is no-store, which the same specification defines by saying that "The no-store response directive indicates that a cache MUST NOT store any part of either the immediate request or the response and MUST NOT use the response to satisfy any other request." It warns against exactly the reading a platform team reaches for: "This directive is not a reliable or sufficient mechanism for ensuring privacy. In particular, malicious or compromised caches might not recognize or obey this directive, and communications networks might be vulnerable to eavesdropping." Treat it as a directive a compliant cache honors, never as proof a key response was never stored anywhere.

Playlist types, concurrency and geo restriction as engineering practice

Two playlist types set two different expectations for a client. The specification states both in one sentence: "If the EXT-X-PLAYLIST-TYPE value is EVENT, Media Segments can only be added to the end of the Media Playlist. If the EXT-X-PLAYLIST-TYPE value is Video On Demand (VOD), the Media Playlist cannot change." A client keeps reloading a live playlist unless it is VOD or an ended EVENT playlist, a correctness rule a caching layer can invert when a per-user manifest is treated like a static VOD file. The spec's one caching hint for a segment about to disappear is the sourced anchor for a related failure mode: a server planning to remove a segment "SHOULD ensure that the HTTP response contains an Expires header that reflects the planned time-to-live." A manifest cached longer than the segments it references actually live means every new viewer lands pointed at segments the origin has already removed, a failure that looks like a CDN outage rather than the mismatched TTL that caused it. One older tag, EXT-X-ALLOW-CACHE, was removed in version 7 and is named only so it is not mistaken for a live control.

RFC 8216 has almost nothing to say about concurrency, session tracking or geo restriction. What little it says covers client behavior, not platform enforcement: a client should keep its concurrent download count small to avoid contributing to a denial of service condition, and requests carrying session cookies must follow ordinary cookie restriction and expiry rules. The two controls a platform actually relies on, concurrency limiting and regional entitlement, have no basis in either RFC and live entirely in the authorization layer already described. A concurrent-session limit is enforced by the key server, tracking which sessions hold a valid key for an account and refusing a request once the account is over its limit. Geo restriction works the same way: the key endpoint resolves the requester's location, usually from the IP address on the key request, and answers or refuses accordingly.

The DRM boundary: what AES-128 does not give you

A general-purpose key server sitting beside a separate licensed hardware security module in a media server rack.

AES-128 makes an unauthenticated fetch useless. It does not make a stream un-copyable once a legitimate viewer has decrypted it, and RFC 8216 never claims otherwise: its cryptographic apparatus stops at handing an entitled client a key, with no output-protection language of any kind, no secure video path, no screen-capture control and no device-binding concept anywhere in the document. That is the line this guide calls the DRM boundary. Below it, encryption plus an entitlement check stops the casual case; above it, once content is decrypted inside a player the platform does not fully control, only a licensed DRM system offers anything more.

RFC 8216 itself points at that boundary, allowing a server to offer the same segment under more than one key format at once, since the spec confirms that "A Media Segment can only be encrypted with one encryption METHOD, using one encryption key and IV. However, a server MAY offer multiple ways to retrieve that key by providing multiple EXT-X-KEY tags, each with a different KEYFORMAT attribute value." That is the mechanism a multi-DRM deployment uses: one segment, several KEYFORMAT tags, one per licensing system. Apple's own HTTP Live Streaming documentation locates keys inside the index file a client already has to load: "The index file, in turn, specifies the location of the available media files, decryption keys, and any alternate streams available" and for an ongoing broadcast the client keeps reloading that index and picking up new keys: "During ongoing broadcasts, load a new version of the index file periodically. Look for new media files and encryption keys in the updated index and add these URLs to the playback queue."

FairPlay Streaming is Apple's own answer to the boundary this section describes, named here as a platform fact rather than a recommendation. Apple's FairPlay Streaming page states its purpose: "Secure the delivery of streaming media to devices through the HTTP Live Streaming (HLS) protocol" and frames the collaboration it enables plainly: "Using FairPlay Streaming (FPS) technology, content providers, encoding vendors, and delivery networks can encrypt content, securely exchange keys, and protect playback on Apple platforms." Running it is a commitment rather than a checkbox: "The FairPlay Streaming Server SDK contains an implementation guide, reference information, and development keys for Key Server Module (KSM) implementors" and production access is gated on the business itself: "Your request will be approved only if your team provides a streaming service to consumers." Widevine and PlayReady are the other two platform DRM systems in wide use, named only so a reader knows where they fit; neither vendor's documentation was consulted here, so nothing above is a comparison among the three.

One stream traced end to end

A working deployment chains every control above: a session-bound signed manifest, signed segments that expire the same way, a key request carrying the session token and checked against entitlement and a key response marked Cache-Control: private with Vary on the session header, so the edge never reuses it across viewers even without an Authorization header. What follows is where the chain breaks in practice, Pharos Production experience rather than anything either RFC requires. A key response cached under the origin's default TTL serves one viewer's key to the next. A repeated IV under a fixed key weakens the cipher against exactly the analysis AES was chosen to resist. A token TTL shorter than the segment duration makes a player refetch mid-segment and stall. Clock skew between the token issuer and the edge rejects a token that is still valid, or accepts one already expired. A manifest cached across users hands every viewer the same signed segment URLs regardless of entitlement, and a CDN checking a signed URL on a normal request but not a range request lets a client bypass the whole scheme.

How Pharos Production helps

Where the DRM boundary sits for a given catalog is a decision a platform team can make for itself, not a default it inherits from whichever control was easiest to add. Four questions settle it.

What is actually at risk, and for how long. A live pay-per-view event or a title still inside a theatrical or first-window license carries a real, current commercial cost if a decrypted copy circulates; a back-catalog title already available on ad-supported services elsewhere does not carry that cost, whatever its encryption looks like on paper.

What devices the catalog has to reach. FairPlay, Widevine and PlayReady between them cover the device families a platform is likely to target, but each is a build-and-run commitment against its own SDK and its own approval process, not a setting added to an existing key server.

Whether output protection is a contractual requirement rather than an engineering preference. A content license that mandates HDCP or a secure video path settles the question outright, because AES-128 alone cannot satisfy that clause: RFC 8216 carries no output-protection language of any kind. Where no license makes that demand, the question stays open to a judgment about cost against risk.

What the loss actually costs against what running a DRM system costs. A working key endpoint with a real entitlement check, the control this guide spends most of its length on, stops the casual case for a fraction of the engineering effort a licensed DRM integration requires. A platform that has not gotten that control right yet has cheaper work to do before a DRM system adds anything.

Our streaming software development guide covers the wider pipeline this sits inside, the transcoding ladder, adaptive bitrate packaging and the platform choices around delivery. Our streaming software development team builds the key server, the signing layer and the entitlement checks this guide describes, on top of whatever transcoding and CDN stack a catalog already runs.

Sources: RFC 8216, HTTP Live Streaming, on rfc-editor.org, also published on the IETF Datatracker; RFC 9111, HTTP Caching, on rfc-editor.org; Apple Inc. developer documentation for HTTP Live Streaming and for FairPlay Streaming, on developer.apple.com. Read on 21 September 2026. Engineering guidance, not legal advice.

FAQ

Last updated:

Quick answers to common questions about custom software development, pricing, process and technology.

  • Does HLS encryption stop screen recording or stream ripping?

    No, and RFC 8216 never claims it does. AES-128 and SAMPLE-AES protect a segment in transit and at rest against anyone who cannot obtain the key.

    Once a legitimate, entitled viewer's player has decrypted a segment for playback, the specification offers no output protection at all: no secure video path, no HDCP requirement, no screen-capture control. Stopping capture after decryption is a platform DRM question, specifically what FairPlay, Widevine or PlayReady enforce on the device, not a question HLS's own encryption tags answer. A team that expects AES-128 alone to stop ripping is expecting a control the specification was never asked to provide.

  • Should the key endpoint live on the same domain as the CDN?

    Our practice at Pharos Production is to keep it separate, though no source read for this guide says the specification requires that. A key endpoint that shares an origin or a caching layer with segment delivery inherits that layer's default cache behavior unless every response is configured individually, and a shared cache serving segments has every incentive to cache aggressively.

    A dedicated, low-traffic key service with an explicit Cache-Control and Vary policy on every response is easier to audit than one code path that also has to serve gigabytes of video.

  • How short should a key request token actually be?

    Short enough that a copied token is worthless before anyone can reuse it, which in our experience means minutes rather than hours for a live event and a session-length window for on-demand playback. The token should be bound to the session that requested the manifest, not just to the account, so a token copied out of one device's network traffic does not work from a second device.

    None of this is stated in RFC 8216, which never mentions a token lifetime at all; it is Pharos Production practice built entirely on top of the one sentence the spec offers about securing key delivery.

  • What changes for a live event versus an on-demand catalog?

    The key rotation cadence and the entitlement check both get stricter for live. An on-demand title can afford one key for its whole runtime, because the content is not time-sensitive once it airs; a live event, especially a pay-per-view one, benefits from rotating keys within the broadcast so a key leaked mid-event does not cover the whole thing.

    Session and concurrency limits matter more for live too, since the value of an account-sharing workaround is highest in the first minutes of a broadcast. On-demand content leans harder on signed segment URLs with longer expiries and less on rapid key rotation.

  • Is SAMPLE-AES more secure than full-segment AES-128?

    Not necessarily, and the two answer different engineering questions rather than different security levels. AES-128 encrypts the whole segment; SAMPLE-AES encrypts the samples inside it and leaves the container readable, which some players and some fragmented MP4 workflows need for reasons unrelated to security, such as parsing timing information without decrypting first.

    Neither method is described in RFC 8216 as stronger than the other. The choice is a compatibility decision with the packaging pipeline and the target player, made alongside the security posture rather than instead of it.

  • What is the single most common way secure HLS delivery actually fails?

    In our experience it is a manifest or a key response caught by the CDN's default caching behavior rather than by an explicit policy, because that failure is invisible until someone notices one viewer's session working on a device that never authenticated. A signed segment scheme, a rotating key and a DRM decision can all be correct and the deployment still leaks, because nobody set Cache-Control and Vary on the one response that carries the actual secret.

    The fix is rarely a new control; it is auditing which responses in the pipeline inherited a cache policy nobody chose on purpose.

I work with startup founders who need a dedicated software development team but don’t want to gamble on hiring, random outsourcing, or opaque delivery.
Most founders face the same problem sooner or later.
Early technical and team decisions lock the product into tech debt, slow delivery, missed milestones and constant re-hiring. By the time this becomes visible, fixing it is already expensive.

As a CTO and software architect, I help founders design, build and run dedicated development teams that work as a true extension of the startup. Not as a black-box vendor.

My focus is on complex products where mistakes are costly:

  • Web3 and blockchain platforms
  • FinTech and regulated products
  • High-load startup systems
  • MVP → scale transitions

We don’t do body-shopping.
We don’t sell generic outsourcing.

Instead, we help founders:

  • build the right team structure from day one
  • keep technical ownership and transparency
  • scale delivery without losing control
  • avoid vendor lock-in and hidden risks

Teams are aligned with the product roadmap, business goals and long-term architecture. Not just short-term velocity.

Dmytro Nasyrov, Founder and CTO at Pharos Production
Dmytro Nasyrov Founder & CTO Let's work together!

Your business results matter

Achieve them with minimized risk through our bespoke innovation capabilities

Your contact details
Please enter your name
Please enter a valid email address
Please enter your message

We use your details only to reply to your request. Data Privacy and Legal Notice

We typically reply within 24 hours

What happens next?

  1. Contact us

    Contact us today to discuss your project. We're ready to review your request promptly and guide you on the best next steps for collaboration

    Same day
  2. NDA

    We're committed to keeping your information confidential, so we'll sign a Non-Disclosure Agreement

    1 day
  3. Plan the Goals

    After we chat about your goals and needs, we'll craft a comprehensive proposal detailing the project scope, team, timeline and budget

    3-5 days
  4. Finalize the Details

    Let's connect on Google Meet to go through the proposal and confirm all the details together!

    1-2 days
  5. Sign the Contract

    As soon as the contract is signed, our dedicated team will jump into action on your project!

    Same day