Mediabunny 1.49.0 · Vanilagy · MPL-2.0 Unmodified source distributed with the installed package. https://www.npmjs.com/package/mediabunny/v/1.49.0 Mozilla Public License Version 2.0 ================================== 1. Definitions -------------- 1.1. "Contributor" means each individual or legal entity that creates, contributes to the creation of, or owns Covered Software. 1.2. "Contributor Version" means the combination of the Contributions of others (if any) used by a Contributor and that particular Contributor's Contribution. 1.3. "Contribution" means Covered Software of a particular Contributor. 1.4. "Covered Software" means Source Code Form to which the initial Contributor has attached the notice in Exhibit A, the Executable Form of such Source Code Form, and Modifications of such Source Code Form, in each case including portions thereof. 1.5. "Incompatible With Secondary Licenses" means (a) that the initial Contributor has attached the notice described in Exhibit B to the Covered Software; or (b) that the Covered Software was made available under the terms of version 1.1 or earlier of the License, but not also under the terms of a Secondary License. 1.6. "Executable Form" means any form of the work other than Source Code Form. 1.7. "Larger Work" means a work that combines Covered Software with other material, in a separate file or files, that is not Covered Software. 1.8. "License" means this document. 1.9. "Licensable" means having the right to grant, to the maximum extent possible, whether at the time of the initial grant or subsequently, any and all of the rights conveyed by this License. 1.10. "Modifications" means any of the following: (a) any file in Source Code Form that results from an addition to, deletion from, or modification of the contents of Covered Software; or (b) any new file in Source Code Form that contains any Covered Software. 1.11. "Patent Claims" of a Contributor means any patent claim(s), including without limitation, method, process, and apparatus claims, in any patent Licensable by such Contributor that would be infringed, but for the grant of the License, by the making, using, selling, offering for sale, having made, import, or transfer of either its Contributions or its Contributor Version. 1.12. "Secondary License" means either the GNU General Public License, Version 2.0, the GNU Lesser General Public License, Version 2.1, the GNU Affero General Public License, Version 3.0, or any later versions of those licenses. 1.13. "Source Code Form" means the form of the work preferred for making modifications. 1.14. "You" (or "Your") means an individual or a legal entity exercising rights under this License. For legal entities, "You" includes any entity that controls, is controlled by, or is under common control with You. For purposes of this definition, "control" means (a) the power, direct or indirect, to cause the direction or management of such entity, whether by contract or otherwise, or (b) ownership of more than fifty percent (50%) of the outstanding shares or beneficial ownership of such entity. 2. License Grants and Conditions -------------------------------- 2.1. Grants Each Contributor hereby grants You a world-wide, royalty-free, non-exclusive license: (a) under intellectual property rights (other than patent or trademark) Licensable by such Contributor to use, reproduce, make available, modify, display, perform, distribute, and otherwise exploit its Contributions, either on an unmodified basis, with Modifications, or as part of a Larger Work; and (b) under Patent Claims of such Contributor to make, use, sell, offer for sale, have made, import, and otherwise transfer either its Contributions or its Contributor Version. 2.2. Effective Date The licenses granted in Section 2.1 with respect to any Contribution become effective for each Contribution on the date the Contributor first distributes such Contribution. 2.3. Limitations on Grant Scope The licenses granted in this Section 2 are the only rights granted under this License. No additional rights or licenses will be implied from the distribution or licensing of Covered Software under this License. Notwithstanding Section 2.1(b) above, no patent license is granted by a Contributor: (a) for any code that a Contributor has removed from Covered Software; or (b) for infringements caused by: (i) Your and any other third party's modifications of Covered Software, or (ii) the combination of its Contributions with other software (except as part of its Contributor Version); or (c) under Patent Claims infringed by Covered Software in the absence of its Contributions. This License does not grant any rights in the trademarks, service marks, or logos of any Contributor (except as may be necessary to comply with the notice requirements in Section 3.4). 2.4. Subsequent Licenses No Contributor makes additional grants as a result of Your choice to distribute the Covered Software under a subsequent version of this License (see Section 10.2) or under the terms of a Secondary License (if permitted under the terms of Section 3.3). 2.5. Representation Each Contributor represents that the Contributor believes its Contributions are its original creation(s) or it has sufficient rights to grant the rights to its Contributions conveyed by this License. 2.6. Fair Use This License is not intended to limit any rights You have under applicable copyright doctrines of fair use, fair dealing, or other equivalents. 2.7. Conditions Sections 3.1, 3.2, 3.3, and 3.4 are conditions of the licenses granted in Section 2.1. 3. Responsibilities ------------------- 3.1. Distribution of Source Form All distribution of Covered Software in Source Code Form, including any Modifications that You create or to which You contribute, must be under the terms of this License. You must inform recipients that the Source Code Form of the Covered Software is governed by the terms of this License, and how they can obtain a copy of this License. You may not attempt to alter or restrict the recipients' rights in the Source Code Form. 3.2. Distribution of Executable Form If You distribute Covered Software in Executable Form then: (a) such Covered Software must also be made available in Source Code Form, as described in Section 3.1, and You must inform recipients of the Executable Form how they can obtain a copy of such Source Code Form by reasonable means in a timely manner, at a charge no more than the cost of distribution to the recipient; and (b) You may distribute such Executable Form under the terms of this License, or sublicense it under different terms, provided that the license for the Executable Form does not attempt to limit or alter the recipients' rights in the Source Code Form under this License. 3.3. Distribution of a Larger Work You may create and distribute a Larger Work under terms of Your choice, provided that You also comply with the requirements of this License for the Covered Software. If the Larger Work is a combination of Covered Software with a work governed by one or more Secondary Licenses, and the Covered Software is not Incompatible With Secondary Licenses, this License permits You to additionally distribute such Covered Software under the terms of such Secondary License(s), so that the recipient of the Larger Work may, at their option, further distribute the Covered Software under the terms of either this License or such Secondary License(s). 3.4. Notices You may not remove or alter the substance of any license notices (including copyright notices, patent notices, disclaimers of warranty, or limitations of liability) contained within the Source Code Form of the Covered Software, except that You may alter any license notices to the extent required to remedy known factual inaccuracies. 3.5. Application of Additional Terms You may choose to offer, and to charge a fee for, warranty, support, indemnity or liability obligations to one or more recipients of Covered Software. However, You may do so only on Your own behalf, and not on behalf of any Contributor. You must make it absolutely clear that any such warranty, support, indemnity, or liability obligation is offered by You alone, and You hereby agree to indemnify every Contributor for any liability incurred by such Contributor as a result of warranty, support, indemnity or liability terms You offer. You may include additional disclaimers of warranty and limitations of liability specific to any jurisdiction. 4. Inability to Comply Due to Statute or Regulation --------------------------------------------------- If it is impossible for You to comply with any of the terms of this License with respect to some or all of the Covered Software due to statute, judicial order, or regulation then You must: (a) comply with the terms of this License to the maximum extent possible; and (b) describe the limitations and the code they affect. Such description must be placed in a text file included with all distributions of the Covered Software under this License. Except to the extent prohibited by statute or regulation, such description must be sufficiently detailed for a recipient of ordinary skill to be able to understand it. 5. Termination -------------- 5.1. The rights granted under this License will terminate automatically if You fail to comply with any of its terms. However, if You become compliant, then the rights granted under this License from a particular Contributor are reinstated (a) provisionally, unless and until such Contributor explicitly and finally terminates Your grants, and (b) on an ongoing basis, if such Contributor fails to notify You of the non-compliance by some reasonable means prior to 60 days after You have come back into compliance. Moreover, Your grants from a particular Contributor are reinstated on an ongoing basis if such Contributor notifies You of the non-compliance by some reasonable means, this is the first time You have received notice of non-compliance with this License from such Contributor, and You become compliant prior to 30 days after Your receipt of the notice. 5.2. If You initiate litigation against any entity by asserting a patent infringement claim (excluding declaratory judgment actions, counter-claims, and cross-claims) alleging that a Contributor Version directly or indirectly infringes any patent, then the rights granted to You by any and all Contributors for the Covered Software under Section 2.1 of this License shall terminate. 5.3. In the event of termination under Sections 5.1 or 5.2 above, all end user license agreements (excluding distributors and resellers) which have been validly granted by You or Your distributors under this License prior to termination shall survive termination. ************************************************************************ * * * 6. Disclaimer of Warranty * * ------------------------- * * * * Covered Software is provided under this License on an "as is" * * basis, without warranty of any kind, either expressed, implied, or * * statutory, including, without limitation, warranties that the * * Covered Software is free of defects, merchantable, fit for a * * particular purpose or non-infringing. The entire risk as to the * * quality and performance of the Covered Software is with You. * * Should any Covered Software prove defective in any respect, You * * (not any Contributor) assume the cost of any necessary servicing, * * repair, or correction. This disclaimer of warranty constitutes an * * essential part of this License. No use of any Covered Software is * * authorized under this License except under this disclaimer. * * * ************************************************************************ ************************************************************************ * * * 7. Limitation of Liability * * -------------------------- * * * * Under no circumstances and under no legal theory, whether tort * * (including negligence), contract, or otherwise, shall any * * Contributor, or anyone who distributes Covered Software as * * permitted above, be liable to You for any direct, indirect, * * special, incidental, or consequential damages of any character * * including, without limitation, damages for lost profits, loss of * * goodwill, work stoppage, computer failure or malfunction, or any * * and all other commercial damages or losses, even if such party * * shall have been informed of the possibility of such damages. This * * limitation of liability shall not apply to liability for death or * * personal injury resulting from such party's negligence to the * * extent applicable law prohibits such limitation. Some * * jurisdictions do not allow the exclusion or limitation of * * incidental or consequential damages, so this exclusion and * * limitation may not apply to You. * * * ************************************************************************ 8. Litigation ------------- Any litigation relating to this License may be brought only in the courts of a jurisdiction where the defendant maintains its principal place of business and such litigation shall be governed by laws of that jurisdiction, without reference to its conflict-of-law provisions. Nothing in this Section shall prevent a party's ability to bring cross-claims or counter-claims. 9. Miscellaneous ---------------- This License represents the complete agreement concerning the subject matter hereof. If any provision of this License is held to be unenforceable, such provision shall be reformed only to the extent necessary to make it enforceable. Any law or regulation which provides that the language of a contract shall be construed against the drafter shall not be used to construe this License against a Contributor. 10. Versions of the License --------------------------- 10.1. New Versions Mozilla Foundation is the license steward. Except as provided in Section 10.3, no one other than the license steward has the right to modify or publish new versions of this License. Each version will be given a distinguishing version number. 10.2. Effect of New Versions You may distribute the Covered Software under the terms of the version of the License under which You originally received the Covered Software, or under the terms of any subsequent version published by the license steward. 10.3. Modified Versions If you create software not governed by this License, and you want to create a new license for such software, you may create and use a modified version of this License if you rename the license and remove any references to the name of the license steward (except to note that such modified license differs from this License). 10.4. Distributing Source Code Form that is Incompatible With Secondary Licenses If You choose to distribute Source Code Form that is Incompatible With Secondary Licenses under the terms of this version of the License, the notice described in Exhibit B of this License must be attached. Exhibit A - Source Code Form License Notice ------------------------------------------- This Source Code Form is subject to the terms of the Mozilla Public License, v. 2.0. If a copy of the MPL was not distributed with this file, You can obtain one at https://mozilla.org/MPL/2.0/. If it is not possible or desirable to put the notice in a particular file, then You may include the notice in a location (such as a LICENSE file in a relevant directory) where a recipient would be likely to look for such a notice. You may add additional accurate notices of copyright ownership. Exhibit B - "Incompatible With Secondary Licenses" Notice --------------------------------------------------------- This Source Code Form is "Incompatible With Secondary Licenses", as defined by the Mozilla Public License, v. 2.0. ===== src/output.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { assert, AsyncMutex, EventEmitter, isIso639Dash2LanguageCode, MaybePromise, Rotation, toArray } from './misc'; import { MetadataTags, TrackDisposition, validateMetadataTags, validateTrackDisposition } from './metadata'; import { Muxer } from './muxer'; import { OutputFormat } from './output-format'; import { AudioSource, MediaSource, SubtitleSource, VideoSource } from './media-source'; import { PathedTarget, Target, TargetRequest } from './target'; import { Writer } from './writer'; import { Logging } from './logging'; /** * List of all track types. * @group Miscellaneous * @public */ export const ALL_TRACK_TYPES = ['video', 'audio', 'subtitle'] as const; /** * Union type of all track types. * @group Miscellaneous * @public */ export type TrackType = typeof ALL_TRACK_TYPES[number]; /** * Represents a track added to an {@link Output}. * @group Output files * @public */ export abstract class OutputTrack { /** @internal */ readonly id: number; /** The {@link Output} this track belongs to. */ readonly output: Output; /** The type of this track. */ readonly type: TrackType; /** The media source providing data for this track. */ readonly source: MediaSource; /** The metadata associated with this track. */ readonly metadata: BaseTrackMetadata; /** @internal */ protected constructor( id: number, output: Output, type: TrackType, source: MediaSource, metadata: BaseTrackMetadata, ) { this.id = id; this.output = output; this.type = type; this.source = source; this.metadata = metadata; } /** Returns true if and only if this track is a video track. */ isVideoTrack(): this is OutputVideoTrack { return this.type === 'video'; } /** Returns true if and only if this track is an audio track. */ isAudioTrack(): this is OutputAudioTrack { return this.type === 'audio'; } /** Returns true if and only if this track is a subtitle track. */ isSubtitleTrack(): this is OutputSubtitleTrack { return this.type === 'subtitle'; } /** * Returns true if and only if this track can be paired with the given other track. Pairability can be set using * the {@link BaseTrackMetadata.group} option. */ canBePairedWith(other: OutputTrack) { if (!(other instanceof OutputTrack)) { throw new TypeError('other must be an OutputTrack.'); } if (this === other) { return false; } const thisGroups = toArray(this.metadata.group!); const otherGroups = toArray(other.metadata.group!); for (const aGroup of thisGroups) { const pairableInSameGroup = this.type !== other.type && otherGroups.some(bGroup => aGroup === bGroup); if (pairableInSameGroup) { return true; } const pairableAcrossGroups = otherGroups.some( bGroup => aGroup._pairedGroups.has(bGroup), ); if (pairableAcrossGroups) { return true; } } return false; } } /** * An {@link OutputTrack} providing video data, created using {@link Output.addVideoTrack}. * @group Output files * @public */ export class OutputVideoTrack extends OutputTrack { declare readonly type: 'video'; declare readonly source: VideoSource; declare readonly metadata: VideoTrackMetadata; /** @internal */ constructor(id: number, output: Output, source: VideoSource, metadata: VideoTrackMetadata) { super(id, output, 'video', source, metadata); } } /** * An {@link OutputTrack} providing audio data, created using {@link Output.addAudioTrack}. * @group Output files * @public */ export class OutputAudioTrack extends OutputTrack { declare readonly type: 'audio'; declare readonly source: AudioSource; declare readonly metadata: AudioTrackMetadata; /** @internal */ constructor(id: number, output: Output, source: AudioSource, metadata: AudioTrackMetadata) { super(id, output, 'audio', source, metadata); } } /** * An {@link OutputTrack} providing subtitle data, created using {@link Output.addSubtitleTrack}. * @group Output files * @public */ export class OutputSubtitleTrack extends OutputTrack { declare readonly type: 'subtitle'; declare readonly source: SubtitleSource; declare readonly metadata: SubtitleTrackMetadata; /** @internal */ constructor(id: number, output: Output, source: SubtitleSource, metadata: SubtitleTrackMetadata) { super(id, output, 'subtitle', source, metadata); } } /** * Used to define pairability between {@link OutputTrack} instances. First create the group, then assign tracks to it * via {@link BaseTrackMetadata.group}. * * Two tracks are considered _pairable_ if they are in the same group but have a different {@link TrackType}, or if they * are in different groups that are paired with each other. Groups can be paired with each other using the * {@link OutputTrackGroup.pairWith} method. * * @group Output files * @public */ export class OutputTrackGroup { /** @internal */ _pairedGroups = new Set(); /** Creates a new {@link OutputTrackGroup}. */ constructor() { // The object's identity is the state } /** * Marks this group as being pairable with another group, symmetrically. Output tracks where each track is assigned * to one half of a group pairing are then considered pairable. * * You cannot pair a group with itself. */ pairWith(other: OutputTrackGroup) { if (!(other instanceof OutputTrackGroup)) { throw new TypeError('other must be an OutputTrackGroup.'); } if (this === other) { throw new TypeError('Cannot pair a group with itself.'); } this._pairedGroups.add(other); other._pairedGroups.add(this); } } /** * Base track metadata, applicable to all tracks. * @group Output files * @public */ export type BaseTrackMetadata = { /** The three-letter, ISO 639-2/T language code specifying the language of this track. */ languageCode?: string; /** A user-defined name for this track, like "English" or "Director Commentary". */ name?: string; /** The track's disposition, i.e. information about its intended usage. */ disposition?: Partial; /** * The maximum amount of encoded packets that will be added to this track. Setting this field provides the muxer * with an additional signal that it can use to preallocate space in the file. * * When this field is set, it is an error to provide more packets than whatever this field specifies. * * Predicting the maximum packet count requires considering both the maximum duration as well as the codec. * - For video codecs, you can assume one packet per frame. * - For audio codecs, there is one packet for each "audio chunk", the duration of which depends on the codec. For * simplicity, you can assume each packet is roughly 10 ms or 512 samples long, whichever is shorter. * - For subtitles, assume each cue and each gap in the subtitles adds a packet. * * If you're not fully sure, make sure to add a buffer of around 33% to make sure you stay below the maximum. */ maximumPacketCount?: number; /** * Whether the timestamps of this track are relative to the Unix epoch (January 1, 1970, 00:00:00 UTC). When `true`, * each timestamp maps to a definitive point in time. */ isRelativeToUnixEpoch?: boolean; /** * Defines the group(s) this track is a part of. Group assignment determines track pairability, determining which * tracks can be presented together with other tracks. This is needed for configuring things like HLS master * playlists. * * Two groups are considered pairable if they are in the same group but are of different {@link TrackType}, or if * they are in two separate groups that have been paired with each other. * * If left blank, a track is automatically assigned to {@link Output.defaultTrackGroup}. */ group?: OutputTrackGroup | OutputTrackGroup[]; }; /** * Additional metadata for video tracks. * @group Output files * @public */ export type VideoTrackMetadata = BaseTrackMetadata & { /** The angle in degrees by which the track's frames should be rotated (clockwise). */ rotation?: Rotation; /** * The expected video frame rate in hertz. If set, all timestamps and durations of this track will be snapped to * this frame rate. You should avoid adding more frames than the rate allows, as this will lead to multiple frames * with the same timestamp. */ frameRate?: number; /** * When true, this track is marked as being made only out of key frames (I-frames). It is an error to add a non-key * frame to this track. */ hasOnlyKeyPackets?: boolean; }; /** * Additional metadata for audio tracks. * @group Output files * @public */ export type AudioTrackMetadata = BaseTrackMetadata & {}; /** * Additional metadata for subtitle tracks. * @group Output files * @public */ export type SubtitleTrackMetadata = BaseTrackMetadata & {}; const validateBaseTrackMetadata = (metadata: BaseTrackMetadata) => { if (!metadata || typeof metadata !== 'object') { throw new TypeError('metadata must be an object.'); } if (metadata.languageCode !== undefined && !isIso639Dash2LanguageCode(metadata.languageCode)) { throw new TypeError('metadata.languageCode, when provided, must be a three-letter, ISO 639-2/T language code.'); } if (metadata.name !== undefined && typeof metadata.name !== 'string') { throw new TypeError('metadata.name, when provided, must be a string.'); } if (metadata.disposition !== undefined) { validateTrackDisposition(metadata.disposition); } if ( metadata.maximumPacketCount !== undefined && (!Number.isInteger(metadata.maximumPacketCount) || metadata.maximumPacketCount < 0) ) { throw new TypeError('metadata.maximumPacketCount, when provided, must be a non-negative integer.'); } if ( metadata.group !== undefined && !(metadata.group instanceof OutputTrackGroup) && (!Array.isArray(metadata.group) || metadata.group.some(group => !(group instanceof OutputTrackGroup))) ) { throw new TypeError( 'metadata.group, when provided, must be an OutputTrackGroup instance or an array of' + ' OutputTrackGroup instances.', ); } }; /** * The options for creating an Output object. * @group Output files * @public */ export type OutputOptions< F extends OutputFormat = OutputFormat, T extends Target = Target, > = { /** The format of the output file. */ format: F; /** The target to which the file will be written. */ target: T | PathedTarget; /** * Optional; the target to which the track initialization data will be written. Most formats do not make use of * this, but some do, such as {@link CmafOutputFormat}. * * When this is a function, it will only be called if an init target is needed. */ initTarget?: T | (() => MaybePromise); /** * Optional; a callback to be called at the end of {@link Output.finalize}. Can be used to run logic once the * output has completed. If a promise is returned, it will be awaited internally by {@link Output.finalize}. */ onFinalize?: () => MaybePromise; }; /** * Describes the events that an {@link Output} emits, with each key being an event name and its value being the * event data. * * @group Output files * @public */ export type OutputEvents = { /** Emitted whenever a {@link Target} is obtained by the output. Useful to track writes. */ target: { /** The target that was obtained. */ target: Target; /** The request that led to the target being obtained, or `null` if the output is not pathed. */ request: TargetRequest | null; /** Whether the target is the root file of the media. */ isRoot: boolean; }; }; /** * Main class orchestrating the creation of new media files. * @group Output files * @public */ export class Output< F extends OutputFormat = OutputFormat, T extends Target = Target, > extends EventEmitter { /** The format of the output file. */ readonly format: F; /** @internal */ _target: T | PathedTarget; /** The current state of the output. */ state: 'pending' | 'started' | 'canceled' | 'finalizing' | 'finalized' = 'pending'; /** * The {@link OutputTrackGroup} that all tracks are assigned to by default unless otherwise specified by * {@link BaseTrackMetadata.group}. */ readonly defaultTrackGroup = new OutputTrackGroup(); /** @internal */ private _initTarget: T | (() => MaybePromise) | null; /** @internal */ _onFinalize: (() => MaybePromise) | null = null; /** @internal */ _muxer: Muxer; /** @internal */ _unfinalizedTargets = new Set(); /** @internal */ _rootWriterPromise: Promise | null = null; /** @internal */ _tracks: OutputTrack[] = []; /** @internal */ _startPromise: Promise | null = null; /** @internal */ _cancelPromise: Promise | null = null; /** @internal */ _finalizePromise: Promise | null = null; /** @internal */ _mutex = new AsyncMutex(); /** @internal */ _metadataTags: MetadataTags = {}; /** @internal */ _rootTarget: T | null = null; /** @internal */ _rootTargetPromise: Promise | null = null; /** * This field is used to synchronize multiple MediaStreamTracks. They use the same time coordinate system across * tracks, and to ensure correct audio-video sync, we must use the same offset for all of them. The reason an offset * is needed at all is because the timestamps typically don't start at zero. * @internal */ _firstMediaStreamTimestamp: number | null = null; /** * The target to which the root file will be written. Throws when using {@link PathedTarget} with an async callback; * prefer the `'target'` event for those cases. */ get target(): T { const errorMessage = 'Output.target cannot be used when using PathedTarget with an async callback.' + ' Use the \'target\' event instead.'; // We use this field to make sure we can reliably throw in the `target` getter whenever retrieving the target // requires awaiting a promise. We do this so there is no different behavior based on order: if the target has // already been retrieved via the normal internal operations, and then somebody calls the `target` getter, even // if the target is now available, the getter should still throw to be consistent in behavior and in definition. if (this._rootTargetPromise) { throw new TypeError(errorMessage); } const rootTargetResult = this._getRootTarget(); if (rootTargetResult instanceof Promise) { throw new TypeError(errorMessage); } return rootTargetResult; } /** * Creates a new instance of {@link Output} which can then be used to create a new media file according to the * specified {@link OutputOptions}. */ constructor(options: OutputOptions) { super(); if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (!(options.format instanceof OutputFormat)) { throw new TypeError('options.format must be an OutputFormat.'); } if (!(options.target instanceof Target || options.target instanceof PathedTarget)) { throw new TypeError('options.target must be a Target or a PathedTarget.'); } if (options.target instanceof Target) { this._rememberTarget(options.target); } if ( options.initTarget !== undefined && !(options.initTarget instanceof Target) && typeof options.initTarget !== 'function' ) { throw new Error( 'options.initTarget, when provided, must be a Target or a function that returns or resolves to' + ' a Target.', ); } if (options.onFinalize !== undefined && typeof options.onFinalize !== 'function') { throw new TypeError('options.onFinalize, when provided, must be a function.'); } this.format = options.format; this._target = options.target; this._onFinalize = options.onFinalize ?? null; this._initTarget = options.initTarget ?? null; if (this._initTarget instanceof Target) { this._rememberTarget(this._initTarget); } this._muxer = options.format._createMuxer(this); } /** @internal */ _getTargetValidated(request: TargetRequest): MaybePromise { assert(this._target instanceof PathedTarget); const result = this._target.getTarget(request); const handleResult = (result: T) => { if (!(result instanceof Target)) { throw new TypeError('getTarget must return a Target.'); } return result; }; if (result instanceof Promise) { return result.then(handleResult); } else { return handleResult(result); } } /** @internal */ async _getTarget(request: TargetRequest) { assert(this._target instanceof PathedTarget); const target = await this._getTargetValidated(request); this._emit('target', { target, request, isRoot: request.isRoot }); if (this.state === 'canceled') { await target._close(); } else { this._rememberTarget(target); } return target; } /** @internal */ _rememberTarget(target: Target) { this._unfinalizedTargets.add(target); target.on('finalized', () => this._unfinalizedTargets.delete(target), { once: true }); } /** @internal */ async _getInitTarget(): Promise { assert(this._initTarget !== null); if (this._initTarget instanceof Target) { return this._initTarget; } const target = await this._initTarget(); if (this.state === 'canceled') { await target._close(); } else { this._rememberTarget(target); } return target; } /** @internal */ _hasInitTarget() { return this._initTarget !== null; } /** @internal */ _getRootTarget(): MaybePromise { if (this._rootTarget) { return this._rootTarget; } if (this._rootTargetPromise) { return this._rootTargetPromise; } if (this._target instanceof Target) { this._emit('target', { target: this._target, request: null, isRoot: true }); this._rootTarget = this._target; return this._target; } const request: TargetRequest = { path: this._target.rootPath, isRoot: true, mimeType: this.format.mimeType, }; const result = this._getTargetValidated(request); const handleResult = (target: T) => { if (this.state === 'canceled') { // Promise thrown away here, but no way to surface it to the user really void target._close(); } else { this._rememberTarget(target); } this._emit('target', { target, request, isRoot: true }); this._rootTarget = target; return target; }; if (result instanceof Promise) { return this._rootTargetPromise = result.then(handleResult); } else { return handleResult(result); } } /** @internal */ _getRootWriter(isMonotonic: boolean | ((target: Target) => boolean)) { return this._rootWriterPromise ??= (async () => { const target = await this._getRootTarget(); const writer = new Writer(target, typeof isMonotonic === 'boolean' ? isMonotonic : isMonotonic(target)); writer.start(); return writer; })(); } /** Adds a video track to the output with the given source. Can only be called before the output is started. */ addVideoTrack(source: VideoSource, metadata: VideoTrackMetadata = {}) { if (!(source instanceof VideoSource)) { throw new TypeError('source must be a VideoSource.'); } validateBaseTrackMetadata(metadata); if (metadata.rotation !== undefined && ![0, 90, 180, 270].includes(metadata.rotation)) { throw new TypeError(`Invalid video rotation: ${metadata.rotation}. Has to be 0, 90, 180 or 270.`); } if (!this.format.supportsVideoRotationMetadata && metadata.rotation) { throw new Error(`${this.format._name} does not support video rotation metadata.`); } if ( metadata.frameRate !== undefined && (!Number.isFinite(metadata.frameRate) || metadata.frameRate <= 0) ) { throw new TypeError( `Invalid video frame rate: ${metadata.frameRate}. Must be a positive number.`, ); } const metadataCopy = { ...metadata }; metadataCopy.group ??= this.defaultTrackGroup; return this._addTrack(new OutputVideoTrack( this._tracks.length + 1, this, source, metadataCopy, )); } /** Adds an audio track to the output with the given source. Can only be called before the output is started. */ addAudioTrack(source: AudioSource, metadata: AudioTrackMetadata = {}) { if (!(source instanceof AudioSource)) { throw new TypeError('source must be an AudioSource.'); } validateBaseTrackMetadata(metadata); const metadataCopy = { ...metadata }; metadataCopy.group ??= this.defaultTrackGroup; return this._addTrack(new OutputAudioTrack( this._tracks.length + 1, this, source, metadataCopy, )); } /** Adds a subtitle track to the output with the given source. Can only be called before the output is started. */ addSubtitleTrack(source: SubtitleSource, metadata: SubtitleTrackMetadata = {}) { if (!(source instanceof SubtitleSource)) { throw new TypeError('source must be a SubtitleSource.'); } validateBaseTrackMetadata(metadata); const metadataCopy = { ...metadata }; metadataCopy.group ??= this.defaultTrackGroup; return this._addTrack(new OutputSubtitleTrack( this._tracks.length + 1, this, source, metadataCopy, )); } /** * Sets descriptive metadata tags about the media file, such as title, author, date, or cover art. When called * multiple times, only the metadata from the last call will be used. * * Can only be called before the output is started. */ setMetadataTags(tags: MetadataTags) { validateMetadataTags(tags); if (this.state !== 'pending') { throw new Error('Cannot set metadata tags after output has been started or canceled.'); } this._metadataTags = tags; } /** @internal */ private _addTrack(track: T) { if (this.state !== 'pending') { throw new Error('Cannot add track after output has been started or canceled.'); } if (track.source._connectedTrack) { throw new Error('Source is already used for a track.'); } // Verify maximum track count constraints const supportedTrackCounts = this.format.getSupportedTrackCounts(); const presentTracksOfThisType = this._tracks.reduce( (count, t) => count + (t.type === track.type ? 1 : 0), 0, ); const maxCount = supportedTrackCounts[track.type].max; if (presentTracksOfThisType === maxCount) { throw new Error( maxCount === 0 ? `${this.format._name} does not support ${track.type} tracks.` : (`${this.format._name} does not support more than ${maxCount} ${track.type} track` + `${maxCount === 1 ? '' : 's'}.`), ); } const maxTotalCount = supportedTrackCounts.total.max; if (this._tracks.length === maxTotalCount) { throw new Error( `${this.format._name} does not support more than ${maxTotalCount} tracks` + `${maxTotalCount === 1 ? '' : 's'} in total.`, ); } if (track.isVideoTrack()) { const supportedVideoCodecs = this.format.getSupportedVideoCodecs(); if (supportedVideoCodecs.length === 0) { throw new Error( `${this.format._name} does not support video tracks.` + this.format._codecUnsupportedHint(track.source._codec), ); } else if (!supportedVideoCodecs.includes(track.source._codec)) { throw new Error( `Codec '${track.source._codec}' cannot be contained within ${this.format._name}. Supported` + ` video codecs are: ${supportedVideoCodecs.map(codec => `'${codec}'`).join(', ')}.` + this.format._codecUnsupportedHint(track.source._codec), ); } } else if (track.isAudioTrack()) { const supportedAudioCodecs = this.format.getSupportedAudioCodecs(); if (supportedAudioCodecs.length === 0) { throw new Error( `${this.format._name} does not support audio tracks.` + this.format._codecUnsupportedHint(track.source._codec), ); } else if (!supportedAudioCodecs.includes(track.source._codec)) { throw new Error( `Codec '${track.source._codec}' cannot be contained within ${this.format._name}. Supported` + ` audio codecs are: ${supportedAudioCodecs.map(codec => `'${codec}'`).join(', ')}.` + this.format._codecUnsupportedHint(track.source._codec), ); } } else if (track.isSubtitleTrack()) { const supportedSubtitleCodecs = this.format.getSupportedSubtitleCodecs(); if (supportedSubtitleCodecs.length === 0) { throw new Error( `${this.format._name} does not support subtitle tracks.` + this.format._codecUnsupportedHint(track.source._codec), ); } else if (!supportedSubtitleCodecs.includes(track.source._codec)) { throw new Error( `Codec '${track.source._codec}' cannot be contained within ${this.format._name}. Supported` + ` subtitle codecs are: ${supportedSubtitleCodecs.map(codec => `'${codec}'`).join(', ')}.` + this.format._codecUnsupportedHint(track.source._codec), ); } } this._tracks.push(track); track.source._connectedTrack = track; return track; } /** * Starts the creation of the output file. This method should be called after all tracks have been added. Only after * the output has started can media samples be added to the tracks. * * @returns A promise that resolves when the output has successfully started and is ready to receive media samples. */ async start() { // Verify minimum track count constraints const supportedTrackCounts = this.format.getSupportedTrackCounts(); for (const trackType of ALL_TRACK_TYPES) { const presentTracksOfThisType = this._tracks.reduce( (count, track) => count + (track.type === trackType ? 1 : 0), 0, ); const minCount = supportedTrackCounts[trackType].min; if (presentTracksOfThisType < minCount) { throw new Error( minCount === supportedTrackCounts[trackType].max ? (`${this.format._name} requires exactly ${minCount} ${trackType}` + ` track${minCount === 1 ? '' : 's'}.`) : (`${this.format._name} requires at least ${minCount} ${trackType}` + ` track${minCount === 1 ? '' : 's'}.`), ); } } const totalMinCount = supportedTrackCounts.total.min; if (this._tracks.length < totalMinCount) { throw new Error( totalMinCount === supportedTrackCounts.total.max ? (`${this.format._name} requires exactly ${totalMinCount} track` + `${totalMinCount === 1 ? '' : 's'}.`) : (`${this.format._name} requires at least ${totalMinCount} track` + `${totalMinCount === 1 ? '' : 's'}.`), ); } if (this.state === 'canceled') { throw new Error('Output has been canceled.'); } if (this._startPromise) { Logging._warn('Output has already been started.'); return this._startPromise; } return this._startPromise = (async () => { this.state = 'started'; const release = await this._mutex.acquire(); try { await this._muxer.start(); const promises = this._tracks.map(track => track.source._start()); await Promise.all(promises); } finally { release(); } })(); } /** * Resolves with the full MIME type of the output file, including track codecs. * * The returned promise will resolve only once the precise codec strings of all tracks are known. */ getMimeType() { return this._muxer.getMimeType(); } /** * Cancels the creation of the output file, releasing internal resources like encoders and preventing further * samples from being added. * * @returns A promise that resolves once all internal resources have been released. */ async cancel() { if (this._cancelPromise) { Logging._warn('Output has already been canceled.'); return this._cancelPromise; } else if (this.state === 'finalizing' || this.state === 'finalized') { // Don't wanna warn when finalizing since that shows a warning when finalization fails and then cancel // is called if (this.state === 'finalized') { Logging._warn('Output has already been finalized.'); } return; } return this._cancelPromise = (async () => { this.state = 'canceled'; const release = await this._mutex.acquire(); try { const promises = this._tracks.map(x => x.source._flushOrWaitForOngoingClose(true)); // Force close await Promise.all(promises); await Promise.all([...this._unfinalizedTargets].map(target => target._close())); this._unfinalizedTargets.clear(); } finally { release(); } })(); } /** * Finalizes the output file. This method must be called after all media samples across all tracks have been added. * Once the Promise returned by this method completes, the output file is ready. */ async finalize() { if (this.state === 'pending') { throw new Error('Cannot finalize before starting.'); } if (this.state === 'canceled') { throw new Error('Cannot finalize after canceling.'); } if (this._finalizePromise) { Logging._warn('Output has already been finalized.'); return this._finalizePromise; } return this._finalizePromise = (async () => { this.state = 'finalizing'; const release = await this._mutex.acquire(); try { const promises = this._tracks.map(x => x.source._flushOrWaitForOngoingClose(false)); await Promise.all(promises); await this._muxer.finalize(); if (this._rootWriterPromise) { const rootWriter = await this._rootWriterPromise; if (!rootWriter.finalized) { await rootWriter.flush(); await rootWriter.finalize(); } } if (this._onFinalize) { await this._onFinalize(); } this.state = 'finalized'; } finally { release(); } })(); } } ===== src/media-sink.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { parsePcmCodec, PCM_AUDIO_CODECS, PcmAudioCodec, VideoCodec, AudioCodec } from './codec'; import { AvcNalUnitType, concatAvcNalUnits, deserializeAvcDecoderConfigurationRecord, determineVideoPacketType, extractNalUnitTypeForAvc, extractNalUnitTypeForHevc, HevcNalUnitType, iterateAvcNalUnits, iterateHevcNalUnits, parseAvcSps, sanitizeHevcPacketForChromium, } from './codec-data'; import { CustomVideoDecoder, customVideoDecoders, CustomAudioDecoder, customAudioDecoders } from './custom-coder'; import { InputDisposedError } from './input'; import { InputAudioTrack, InputTrack, InputVideoTrack } from './input-track'; import { AnyIterable, assert, assertNever, CallSerializer, getInt24, getUint24, insertSorted, isChromium, isFirefox, isNumber, isWebKit, last, mapAsyncGenerator, promiseWithResolvers, Rotation, toAsyncIterator, toDataView, toUint8Array, validateAnyIterable, } from './misc'; import { EncodedPacket } from './packet'; import { fromAlaw, fromUlaw } from './pcm'; import { AudioSample, clampCropRectangle, CropRectangle, validateCropRectangle, VideoSample, VideoSamplePixelFormat, } from './sample'; import { Logging } from './logging'; /** * Additional options for controlling packet retrieval. * @group Media sinks * @public */ export type PacketRetrievalOptions = { /** * When set to `true`, only packet metadata (like timestamp) will be retrieved - the actual packet data will not * be loaded. */ metadataOnly?: boolean; /** * When set to `true`, key packets will be verified upon retrieval by looking into the packet's bitstream. * If not enabled, the packet types will be determined solely by what's stored in the containing file and may be * incorrect, potentially leading to decoder errors. Since determining a packet's actual type requires looking into * its data, this option cannot be enabled together with `metadataOnly`. */ verifyKeyPackets?: boolean; /** * When querying packets in live media that are in the future relative to the current live edge, Mediabunny will, * by default, wait for the stream to advance until the query can be satisfied. In a sense, Mediabunny simply treats * live streams as media files that are still being written, and any read that depends on future information will * wait until it can be fulfilled. * * If you want to query packets based only on the currently known information, set this field to `true` - this way, * Mediabunny will never wait for the live stream to catch up. * * For non-live media, this field has no effect. */ skipLiveWait?: boolean; }; const validatePacketRetrievalOptions = (options: PacketRetrievalOptions) => { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.metadataOnly !== undefined && typeof options.metadataOnly !== 'boolean') { throw new TypeError('options.metadataOnly, when defined, must be a boolean.'); } if (options.verifyKeyPackets !== undefined && typeof options.verifyKeyPackets !== 'boolean') { throw new TypeError('options.verifyKeyPackets, when defined, must be a boolean.'); } if (options.verifyKeyPackets && options.metadataOnly) { throw new TypeError('options.verifyKeyPackets and options.metadataOnly cannot be enabled together.'); } if (options.skipLiveWait !== undefined && typeof options.skipLiveWait !== 'boolean') { throw new TypeError('options.skipLiveWait, when defined, must be a boolean.'); } }; const validateTimestamp = (timestamp: number) => { if (!isNumber(timestamp)) { throw new TypeError('timestamp must be a number.'); // It can be non-finite, that's fine } }; const maybeFixPacketType = ( track: InputTrack, promise: Promise, options: PacketRetrievalOptions, ) => { if (options.verifyKeyPackets) { return promise.then(async (packet) => { if (!packet || packet.type === 'delta') { return packet; } const determinedType = await track.determinePacketType(packet); if (determinedType) { // @ts-expect-error Technically readonly packet.type = determinedType; } return packet; }); } else { return promise; } }; /** * Sink for retrieving encoded packets from an input track. * @group Media sinks * @public */ export class EncodedPacketSink { /** @internal */ _track: InputTrack; /** Creates a new {@link EncodedPacketSink} for the given {@link InputTrack}. */ constructor(track: InputTrack) { if (!(track instanceof InputTrack)) { throw new TypeError('track must be an InputTrack.'); } this._track = track; } /** * Retrieves the track's first packet (in decode order), or null if it has no packets. The first packet is very * likely to be a key packet, but it doesn't have to be. */ async getFirstPacket(options: PacketRetrievalOptions = {}) { validatePacketRetrievalOptions(options); if (this._track.input._disposed) { throw new InputDisposedError(); } return maybeFixPacketType(this._track, this._track._backing.getFirstPacket(options), options); } /** Retrieves the track's first key packet (in decode order), or null if it has no key packets. */ async getFirstKeyPacket(options: PacketRetrievalOptions = {}) { validatePacketRetrievalOptions(options); const firstPacket = await this.getFirstPacket(options); if (!firstPacket) { return null; } if (firstPacket.type === 'key') { // Great return firstPacket; } return this.getNextKeyPacket(firstPacket, options); } /** * Retrieves the packet corresponding to the given timestamp, in seconds. More specifically, returns the last packet * (in presentation order) with a start timestamp less than or equal to the given timestamp. This method can be * used to retrieve a track's last packet using `getPacket(Infinity)`. The method returns null if the timestamp * is before the first packet in the track. * * @param timestamp - The timestamp used for retrieval, in seconds. */ async getPacket(timestamp: number, options: PacketRetrievalOptions = {}) { validateTimestamp(timestamp); validatePacketRetrievalOptions(options); if (this._track.input._disposed) { throw new InputDisposedError(); } return maybeFixPacketType(this._track, this._track._backing.getPacket(timestamp, options), options); } /** * Retrieves the packet following the given packet (in decode order), or null if the given packet is the * last packet. */ async getNextPacket(packet: EncodedPacket, options: PacketRetrievalOptions = {}) { if (!(packet instanceof EncodedPacket)) { throw new TypeError('packet must be an EncodedPacket.'); } validatePacketRetrievalOptions(options); if (this._track.input._disposed) { throw new InputDisposedError(); } return maybeFixPacketType(this._track, this._track._backing.getNextPacket(packet, options), options); } /** * Retrieves the key packet corresponding to the given timestamp, in seconds. More specifically, returns the last * key packet (in presentation order) with a start timestamp less than or equal to the given timestamp. A key packet * is a packet that doesn't require previous packets to be decoded. This method can be used to retrieve a track's * last key packet using `getKeyPacket(Infinity)`. The method returns null if the timestamp is before the first * key packet in the track. * * To ensure that the returned packet is guaranteed to be a real key frame, enable `options.verifyKeyPackets`. * * @param timestamp - The timestamp used for retrieval, in seconds. */ async getKeyPacket(timestamp: number, options: PacketRetrievalOptions = {}): Promise { validateTimestamp(timestamp); validatePacketRetrievalOptions(options); if (this._track.input._disposed) { throw new InputDisposedError(); } if (!options.verifyKeyPackets) { return this._track._backing.getKeyPacket(timestamp, options); } const packet = await this._track._backing.getKeyPacket(timestamp, options); if (!packet) { return packet; } assert(packet.type === 'key'); const determinedType = await this._track.determinePacketType(packet); if (determinedType === 'delta') { // Try returning the previous key packet (in hopes that it's actually a key packet) return this.getKeyPacket(packet.timestamp - 1 / await this._track.getTimeResolution(), options); } return packet; } /** * Retrieves the key packet following the given packet (in decode order), or null if the given packet is the last * key packet. * * To ensure that the returned packet is guaranteed to be a real key frame, enable `options.verifyKeyPackets`. */ async getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions = {}): Promise { if (!(packet instanceof EncodedPacket)) { throw new TypeError('packet must be an EncodedPacket.'); } validatePacketRetrievalOptions(options); if (this._track.input._disposed) { throw new InputDisposedError(); } if (!options.verifyKeyPackets) { return this._track._backing.getNextKeyPacket(packet, options); } const nextPacket = await this._track._backing.getNextKeyPacket(packet, options); if (!nextPacket) { return nextPacket; } assert(nextPacket.type === 'key'); const determinedType = await this._track.determinePacketType(nextPacket); if (determinedType === 'delta') { // Try returning the next key packet (in hopes that it's actually a key packet) return this.getNextKeyPacket(nextPacket, options); } return nextPacket; } /** * Creates an async iterator that yields the packets in this track in decode order. To enable fast iteration, this * method will intelligently preload packets based on the speed of the consumer. * * @param startPacket - (optional) The packet from which iteration should begin. This packet will also be yielded. * @param endPacket - (optional) The packet at which iteration should end. This packet will _not_ be yielded. */ packets( startPacket?: EncodedPacket, endPacket?: EncodedPacket, options: PacketRetrievalOptions = {}, ): AsyncGenerator { if (startPacket !== undefined && !(startPacket instanceof EncodedPacket)) { throw new TypeError('startPacket must be an EncodedPacket.'); } if (startPacket !== undefined && startPacket.isMetadataOnly && !options?.metadataOnly) { throw new TypeError('startPacket can only be metadata-only if options.metadataOnly is enabled.'); } if (endPacket !== undefined && !(endPacket instanceof EncodedPacket)) { throw new TypeError('endPacket must be an EncodedPacket.'); } validatePacketRetrievalOptions(options); if (this._track.input._disposed) { throw new InputDisposedError(); } const packetQueue: EncodedPacket[] = []; let { promise: queueNotEmpty, resolve: onQueueNotEmpty } = promiseWithResolvers(); let { promise: queueDequeue, resolve: onQueueDequeue } = promiseWithResolvers(); let ended = false; let terminated = false; // This stores errors that are "out of band" in the sense that they didn't occur in the normal flow of this // method but instead in a different context. This error should not go unnoticed and must be bubbled up to // the consumer. let outOfBandError = null as Error | null; const timestamps: number[] = []; // The queue should always be big enough to hold 1 second worth of packets const maxQueueSize = () => Math.max(2, timestamps.length); // The following is the "pump" process that keeps pumping packets into the queue (async () => { let packet = startPacket ?? await this.getFirstPacket(options); while (packet && !terminated && !this._track.input._disposed) { if (endPacket && packet.sequenceNumber >= endPacket?.sequenceNumber) { break; } if (packetQueue.length > maxQueueSize()) { ({ promise: queueDequeue, resolve: onQueueDequeue } = promiseWithResolvers()); await queueDequeue; continue; } packetQueue.push(packet); onQueueNotEmpty(); ({ promise: queueNotEmpty, resolve: onQueueNotEmpty } = promiseWithResolvers()); packet = await this.getNextPacket(packet, options); } ended = true; onQueueNotEmpty(); })().catch((error: Error) => { if (!outOfBandError) { outOfBandError = error; onQueueNotEmpty(); } }); const track = this._track; return { async next() { while (true) { if (track.input._disposed) { throw new InputDisposedError(); } else if (terminated) { return { value: undefined, done: true }; } else if (outOfBandError) { throw outOfBandError; } else if (packetQueue.length > 0) { const value = packetQueue.shift()!; const now = performance.now(); timestamps.push(now); while (timestamps.length > 0 && now - timestamps[0]! >= 1000) { timestamps.shift(); } onQueueDequeue(); return { value, done: false }; } else if (ended) { return { value: undefined, done: true }; } else { await queueNotEmpty; } } }, async return() { terminated = true; onQueueDequeue(); onQueueNotEmpty(); return { value: undefined, done: true }; }, async throw(error) { throw error; }, [Symbol.asyncIterator]() { return this; }, }; } } abstract class DecoderWrapper< MediaSample extends VideoSample | AudioSample, > { constructor( public onSample: (sample: MediaSample) => unknown, public onError: (error: Error) => unknown, ) {} abstract getDecodeQueueSize(): number; abstract decode(packet: EncodedPacket): void; abstract flush(): Promise; abstract close(): void; } /** * Base class for decoded media sample sinks. * @group Media sinks * @public */ export abstract class BaseMediaSampleSink< MediaSample extends VideoSample | AudioSample, > { /** @internal */ abstract _track: InputTrack; /** @internal */ abstract _createDecoder( onSample: (sample: MediaSample) => unknown, onError: (error: Error) => unknown ): Promise>; /** @internal */ abstract _createPacketSink(): EncodedPacketSink; /** @internal */ protected mediaSamplesInRange( startTimestamp = -Infinity, endTimestamp = Infinity, options: PacketRetrievalOptions, ): AsyncGenerator { validateTimestamp(startTimestamp); validateTimestamp(endTimestamp); const sampleQueue: MediaSample[] = []; let firstSampleQueued = false; let lastSample: MediaSample | null = null; let { promise: queueNotEmpty, resolve: onQueueNotEmpty } = promiseWithResolvers(); let { promise: queueDequeue, resolve: onQueueDequeue } = promiseWithResolvers(); let decoderIsFlushed = false; let ended = false; let terminated = false; // This stores errors that are "out of band" in the sense that they didn't occur in the normal flow of this // method but instead in a different context. This error should not go unnoticed and must be bubbled up to // the consumer. let outOfBandError = null as Error | null; const packetRetrievalOptions: PacketRetrievalOptions = { ...options, verifyKeyPackets: true, metadataOnly: false, }; // The following is the "pump" process that keeps pumping packets into the decoder (async () => { const decoder = await this._createDecoder((sample) => { onQueueDequeue(); if (sample.timestamp >= endTimestamp) { ended = true; } if (ended) { sample.close(); return; } if (lastSample) { if (sample.timestamp > startTimestamp) { // We don't know ahead of time what the first first is. This is because the first first is the // last first whose timestamp is less than or equal to the start timestamp. Therefore we need to // wait for the first first after the start timestamp, and then we'll know that the previous // first was the first first. sampleQueue.push(lastSample); firstSampleQueued = true; } else { lastSample.close(); } } if (sample.timestamp >= startTimestamp) { sampleQueue.push(sample); firstSampleQueued = true; } lastSample = firstSampleQueued ? null : sample; if (sampleQueue.length > 0) { onQueueNotEmpty(); ({ promise: queueNotEmpty, resolve: onQueueNotEmpty } = promiseWithResolvers()); } }, (error) => { if (!outOfBandError) { outOfBandError = error; onQueueNotEmpty(); } }); const packetSink = this._createPacketSink(); const keyPacket = await packetSink.getKeyPacket(startTimestamp, packetRetrievalOptions) ?? await packetSink.getFirstKeyPacket(packetRetrievalOptions); let currentPacket: EncodedPacket | null = keyPacket; // B-frames make it exceedingly difficult to properly define an upper bound for packet iteration if an end // timestamp is set, so we just don't do it. The case that makes it especially tricky is when the frames // following a key frame have a lower timestamp than the keyframe; something that quite frequently happens // in HEVC streams. The price to pay for not upper-bounding the packet iterator is a slight increase in // decoder work at the end of the range, but the added correctness and reliability makes this tradeoff worth // it. const endPacket = undefined; const packets = packetSink.packets(keyPacket ?? undefined, endPacket, packetRetrievalOptions); await packets.next(); // Skip the start packet as we already have it while (currentPacket && !ended && !this._track.input._disposed) { const maxQueueSize = computeMaxQueueSize(sampleQueue.length); if (sampleQueue.length + decoder.getDecodeQueueSize() > maxQueueSize) { ({ promise: queueDequeue, resolve: onQueueDequeue } = promiseWithResolvers()); await queueDequeue; continue; } decoder.decode(currentPacket); const packetResult = await packets.next(); if (packetResult.done) { break; } currentPacket = packetResult.value; } await packets.return(); if (!terminated && !this._track.input._disposed) { await decoder.flush(); } decoder.close(); if (!firstSampleQueued && lastSample) { sampleQueue.push(lastSample); } decoderIsFlushed = true; onQueueNotEmpty(); // To unstuck the generator })().catch((error: Error) => { if (!outOfBandError) { outOfBandError = error; onQueueNotEmpty(); } }); const track = this._track; const closeSamples = () => { lastSample?.close(); for (const sample of sampleQueue) { sample.close(); } }; return { async next() { while (true) { if (track.input._disposed) { closeSamples(); throw new InputDisposedError(); } else if (terminated) { return { value: undefined, done: true }; } else if (outOfBandError) { closeSamples(); throw outOfBandError; } else if (sampleQueue.length > 0) { const value = sampleQueue.shift()!; onQueueDequeue(); return { value, done: false }; } else if (!decoderIsFlushed) { await queueNotEmpty; } else { return { value: undefined, done: true }; } } }, async return() { terminated = true; ended = true; onQueueDequeue(); onQueueNotEmpty(); closeSamples(); return { value: undefined, done: true }; }, async throw(error) { throw error; }, [Symbol.asyncIterator]() { return this; }, }; } /** @internal */ protected mediaSamplesAtTimestamps( timestamps: AnyIterable, options: PacketRetrievalOptions, ): AsyncGenerator { validateAnyIterable(timestamps); const timestampIterator = toAsyncIterator(timestamps); const timestampsOfInterest: number[] = []; const sampleQueue: (MediaSample | null)[] = []; let { promise: queueNotEmpty, resolve: onQueueNotEmpty } = promiseWithResolvers(); let { promise: queueDequeue, resolve: onQueueDequeue } = promiseWithResolvers(); let decoderIsFlushed = false; let terminated = false; // This stores errors that are "out of band" in the sense that they didn't occur in the normal flow of this // method but instead in a different context. This error should not go unnoticed and must be bubbled up to // the consumer. let outOfBandError = null as Error | null; const pushToQueue = (sample: MediaSample | null) => { sampleQueue.push(sample); onQueueNotEmpty(); ({ promise: queueNotEmpty, resolve: onQueueNotEmpty } = promiseWithResolvers()); }; const retrievalOptions: PacketRetrievalOptions = { ...options, verifyKeyPackets: true, metadataOnly: false, }; // The following is the "pump" process that keeps pumping packets into the decoder (async () => { const decoder = await this._createDecoder((sample) => { onQueueDequeue(); if (terminated) { sample.close(); return; } let sampleUses = 0; while ( timestampsOfInterest.length > 0 && sample.timestamp - timestampsOfInterest[0]! > -1e-10 // Give it a little epsilon ) { sampleUses++; timestampsOfInterest.shift(); } if (sampleUses > 0) { for (let i = 0; i < sampleUses; i++) { // Clone the sample if we need to emit it multiple times pushToQueue((i < sampleUses - 1 ? sample.clone() : sample) as MediaSample); } } else { sample.close(); } }, (error) => { if (!outOfBandError) { outOfBandError = error; onQueueNotEmpty(); } }); const packetSink = this._createPacketSink(); let lastPacket: EncodedPacket | null = null; let lastKeyPacket: EncodedPacket | null = null; // The end sequence number (inclusive) in the next batch of packets that will be decoded. The batch starts // at the last key frame and goes until this sequence number. let maxSequenceNumber = -1; const decodePackets = async () => { assert(lastKeyPacket); // Start at the current key packet let currentPacket = lastKeyPacket; decoder.decode(currentPacket); while (currentPacket.sequenceNumber < maxSequenceNumber) { const maxQueueSize = computeMaxQueueSize(sampleQueue.length); while (sampleQueue.length + decoder.getDecodeQueueSize() > maxQueueSize && !terminated) { ({ promise: queueDequeue, resolve: onQueueDequeue } = promiseWithResolvers()); await queueDequeue; } if (terminated) { break; } const nextPacket = await packetSink.getNextPacket(currentPacket, retrievalOptions); assert(nextPacket); decoder.decode(nextPacket); currentPacket = nextPacket; } maxSequenceNumber = -1; }; const flushDecoder = async () => { await decoder.flush(); // We don't expect this list to have any elements in it anymore, but in case it does, let's emit // nulls for every remaining element, then clear it. for (let i = 0; i < timestampsOfInterest.length; i++) { pushToQueue(null); } timestampsOfInterest.length = 0; }; for await (const timestamp of timestampIterator) { validateTimestamp(timestamp); if (terminated || this._track.input._disposed) { break; } const targetPacket = await packetSink.getPacket(timestamp, retrievalOptions); const keyPacket = targetPacket && await packetSink.getKeyPacket(timestamp, retrievalOptions); if (!keyPacket) { if (maxSequenceNumber !== -1) { await decodePackets(); await flushDecoder(); } pushToQueue(null); lastPacket = null; continue; } // Check if the key packet has changed or if we're going back in time if ( lastPacket && ( keyPacket.sequenceNumber !== lastKeyPacket!.sequenceNumber || targetPacket.timestamp < lastPacket.timestamp ) ) { await decodePackets(); await flushDecoder(); // Always flush here, improves decoder compatibility } timestampsOfInterest.push(targetPacket.timestamp); maxSequenceNumber = Math.max(targetPacket.sequenceNumber, maxSequenceNumber); lastPacket = targetPacket; lastKeyPacket = keyPacket; } if (!terminated && !this._track.input._disposed) { if (maxSequenceNumber !== -1) { // We still need to decode packets await decodePackets(); } await flushDecoder(); } decoder.close(); decoderIsFlushed = true; onQueueNotEmpty(); // To unstuck the generator })().catch((error: Error) => { if (!outOfBandError) { outOfBandError = error; onQueueNotEmpty(); } }); const track = this._track; const closeSamples = () => { for (const sample of sampleQueue) { sample?.close(); } }; return { async next() { while (true) { if (track.input._disposed) { closeSamples(); throw new InputDisposedError(); } else if (terminated) { return { value: undefined, done: true }; } else if (outOfBandError) { closeSamples(); throw outOfBandError; } else if (sampleQueue.length > 0) { const value = sampleQueue.shift(); assert(value !== undefined); onQueueDequeue(); return { value, done: false }; } else if (!decoderIsFlushed) { await queueNotEmpty; } else { return { value: undefined, done: true }; } } }, async return() { terminated = true; onQueueDequeue(); onQueueNotEmpty(); closeSamples(); return { value: undefined, done: true }; }, async throw(error) { throw error; }, [Symbol.asyncIterator]() { return this; }, }; } } const computeMaxQueueSize = (decodedSampleQueueSize: number) => { // If we have decoded samples lying around, limit the total queue size to a small value (decoded samples can use up // a lot of memory). If not, we're fine with a much bigger queue of encoded packets waiting to be decoded. In fact, // some decoders only start flushing out decoded chunks when the packet queue is large enough. return decodedSampleQueueSize === 0 ? 40 : 8; }; class VideoDecoderWrapper extends DecoderWrapper { decoder: VideoDecoder | null = null; customDecoder: CustomVideoDecoder | null = null; customDecoderCallSerializer = new CallSerializer(); customDecoderQueueSize = 0; inputTimestamps: number[] = []; // Timestamps input into the decoder, sorted. sampleQueue: VideoSample[] = []; // Safari-specific thing, check usage. currentPacketIndex = 0; raslSkipped = false; // For HEVC stuff // Alpha stuff alphaDecoder: VideoDecoder | null = null; alphaHadKeyframe = false; colorQueue: VideoFrame[] = []; alphaQueue: (VideoFrame | null)[] = []; merger: ColorAlphaMerger | null = null; decodedAlphaChunkCount = 0; alphaDecoderQueueSize = 0; /** Each value is the number of decoded alpha chunks at which a null alpha frame should be added. */ nullAlphaFrameQueue: number[] = []; currentAlphaPacketIndex = 0; alphaRaslSkipped = false; // For HEVC stuff frameHandlerSerializer = new CallSerializer(); constructor( onSample: (sample: VideoSample) => unknown, onError: (error: Error) => unknown, public codec: VideoCodec, public decoderConfig: VideoDecoderConfig, public rotation: Rotation, public timeResolution: number, ) { super(onSample, onError); const MatchingCustomDecoder = customVideoDecoders.find(x => x.supports(codec, decoderConfig)); if (MatchingCustomDecoder) { // @ts-expect-error "Can't create instance of abstract class 🤓" this.customDecoder = new MatchingCustomDecoder() as CustomVideoDecoder; // @ts-expect-error It's technically readonly this.customDecoder.codec = codec; // @ts-expect-error It's technically readonly this.customDecoder.config = decoderConfig; // @ts-expect-error It's technically readonly this.customDecoder.onSample = (sample) => { if (!(sample instanceof VideoSample)) { throw new TypeError('The argument passed to onSample must be a VideoSample.'); } this.finalizeAndEmitSample(sample); }; void this.customDecoderCallSerializer.call(() => this.customDecoder!.init()); } else { const colorHandler = (frame: VideoFrame) => { this.frameHandlerSerializer.call(async () => { if (this.alphaQueue.length > 0) { // Even when no alpha data is present (most of the time), there will be nulls in this queue const alphaFrame = this.alphaQueue.shift(); assert(alphaFrame !== undefined); await this.mergeAlpha(frame, alphaFrame); } else { this.colorQueue.push(frame); } }).catch((error: Error) => this.onError(error)); }; if (codec === 'avc' && this.decoderConfig.description && isChromium()) { // Chromium has/had a bug with playing interlaced AVC (https://issues.chromium.org/issues/456919096) // which can be worked around by requesting that software decoding be used. So, here we peek into the // AVC description, if present, and switch to software decoding if we find interlaced content. const record = deserializeAvcDecoderConfigurationRecord(toUint8Array(this.decoderConfig.description)); if (record && record.sequenceParameterSets.length > 0) { const sps = parseAvcSps(record.sequenceParameterSets[0]!); if (sps && sps.frameMbsOnlyFlag === 0) { this.decoderConfig = { ...this.decoderConfig, hardwareAcceleration: 'prefer-software', }; } } } const stack = new Error('Decoding error').stack; this.decoder = new VideoDecoder({ output: (frame) => { try { colorHandler(frame); } catch (error) { this.onError(error as Error); } }, error: (error) => { error.stack = stack; // Provide a more useful stack trace, the default one sucks this.onError(error); }, }); this.decoder.configure(this.decoderConfig); } } getDecodeQueueSize() { if (this.customDecoder) { return this.customDecoderQueueSize; } else { assert(this.decoder); return Math.max( this.decoder.decodeQueueSize, this.alphaDecoder?.decodeQueueSize ?? 0, ); } } decode(packet: EncodedPacket) { if (this.codec === 'hevc' && this.currentPacketIndex > 0 && !this.raslSkipped) { if (this.hasHevcRaslPicture(packet.data)) { return; // Drop } this.raslSkipped = true; } if (this.customDecoder) { this.customDecoderQueueSize++; void this.customDecoderCallSerializer .call(() => this.customDecoder!.decode(packet)) .then(() => this.customDecoderQueueSize--); } else { assert(this.decoder); if (!isWebKit()) { insertSorted(this.inputTimestamps, packet.timestamp, x => x); } if (isChromium() && this.currentPacketIndex === 0) { if (this.codec === 'avc') { // Workaround for https://issues.chromium.org/issues/470109459 const filteredNalUnits: Uint8Array[] = []; for (const loc of iterateAvcNalUnits(packet.data, this.decoderConfig)) { const type = extractNalUnitTypeForAvc(packet.data[loc.offset]!); if (type === AvcNalUnitType.AUD) { // If packets contain an AUD and have NALUs before it, this trips up Chromium's key frame // detector. Clear the NALUs if an AUD is encountered. // https://github.com/Vanilagy/mediabunny/issues/396 filteredNalUnits.length = 0; } // These trip up Chromium's key frame detection, so let's strip them if (!(type >= 20 && type <= 31)) { filteredNalUnits.push(packet.data.subarray(loc.offset, loc.offset + loc.length)); } } const newData = concatAvcNalUnits(filteredNalUnits, this.decoderConfig); packet = new EncodedPacket(newData, packet.type, packet.timestamp, packet.duration); } else if (this.codec === 'hevc') { // Workaround for https://issues.chromium.org/issues/507611247 const sanitizedData = sanitizeHevcPacketForChromium(packet.data, this.decoderConfig); if (sanitizedData) { packet = new EncodedPacket(sanitizedData, packet.type, packet.timestamp, packet.duration); } } } this.decoder.decode(packet.toEncodedVideoChunk()); this.decodeAlphaData(packet); } this.currentPacketIndex++; } decodeAlphaData(packet: EncodedPacket) { if (!packet.sideData.alpha) { // No alpha side data in the packet, most common case this.pushNullAlphaFrame(); return; } if (!this.merger) { this.merger = new ColorAlphaMerger(); } // Check if we need to set up the alpha decoder if (!this.alphaDecoder) { const alphaHandler = (frame: VideoFrame) => { this.frameHandlerSerializer.call(async () => { if (this.colorQueue.length > 0) { const colorFrame = this.colorQueue.shift(); assert(colorFrame !== undefined); await this.mergeAlpha(colorFrame, frame); } else { this.alphaQueue.push(frame); } // Check if any null frames have been queued for this point this.decodedAlphaChunkCount++; while ( this.nullAlphaFrameQueue.length > 0 && this.nullAlphaFrameQueue[0] === this.decodedAlphaChunkCount ) { this.nullAlphaFrameQueue.shift(); if (this.colorQueue.length > 0) { const colorFrame = this.colorQueue.shift(); assert(colorFrame !== undefined); await this.mergeAlpha(colorFrame, null); } else { this.alphaQueue.push(null); } } this.alphaDecoderQueueSize--; }).catch((error: Error) => this.onError(error)); }; const stack = new Error('Decoding error').stack; this.alphaDecoder = new VideoDecoder({ output: (frame) => { try { alphaHandler(frame); } catch (error) { this.onError(error as Error); } }, error: (error) => { error.stack = stack; // Provide a more useful stack trace, the default one sucks this.onError(error); }, }); this.alphaDecoder.configure(this.decoderConfig); } const type = determineVideoPacketType(this.codec, this.decoderConfig, packet.sideData.alpha); // Alpha packets might follow a different key frame rhythm than the main packets. Therefore, before we start // decoding, we must first find a packet that's actually a key frame. Until then, we treat the image as opaque. if (!this.alphaHadKeyframe) { this.alphaHadKeyframe = type === 'key'; } if (this.alphaHadKeyframe) { // Same RASL skipping logic as for color, unlikely to be hit (since who uses HEVC with separate alpha??) but // here for symmetry. if (this.codec === 'hevc' && this.currentAlphaPacketIndex > 0 && !this.alphaRaslSkipped) { if (this.hasHevcRaslPicture(packet.sideData.alpha)) { this.pushNullAlphaFrame(); return; } this.alphaRaslSkipped = true; } this.currentAlphaPacketIndex++; this.alphaDecoder.decode(packet.alphaToEncodedVideoChunk(type ?? packet.type)); this.alphaDecoderQueueSize++; } else { this.pushNullAlphaFrame(); } } pushNullAlphaFrame() { if (this.alphaDecoderQueueSize === 0) { // Easy this.alphaQueue.push(null); } else { // There are still alpha chunks being decoded, so pushing `null` immediately would result in out-of-order // data and be incorrect. Instead, we need to enqueue a "null frame" for when the current decoder workload // has finished. this.nullAlphaFrameQueue.push(this.decodedAlphaChunkCount + this.alphaDecoderQueueSize); } } /** * If we're using HEVC, we need to make sure to skip any RASL slices that follow a non-IDR key frame such as * CRA_NUT. This is because RASL slices cannot be decoded without data before the CRA_NUT. Browsers behave * differently here: Chromium drops the packets, Safari throws a decoder error. Either way, it's not good * and causes bugs upstream. So, let's take the dropping into our own hands. */ hasHevcRaslPicture(packetData: Uint8Array) { for (const loc of iterateHevcNalUnits(packetData, this.decoderConfig)) { const type = extractNalUnitTypeForHevc(packetData[loc.offset]!); if (type === HevcNalUnitType.RASL_N || type === HevcNalUnitType.RASL_R) { return true; } } return false; } /** Handler for the WebCodecs VideoDecoder for ironing out browser differences. */ sampleHandler(sample: VideoSample) { if (isWebKit()) { // For correct B-frame handling, we don't just hand over the frames directly but instead add them to // a queue, because we want to ensure frames are emitted in presentation order. We flush the queue // each time we receive a frame with a timestamp larger than the highest we've seen so far, as we // can sure that is not a B-frame. Typically, WebCodecs automatically guarantees that frames are // emitted in presentation order, but Safari doesn't always follow this rule. if (this.sampleQueue.length > 0 && (sample.timestamp >= last(this.sampleQueue)!.timestamp)) { for (const sample of this.sampleQueue) { this.finalizeAndEmitSample(sample); } this.sampleQueue.length = 0; } insertSorted(this.sampleQueue, sample, x => x.timestamp); } else { // Assign it the next earliest timestamp from the input. We do this because browsers, by spec, are // required to emit decoded frames in presentation order *while* retaining the timestamp of their // originating EncodedVideoChunk. For files with B-frames but no out-of-order timestamps (like a // missing ctts box, for example), this causes a mismatch. We therefore fix the timestamps and // ensure they are sorted by doing this. const timestamp = this.inputTimestamps.shift(); // There's no way we'd have more decoded frames than encoded packets we passed in. Actually, the // correspondence should be 1:1. assert(timestamp !== undefined); sample.setTimestamp(timestamp); this.finalizeAndEmitSample(sample); } } finalizeAndEmitSample(sample: VideoSample) { // Round the timestamps to the time resolution sample.setTimestamp(Math.round(sample.timestamp * this.timeResolution) / this.timeResolution); sample.setDuration(Math.round(sample.duration * this.timeResolution) / this.timeResolution); sample.setRotation(this.rotation); this.onSample(sample); } async mergeAlpha(color: VideoFrame, alpha: VideoFrame | null) { if (!alpha) { // Nothing needs to be merged const finalSample = new VideoSample(color); this.sampleHandler(finalSample); return; } assert(this.merger); // The merger takes ownership of the frames, so no need to close them ourselves const finalFrame = await this.merger.update(color, alpha); const finalSample = new VideoSample(finalFrame); this.sampleHandler(finalSample); } async flush() { if (this.customDecoder) { await this.customDecoderCallSerializer.call(() => this.customDecoder!.flush()); } else { assert(this.decoder); await Promise.all([ this.decoder.flush(), this.alphaDecoder?.flush(), ]); await this.frameHandlerSerializer.currentPromise; this.colorQueue.forEach(x => x.close()); this.colorQueue.length = 0; this.alphaQueue.forEach(x => x?.close()); this.alphaQueue.length = 0; this.alphaHadKeyframe = false; this.decodedAlphaChunkCount = 0; this.alphaDecoderQueueSize = 0; this.nullAlphaFrameQueue.length = 0; this.currentAlphaPacketIndex = 0; this.alphaRaslSkipped = false; } if (isWebKit()) { for (const sample of this.sampleQueue) { this.finalizeAndEmitSample(sample); } this.sampleQueue.length = 0; } this.currentPacketIndex = 0; this.raslSkipped = false; } close() { if (this.customDecoder) { void this.customDecoderCallSerializer.call(() => this.customDecoder!.close()); } else { assert(this.decoder); this.decoder.close(); this.alphaDecoder?.close(); this.colorQueue.forEach(x => x.close()); this.colorQueue.length = 0; this.alphaQueue.forEach(x => x?.close()); this.alphaQueue.length = 0; this.merger?.close(); } for (const sample of this.sampleQueue) { sample.close(); } this.sampleQueue.length = 0; } } let mergerGpuUnavailable = false; /** Utility class that merges together color and alpha information using simple WebGL 2 shaders. */ export class ColorAlphaMerger { static forceCpu = true; canvas: OffscreenCanvas | HTMLCanvasElement | null = null; private gl: WebGL2RenderingContext | null = null; private program: WebGLProgram | null = null; private vao: WebGLVertexArrayObject | null = null; private colorTexture: WebGLTexture | null = null; private alphaTexture: WebGLTexture | null = null; private worker: Worker | null = null; private pendingRequests = new Map>>(); private nextRequestId = 0; constructor() { const canMakeCanvas = typeof OffscreenCanvas !== 'undefined' // eslint-disable-next-line @typescript-eslint/no-deprecated || (typeof document !== 'undefined' && typeof document.createElement === 'function'); if (!ColorAlphaMerger.forceCpu && canMakeCanvas && !mergerGpuUnavailable) { // Try the GPU path. If anything goes wrong, we silently fall back to the CPU path. try { // Canvas will be resized later if (typeof OffscreenCanvas !== 'undefined') { // Prefer OffscreenCanvas for Worker environments this.canvas = new OffscreenCanvas(300, 150); } else { this.canvas = document.createElement('canvas'); } const gl = this.canvas.getContext('webgl2', { premultipliedAlpha: false, }) as unknown as WebGL2RenderingContext | null; // Casting because of some TypeScript weirdness if (!gl) { throw new Error('Couldn\'t acquire WebGL 2 context.'); } this.gl = gl; this.program = this.createProgram(); this.vao = this.createVAO(); this.colorTexture = this.createTexture(); this.alphaTexture = this.createTexture(); this.gl.useProgram(this.program); this.gl.uniform1i(this.gl.getUniformLocation(this.program, 'u_colorTexture'), 0); this.gl.uniform1i(this.gl.getUniformLocation(this.program, 'u_alphaTexture'), 1); } catch (error) { this.gl = null; this.canvas = null; mergerGpuUnavailable = true; Logging._warn('Falling back to CPU for color/alpha merging.', error); } } } async update(color: VideoFrame, alpha: VideoFrame): Promise { if (this.gl) { return this.updateGpu(color, alpha); } else { return this.updateCpu(color, alpha); } } private createProgram(): WebGLProgram { assert(this.gl); const vertexShader = this.createShader(this.gl.VERTEX_SHADER, `#version 300 es in vec2 a_position; in vec2 a_texCoord; out vec2 v_texCoord; void main() { gl_Position = vec4(a_position, 0.0, 1.0); v_texCoord = a_texCoord; } `); const fragmentShader = this.createShader(this.gl.FRAGMENT_SHADER, `#version 300 es precision highp float; uniform sampler2D u_colorTexture; uniform sampler2D u_alphaTexture; in vec2 v_texCoord; out vec4 fragColor; void main() { vec3 color = texture(u_colorTexture, v_texCoord).rgb; float alpha = texture(u_alphaTexture, v_texCoord).r; fragColor = vec4(color, alpha); } `); const program = this.gl.createProgram(); this.gl.attachShader(program, vertexShader); this.gl.attachShader(program, fragmentShader); this.gl.linkProgram(program); return program; } private createShader(type: number, source: string): WebGLShader { assert(this.gl); const shader = this.gl.createShader(type)!; this.gl.shaderSource(shader, source); this.gl.compileShader(shader); return shader; } private createVAO(): WebGLVertexArrayObject { assert(this.gl); assert(this.program); const vao = this.gl.createVertexArray(); this.gl.bindVertexArray(vao); const vertices = new Float32Array([ -1, -1, 0, 1, 1, -1, 1, 1, -1, 1, 0, 0, 1, 1, 1, 0, ]); const buffer = this.gl.createBuffer(); this.gl.bindBuffer(this.gl.ARRAY_BUFFER, buffer); this.gl.bufferData(this.gl.ARRAY_BUFFER, vertices, this.gl.STATIC_DRAW); const positionLocation = this.gl.getAttribLocation(this.program, 'a_position'); const texCoordLocation = this.gl.getAttribLocation(this.program, 'a_texCoord'); this.gl.enableVertexAttribArray(positionLocation); this.gl.vertexAttribPointer(positionLocation, 2, this.gl.FLOAT, false, 16, 0); this.gl.enableVertexAttribArray(texCoordLocation); this.gl.vertexAttribPointer(texCoordLocation, 2, this.gl.FLOAT, false, 16, 8); return vao; } private createTexture(): WebGLTexture { assert(this.gl); const texture = this.gl.createTexture(); this.gl.bindTexture(this.gl.TEXTURE_2D, texture); this.gl.texParameteri(this.gl.TEXTURE_2D, this.gl.TEXTURE_WRAP_S, this.gl.CLAMP_TO_EDGE); this.gl.texParameteri(this.gl.TEXTURE_2D, this.gl.TEXTURE_WRAP_T, this.gl.CLAMP_TO_EDGE); this.gl.texParameteri(this.gl.TEXTURE_2D, this.gl.TEXTURE_MIN_FILTER, this.gl.LINEAR); this.gl.texParameteri(this.gl.TEXTURE_2D, this.gl.TEXTURE_MAG_FILTER, this.gl.LINEAR); return texture; } private updateGpu(color: VideoFrame, alpha: VideoFrame): VideoFrame { assert(this.gl); assert(this.canvas); if (color.displayWidth !== this.canvas.width || color.displayHeight !== this.canvas.height) { this.canvas.width = color.displayWidth; this.canvas.height = color.displayHeight; } this.gl.activeTexture(this.gl.TEXTURE0); this.gl.bindTexture(this.gl.TEXTURE_2D, this.colorTexture); this.gl.texImage2D(this.gl.TEXTURE_2D, 0, this.gl.RGBA, this.gl.RGBA, this.gl.UNSIGNED_BYTE, color); this.gl.activeTexture(this.gl.TEXTURE1); this.gl.bindTexture(this.gl.TEXTURE_2D, this.alphaTexture); this.gl.texImage2D(this.gl.TEXTURE_2D, 0, this.gl.RGBA, this.gl.RGBA, this.gl.UNSIGNED_BYTE, alpha); this.gl.viewport(0, 0, this.canvas.width, this.canvas.height); this.gl.clear(this.gl.COLOR_BUFFER_BIT); this.gl.bindVertexArray(this.vao); this.gl.drawArrays(this.gl.TRIANGLE_STRIP, 0, 4); const finalFrame = new VideoFrame(this.canvas, { timestamp: color.timestamp, duration: color.duration ?? undefined, }); color.close(); alpha.close(); return finalFrame; } private updateCpu(color: VideoFrame, alpha: VideoFrame): Promise { if (!this.worker) { const blob = new Blob( [`(${colorAlphaMergerWorkerCode.toString()})()`], { type: 'application/javascript' }, ); const url = URL.createObjectURL(blob); this.worker = new Worker(url); URL.revokeObjectURL(url); this.worker.addEventListener('message', (event: MessageEvent) => { const data = event.data; const pending = this.pendingRequests.get(data.id); if (!pending) { return; } this.pendingRequests.delete(data.id); if ('error' in data) { pending.reject(new Error(data.error)); } else { pending.resolve(data.frame); } }); this.worker.addEventListener('error', (event) => { const error = new Error(event.message || 'Color/alpha merge worker error.'); for (const pending of this.pendingRequests.values()) { pending.reject(error); } this.pendingRequests.clear(); }); } const id = this.nextRequestId++; const pending = promiseWithResolvers(); this.pendingRequests.set(id, pending); this.worker.postMessage({ id, color, alpha }, { transfer: [color, alpha] }); return pending.promise; } close() { this.gl?.getExtension('WEBGL_lose_context')?.loseContext(); this.gl = null; this.canvas = null; this.worker?.terminate(); this.worker = null; const error = new Error('Color/alpha merger closed.'); for (const pending of this.pendingRequests.values()) { pending.reject(error); } this.pendingRequests.clear(); } } type ColorAlphaMergerWorkerRequest = { id: number; color: VideoFrame; alpha: VideoFrame; }; type ColorAlphaMergerWorkerResponse = | { id: number; frame: VideoFrame } | { id: number; error: string }; const colorAlphaMergerWorkerCode = () => { // These buffers are reused across frames as long as the size matches, since consecutive frames usually share // dimensions let cpuAlphaBuffer: Uint8Array | null = null; let cpuColorBuffer: Uint8Array | null = null; // Serialize execution internally so concurrent requests don't race on the shared cpu*Buffer state. let chain: Promise = Promise.resolve(); self.addEventListener('message', (event: MessageEvent) => { const { id, color, alpha } = event.data; chain = chain.then(async () => { try { const frame = await merge(color, alpha); self.postMessage({ id, frame }, { transfer: [frame] }); } catch (error) { self.postMessage({ id, error: (error as Error).message }); } finally { // We took ownership of the inputs via transfer; close them now that the merge (or its error) is done. color.close(); alpha.close(); } }); }); const merge = async (color: VideoFrame, alpha: VideoFrame): Promise => { const format = color.format as VideoSamplePixelFormat | null; const alphaFormat = alpha.format as VideoSamplePixelFormat | null; if (!format || !alphaFormat) { throw new Error('CPU color/alpha merging requires a known VideoFrame format.'); } // The alpha frame must have the same bit depth as the color frame const colorIs10 = format.includes('P10'); const colorIs12 = format.includes('P12'); const alphaIs10 = alphaFormat.includes('P10'); const alphaIs12 = alphaFormat.includes('P12'); if (alphaIs10 !== colorIs10 || alphaIs12 !== colorIs12) { throw new Error( `CPU color/alpha merging requires the alpha frame to have the same bit depth as the color frame` + ` (color: '${format}', alpha: '${alphaFormat}').`, ); } const width = color.codedWidth; const height = color.codedHeight; if (format === 'RGBX' || format === 'RGBA' || format === 'BGRX' || format === 'BGRA') { return await mergeInterleavedRgba(color, alpha, width, height, format); } else if ( format === 'I420' || format === 'I420P10' || format === 'I420P12' || format === 'I422' || format === 'I422P10' || format === 'I422P12' || format === 'I444' || format === 'I444P10' || format === 'I444P12' ) { return await mergePlanarYuv(color, alpha, width, height, format); } else if (format === 'NV12') { return await mergeNv12(color, alpha, width, height); } throw new Error(`CPU color/alpha merging does not support format '${format}'.`); }; const mergeInterleavedRgba = async ( color: VideoFrame, alpha: VideoFrame, width: number, height: number, format: 'RGBX' | 'RGBA' | 'BGRX' | 'BGRA', ): Promise => { const pixelCount = width * height; const output = new Uint8Array(pixelCount * 4); // Color goes straight into the output buffer via copyTo, no intermediate copy needed await color.copyTo(output); // And now add the alpha data const alphaY = await readAlpha(alpha, width, height, 1); for (let i = 0, j = 3; i < pixelCount; i++, j += 4) { output[j] = alphaY[i]!; } const outputFormat = (format === 'RGBX' || format === 'RGBA') ? 'RGBA' : 'BGRA'; const init = { format: outputFormat, codedWidth: width, codedHeight: height, timestamp: color.timestamp, duration: color.duration ?? undefined, transfer: [output.buffer], } as const; return new VideoFrame(output, init); }; const mergePlanarYuv = async ( color: VideoFrame, alpha: VideoFrame, width: number, height: number, format: | 'I420' | 'I420P10' | 'I420P12' | 'I422' | 'I422P10' | 'I422P12' | 'I444' | 'I444P10' | 'I444P12', ): Promise => { const is10 = format.includes('P10'); const is12 = format.includes('P12'); const bytesPerSample = (is10 || is12) ? 2 : 1; let chromaW: number; let chromaH: number; if (format.startsWith('I420')) { chromaW = Math.ceil(width / 2); chromaH = Math.ceil(height / 2); } else if (format.startsWith('I422')) { chromaW = Math.ceil(width / 2); chromaH = height; } else { chromaW = width; chromaH = height; } const ySamples = width * height; const uvSamples = chromaW * chromaH; const yBytes = ySamples * bytesPerSample; const uvBytes = uvSamples * bytesPerSample; const aBytes = ySamples * bytesPerSample; const outputBytes = yBytes + 2 * uvBytes + aBytes; const output = new Uint8Array(outputBytes); // Write color planes directly into the output buffer via copyTo, no intermediate copy await color.copyTo(output); const alphaY = await readAlpha(alpha, width, height, bytesPerSample); const aOffset = yBytes + 2 * uvBytes; output.set(alphaY, aOffset); const outputFormat = (format.slice(0, 4) + 'A' + format.slice(4)) as VideoPixelFormat; const init = { format: outputFormat, codedWidth: width, codedHeight: height, timestamp: color.timestamp, duration: color.duration ?? undefined, transfer: [output.buffer], }; return new VideoFrame(output, init); }; const mergeNv12 = async ( color: VideoFrame, alpha: VideoFrame, width: number, height: number, ): Promise => { const ySize = width * height; const chromaW = Math.ceil(width / 2); const chromaH = Math.ceil(height / 2); const uvSize = chromaW * chromaH; const sourceSize = color.allocationSize(); if (!cpuColorBuffer || cpuColorBuffer.byteLength !== sourceSize) { cpuColorBuffer = new Uint8Array(sourceSize); } await color.copyTo(cpuColorBuffer); const output = new Uint8Array(ySize + 2 * uvSize + ySize); // Y plane copies straight over output.set(cpuColorBuffer.subarray(0, ySize), 0); // Deinterleave the UV plane into separate U and V planes const uOffset = ySize; const vOffset = ySize + uvSize; const uvStart = ySize; for (let i = 0; i < uvSize; i++) { output[uOffset + i] = cpuColorBuffer[uvStart + i * 2]!; output[vOffset + i] = cpuColorBuffer[uvStart + i * 2 + 1]!; } const alphaY = await readAlpha(alpha, width, height, 1); output.set(alphaY, ySize + 2 * uvSize); const init = { format: 'I420A', codedWidth: width, codedHeight: height, timestamp: color.timestamp, duration: color.duration ?? undefined, transfer: [output.buffer], } as const; return new VideoFrame(output, init); }; const readAlpha = async (alpha: VideoFrame, width: number, height: number, bytesPerSample: number) => { const size = alpha.allocationSize(); if (!cpuAlphaBuffer || cpuAlphaBuffer.byteLength !== size) { cpuAlphaBuffer = new Uint8Array(size); } await alpha.copyTo(cpuAlphaBuffer); const format = alpha.format; if (format === 'RGBA' || format === 'BGRA' || format === 'RGBX' || format === 'BGRX') { // Pack alpha data tightly. Assume alpha is stored in RGB, so sample just from R for simplicity. const rOffset = (format === 'RGBA' || format === 'RGBX') ? 0 : 2; const pixelCount = width * height; for (let i = 0; i < pixelCount; i++) { cpuAlphaBuffer[i] = cpuAlphaBuffer[i * 4 + rOffset]!; } return cpuAlphaBuffer.subarray(0, pixelCount); } else { // For Y-plane-first formats (I*** and NV12), the leading width*height samples are the Y plane return cpuAlphaBuffer.subarray(0, width * height * bytesPerSample); } }; }; /** * Describes additional decoder preferences for video sinks. * @group Media sinks * @public */ export type VideoSinkDecoderOptions = { /** * A hint that configures the hardware acceleration method of the decoder. This is best left on `'no-preference'`, * the default. */ hardwareAcceleration?: 'no-preference' | 'prefer-hardware' | 'prefer-software'; /** * Hint that the selected decoder should be configured to minimize the number of packets that have to be decoded * before video frames are output. */ optimizeForLatency?: boolean; }; const validateVideoSinkDecoderOptions = (decoderOptions: VideoSinkDecoderOptions) => { if (!decoderOptions || typeof decoderOptions !== 'object') { throw new TypeError('decoderOptions must be an object.'); } if ( decoderOptions.hardwareAcceleration !== undefined && !['no-preference', 'prefer-hardware', 'prefer-software'].includes(decoderOptions.hardwareAcceleration) ) { throw new TypeError( 'decoderOptions.hardwareAcceleration, when provided, must be \'no-preference\', \'prefer-hardware\' or' + ' \'prefer-software\'.', ); } if (decoderOptions.optimizeForLatency !== undefined && typeof decoderOptions.optimizeForLatency !== 'boolean') { throw new TypeError('decoderOptions.optimizeForLatency, when provided, must be a boolean.'); } }; /** * A sink that retrieves decoded video samples (video frames) from a video track. * @group Media sinks * @public */ export class VideoSampleSink extends BaseMediaSampleSink { /** @internal */ _track: InputVideoTrack; /** @internal */ _decoderOptions: VideoSinkDecoderOptions; /** Creates a new {@link VideoSampleSink} for the given {@link InputVideoTrack}. */ constructor(videoTrack: InputVideoTrack, decoderOptions: VideoSinkDecoderOptions = {}) { if (!(videoTrack instanceof InputVideoTrack)) { throw new TypeError('videoTrack must be an InputVideoTrack.'); } validateVideoSinkDecoderOptions(decoderOptions); super(); this._track = videoTrack; this._decoderOptions = decoderOptions; } /** @internal */ async _createDecoder( onSample: (sample: VideoSample) => unknown, onError: (error: Error) => unknown, ) { if (!(await this._track.canDecode())) { throw new Error( 'This video track cannot be decoded by this browser. Make sure to check decodability before using' + ' a track.', ); } const codec = await this._track.getCodec(); const rotation = await this._track.getRotation(); let decoderConfig = await this._track.getDecoderConfig(); const timeResolution = await this._track.getTimeResolution(); assert(codec && decoderConfig); decoderConfig = { ...decoderConfig, hardwareAcceleration: this._decoderOptions.hardwareAcceleration, optimizeForLatency: this._decoderOptions.optimizeForLatency, }; return new VideoDecoderWrapper(onSample, onError, codec, decoderConfig, rotation, timeResolution); } /** @internal */ _createPacketSink() { return new EncodedPacketSink(this._track); } /** * Retrieves the video sample (frame) corresponding to the given timestamp, in seconds. More specifically, returns * the last video sample (in presentation order) with a start timestamp less than or equal to the given timestamp. * Returns null if the timestamp is before the track's first timestamp. * * @param timestamp - The timestamp used for retrieval, in seconds. * @param options - Options used for the underlying packet retrieval. */ async getSample(timestamp: number, options: PacketRetrievalOptions = {}) { validateTimestamp(timestamp); for await (const sample of this.mediaSamplesAtTimestamps([timestamp], options)) { return sample; } throw new Error('Internal error: Iterator returned nothing.'); } /** * Creates an async iterator that yields the video samples (frames) of this track in presentation order. This method * will intelligently pre-decode a few frames ahead to enable fast iteration. * * @param startTimestamp - The timestamp in seconds at which to start yielding samples (inclusive). * @param endTimestamp - The timestamp in seconds at which to stop yielding samples (exclusive). * @param options - Options used for the underlying packet retrieval. */ samples(startTimestamp?: number, endTimestamp?: number, options: PacketRetrievalOptions = {}) { return this.mediaSamplesInRange(startTimestamp, endTimestamp, options); } /** * Creates an async iterator that yields a video sample (frame) for each timestamp in the argument. This method * uses an optimized decoding pipeline if these timestamps are monotonically sorted, decoding each packet at most * once, and is therefore more efficient than manually getting the sample for every timestamp. The iterator may * yield null if no frame is available for a given timestamp. * * This method is good for sparse access of media data. If you want primarily sequential media access, prefer * {@link VideoSampleSink.samples} instead. * * @param timestamps - An iterable or async iterable of timestamps in seconds. * @param options - Options used for the underlying packet retrieval. */ samplesAtTimestamps(timestamps: AnyIterable, options: PacketRetrievalOptions = {}) { return this.mediaSamplesAtTimestamps(timestamps, options); } } /** * A canvas with additional timing information (timestamp & duration). * @group Media sinks * @public */ export type WrappedCanvas = { /** A canvas element or offscreen canvas. */ canvas: HTMLCanvasElement | OffscreenCanvas; /** The timestamp of the corresponding video sample, in seconds. */ timestamp: number; /** The duration of the corresponding video sample, in seconds. */ duration: number; }; /** * Options for constructing a CanvasSink. * @group Media sinks * @public */ export type CanvasSinkOptions = { /** * Whether the output canvases should have transparency instead of a black background. Defaults to `false`. Set * this to `true` when using this sink to read transparent videos. */ alpha?: boolean; /** * The width of the output canvas in pixels, defaulting to the display width of the video track. If height is not * set, it will be deduced automatically based on aspect ratio. */ width?: number; /** * The height of the output canvas in pixels, defaulting to the display height of the video track. If width is not * set, it will be deduced automatically based on aspect ratio. */ height?: number; /** * The fitting algorithm in case both width and height are set. * * - `'fill'` will stretch the image to fill the entire box, potentially altering aspect ratio. * - `'contain'` will contain the entire image within the box while preserving aspect ratio. This may lead to * letterboxing. * - `'cover'` will scale the image until the entire box is filled, while preserving aspect ratio. */ fit?: 'fill' | 'contain' | 'cover'; /** * The clockwise rotation by which to rotate the raw video frame. Defaults to the rotation set in the file metadata. * Rotation is applied before resizing. */ rotation?: Rotation; /** * Specifies the rectangular region of the input video to crop to. The crop region will automatically be clamped to * the dimensions of the input video track. Cropping is performed after rotation but before resizing. The crop * region is in the _display pixel space_ of the underlying video data. */ crop?: CropRectangle; /** * When set, specifies the number of canvases in the pool. These canvases will be reused in a ring buffer / * round-robin type fashion. This keeps the amount of allocated VRAM constant and relieves the browser from * constantly allocating/deallocating canvases. A pool size of 0 or `undefined` disables the pool and means a new * canvas is created each time. */ poolSize?: number; /** Additional preferences for the underlying video decoder. */ decoderOptions?: VideoSinkDecoderOptions; }; /** * A sink that renders video samples (frames) of the given video track to canvases. This is often more useful than * directly retrieving frames, as it comes with common preprocessing steps such as resizing or applying rotation * metadata. * * This sink will yield `HTMLCanvasElement`s when in a DOM context, and `OffscreenCanvas`es otherwise. * * @group Media sinks * @public */ export class CanvasSink { /** @internal */ _videoTrack: InputVideoTrack; /** @internal */ _alpha: boolean; /** @internal */ _width!: number; /** @internal */ _height!: number; /** @internal */ _options: CanvasSinkOptions; /** @internal */ _fit: 'fill' | 'contain' | 'cover'; /** @internal */ _rotation: Rotation = 0; /** @internal */ _crop?: { left: number; top: number; width: number; height: number }; /** @internal */ _initPromise: Promise | null = null; /** @internal */ _videoSampleSink: VideoSampleSink; /** @internal */ _canvasPool: (HTMLCanvasElement | OffscreenCanvas | null)[]; /** @internal */ _nextCanvasIndex = 0; /** Creates a new {@link CanvasSink} for the given {@link InputVideoTrack}. */ constructor(videoTrack: InputVideoTrack, options: CanvasSinkOptions = {}) { if (!(videoTrack instanceof InputVideoTrack)) { throw new TypeError('videoTrack must be an InputVideoTrack.'); } if (options && typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.alpha !== undefined && typeof options.alpha !== 'boolean') { throw new TypeError('options.alpha, when provided, must be a boolean.'); } if (options.width !== undefined && (!Number.isInteger(options.width) || options.width <= 0)) { throw new TypeError('options.width, when defined, must be a positive integer.'); } if (options.height !== undefined && (!Number.isInteger(options.height) || options.height <= 0)) { throw new TypeError('options.height, when defined, must be a positive integer.'); } if (options.fit !== undefined && !['fill', 'contain', 'cover'].includes(options.fit)) { throw new TypeError('options.fit, when provided, must be one of "fill", "contain", or "cover".'); } if ( options.width !== undefined && options.height !== undefined && options.fit === undefined ) { throw new TypeError( 'When both options.width and options.height are provided, options.fit must also be provided.', ); } if (options.rotation !== undefined && ![0, 90, 180, 270].includes(options.rotation)) { throw new TypeError('options.rotation, when provided, must be 0, 90, 180 or 270.'); } if (options.crop !== undefined) { validateCropRectangle(options.crop, 'options.'); } if ( options.poolSize !== undefined && (typeof options.poolSize !== 'number' || !Number.isInteger(options.poolSize) || options.poolSize < 0) ) { throw new TypeError('poolSize must be a non-negative integer.'); } if (options.decoderOptions !== undefined) { validateVideoSinkDecoderOptions(options.decoderOptions); } this._videoTrack = videoTrack; this._alpha = options.alpha ?? false; this._options = options; this._fit = options.fit ?? 'fill'; this._videoSampleSink = new VideoSampleSink(videoTrack, options.decoderOptions); this._canvasPool = Array.from({ length: options.poolSize ?? 0 }, () => null); } /** @internal */ _ensureInit() { return this._initPromise ??= (async () => { const options = this._options; const videoTrack = this._videoTrack; const rotation = options.rotation ?? await videoTrack.getRotation(); const squarePixelWidth = await videoTrack.getSquarePixelWidth(); const squarePixelHeight = await videoTrack.getSquarePixelHeight(); const [rotatedWidth, rotatedHeight] = rotation % 180 === 0 ? [squarePixelWidth, squarePixelHeight] : [squarePixelHeight, squarePixelWidth]; let crop = options.crop; if (crop) { crop = clampCropRectangle(crop, rotatedWidth, rotatedHeight); } let [width, height] = crop ? [crop.width, crop.height] : [rotatedWidth, rotatedHeight]; const originalAspectRatio = width / height; // If width and height aren't defined together, deduce the missing value using the aspect ratio if (options.width !== undefined && options.height === undefined) { width = options.width; height = Math.round(width / originalAspectRatio); } else if (options.width === undefined && options.height !== undefined) { height = options.height; width = Math.round(height * originalAspectRatio); } else if (options.width !== undefined && options.height !== undefined) { width = options.width; height = options.height; } this._width = width; this._height = height; this._rotation = rotation; this._crop = crop; })(); } /** @internal */ _videoSampleToWrappedCanvas(sample: VideoSample): WrappedCanvas { const width = this._width; const height = this._height; let canvas = this._canvasPool[this._nextCanvasIndex]; let canvasIsNew = false; if (!canvas) { if (typeof document !== 'undefined') { // Prefer an HTMLCanvasElement canvas = document.createElement('canvas'); canvas.width = width; canvas.height = height; } else { canvas = new OffscreenCanvas(width, height); } if (this._canvasPool.length > 0) { this._canvasPool[this._nextCanvasIndex] = canvas; } canvasIsNew = true; } if (this._canvasPool.length > 0) { this._nextCanvasIndex = (this._nextCanvasIndex + 1) % this._canvasPool.length; } const context = canvas.getContext('2d', { alpha: this._alpha || isFirefox(), // Firefox has VideoFrame glitches with opaque canvases }) as CanvasRenderingContext2D | OffscreenCanvasRenderingContext2D; assert(context); context.resetTransform(); if (!canvasIsNew) { if (!this._alpha && isFirefox()) { context.fillStyle = 'black'; context.fillRect(0, 0, width, height); } else { context.clearRect(0, 0, width, height); } } sample.drawWithFit(context, { fit: this._fit, rotation: this._rotation, crop: this._crop, }); const result = { canvas, timestamp: sample.timestamp, duration: sample.duration, }; sample.close(); return result; } /** * Retrieves a canvas with the video frame corresponding to the given timestamp, in seconds. More specifically, * returns the last video frame (in presentation order) with a start timestamp less than or equal to the given * timestamp. Returns null if the timestamp is before the track's first timestamp. * * @param timestamp - The timestamp used for retrieval, in seconds. * @param options - Options used for the underlying packet retrieval. */ async getCanvas(timestamp: number, options?: PacketRetrievalOptions) { validateTimestamp(timestamp); await this._ensureInit(); const sample = await this._videoSampleSink.getSample(timestamp, options); return sample && this._videoSampleToWrappedCanvas(sample); } /** * Creates an async iterator that yields canvases with the video frames of this track in presentation order. This * method will intelligently pre-decode a few frames ahead to enable fast iteration. * * @param startTimestamp - The timestamp in seconds at which to start yielding canvases (inclusive). * @param endTimestamp - The timestamp in seconds at which to stop yielding canvases (exclusive). * @param options - Options used for the underlying packet retrieval. */ async* canvases(startTimestamp?: number, endTimestamp?: number, options?: PacketRetrievalOptions) { await this._ensureInit(); yield* mapAsyncGenerator( this._videoSampleSink.samples(startTimestamp, endTimestamp, options), sample => this._videoSampleToWrappedCanvas(sample), ); } /** * Creates an async iterator that yields a canvas for each timestamp in the argument. This method uses an optimized * decoding pipeline if these timestamps are monotonically sorted, decoding each packet at most once, and is * therefore more efficient than manually getting the canvas for every timestamp. The iterator may yield null if * no frame is available for a given timestamp. * * This method is good for sparse access of media data. If you want primarily sequential media access, prefer * {@link CanvasSink.canvases} instead. * * @param timestamps - An iterable or async iterable of timestamps in seconds. * @param options - Options used for the underlying packet retrieval. */ async* canvasesAtTimestamps(timestamps: AnyIterable, options?: PacketRetrievalOptions) { await this._ensureInit(); yield* mapAsyncGenerator( this._videoSampleSink.samplesAtTimestamps(timestamps, options), sample => sample && this._videoSampleToWrappedCanvas(sample), ); } } class AudioDecoderWrapper extends DecoderWrapper { decoder: AudioDecoder | null = null; customDecoder: CustomAudioDecoder | null = null; customDecoderCallSerializer = new CallSerializer(); customDecoderQueueSize = 0; // Internal state to accumulate a precise current timestamp based on audio durations, not the (potentially // inaccurate) packet timestamps. currentTimestamp: number | null = null; // Chromium does not respect negative packet timestamps, so we must do the fixin' ourselves expectedFirstTimestamp: number | null = null; timestampOffset = 0; constructor( onSample: (sample: AudioSample) => unknown, onError: (error: Error) => unknown, codec: AudioCodec, decoderConfig: AudioDecoderConfig, ) { super(onSample, onError); const sampleHandler = (sample: AudioSample) => { let sampleTimestamp = sample.timestamp; if (this.expectedFirstTimestamp && this.currentTimestamp === null) { this.timestampOffset = this.expectedFirstTimestamp - sampleTimestamp; ; } sampleTimestamp += this.timestampOffset; if ( this.currentTimestamp === null || Math.abs(sampleTimestamp - this.currentTimestamp) >= sample.duration ) { // We need to sync with the sample timestamp again this.currentTimestamp = sampleTimestamp; } const preciseTimestamp = this.currentTimestamp; this.currentTimestamp += sample.duration; if (sample.numberOfFrames === 0) { // We skip zero-data (empty) AudioSamples. These are sometimes emitted, for example, by Firefox when it // decodes Vorbis (at the start). sample.close(); return; } // Round the timestamp to the sample rate const sampleRate = decoderConfig.sampleRate; sample.setTimestamp(Math.round(preciseTimestamp * sampleRate) / sampleRate); onSample(sample); }; const MatchingCustomDecoder = customAudioDecoders.find(x => x.supports(codec, decoderConfig)); if (MatchingCustomDecoder) { // @ts-expect-error "Can't create instance of abstract class 🤓" this.customDecoder = new MatchingCustomDecoder() as CustomAudioDecoder; // @ts-expect-error It's technically readonly this.customDecoder.codec = codec; // @ts-expect-error It's technically readonly this.customDecoder.config = decoderConfig; // @ts-expect-error It's technically readonly this.customDecoder.onSample = (sample) => { if (!(sample instanceof AudioSample)) { throw new TypeError('The argument passed to onSample must be an AudioSample.'); } sampleHandler(sample); }; void this.customDecoderCallSerializer.call(() => this.customDecoder!.init()); } else { const stack = new Error('Decoding error').stack; this.decoder = new AudioDecoder({ output: (data) => { try { sampleHandler(new AudioSample(data)); } catch (error) { this.onError(error as Error); } }, error: (error) => { error.stack = stack; // Provide a more useful stack trace, the default one sucks this.onError(error); }, }); this.decoder.configure(decoderConfig); } } getDecodeQueueSize() { if (this.customDecoder) { return this.customDecoderQueueSize; } else { assert(this.decoder); return this.decoder.decodeQueueSize; } } decode(packet: EncodedPacket) { if (this.customDecoder) { this.customDecoderQueueSize++; void this.customDecoderCallSerializer .call(() => this.customDecoder!.decode(packet)) .then(() => this.customDecoderQueueSize--); } else { assert(this.decoder); this.expectedFirstTimestamp ??= packet.timestamp; this.decoder.decode(packet.toEncodedAudioChunk()); } } async flush() { if (this.customDecoder) { await this.customDecoderCallSerializer.call(() => this.customDecoder!.flush()); } else { assert(this.decoder); await this.decoder.flush(); } this.currentTimestamp = null; this.expectedFirstTimestamp = null; this.timestampOffset = 0; } close() { if (this.customDecoder) { void this.customDecoderCallSerializer.call(() => this.customDecoder!.close()); } else { assert(this.decoder); this.decoder.close(); } } } // There are a lot of PCM variants not natively supported by the browser and by AudioData. Therefore we need a simple // decoder that maps any input PCM format into a PCM format supported by the browser. class PcmAudioDecoderWrapper extends DecoderWrapper { codec: PcmAudioCodec; inputSampleSize: 1 | 2 | 3 | 4 | 8; readInputValue: (view: DataView, byteOffset: number) => number; outputSampleSize: 1 | 2 | 4; outputFormat: 'u8' | 's16' | 's32' | 'f32'; writeOutputValue: (view: DataView, byteOffset: number, value: number) => void; // Internal state to accumulate a precise current timestamp based on audio durations, not the (potentially // inaccurate) packet timestamps. currentTimestamp: number | null = null; constructor( onSample: (sample: AudioSample) => unknown, onError: (error: Error) => unknown, public decoderConfig: AudioDecoderConfig, ) { super(onSample, onError); assert((PCM_AUDIO_CODECS as readonly string[]).includes(decoderConfig.codec)); this.codec = decoderConfig.codec as PcmAudioCodec; const { dataType, sampleSize, littleEndian } = parsePcmCodec(this.codec); this.inputSampleSize = sampleSize; switch (sampleSize) { case 1: { if (dataType === 'unsigned') { this.readInputValue = (view, byteOffset) => view.getUint8(byteOffset) - 2 ** 7; } else if (dataType === 'signed') { this.readInputValue = (view, byteOffset) => view.getInt8(byteOffset); } else if (dataType === 'ulaw') { this.readInputValue = (view, byteOffset) => fromUlaw(view.getUint8(byteOffset)); } else if (dataType === 'alaw') { this.readInputValue = (view, byteOffset) => fromAlaw(view.getUint8(byteOffset)); } else { assert(false); } }; break; case 2: { if (dataType === 'unsigned') { this.readInputValue = (view, byteOffset) => view.getUint16(byteOffset, littleEndian) - 2 ** 15; } else if (dataType === 'signed') { this.readInputValue = (view, byteOffset) => view.getInt16(byteOffset, littleEndian); } else { assert(false); } }; break; case 3: { if (dataType === 'unsigned') { this.readInputValue = (view, byteOffset) => getUint24(view, byteOffset, littleEndian) - 2 ** 23; } else if (dataType === 'signed') { this.readInputValue = (view, byteOffset) => getInt24(view, byteOffset, littleEndian); } else { assert(false); } }; break; case 4: { if (dataType === 'unsigned') { this.readInputValue = (view, byteOffset) => view.getUint32(byteOffset, littleEndian) - 2 ** 31; } else if (dataType === 'signed') { this.readInputValue = (view, byteOffset) => view.getInt32(byteOffset, littleEndian); } else if (dataType === 'float') { this.readInputValue = (view, byteOffset) => view.getFloat32(byteOffset, littleEndian); } else { assert(false); } }; break; case 8: { if (dataType === 'float') { this.readInputValue = (view, byteOffset) => view.getFloat64(byteOffset, littleEndian); } else { assert(false); } }; break; default: { assertNever(sampleSize); assert(false); }; } switch (sampleSize) { case 1: { if (dataType === 'ulaw' || dataType === 'alaw') { this.outputSampleSize = 2; this.outputFormat = 's16'; this.writeOutputValue = (view, byteOffset, value) => view.setInt16(byteOffset, value, true); } else { this.outputSampleSize = 1; this.outputFormat = 'u8'; this.writeOutputValue = (view, byteOffset, value) => view.setUint8(byteOffset, value + 2 ** 7); } }; break; case 2: { this.outputSampleSize = 2; this.outputFormat = 's16'; this.writeOutputValue = (view, byteOffset, value) => view.setInt16(byteOffset, value, true); }; break; case 3: { this.outputSampleSize = 4; this.outputFormat = 's32'; // From https://www.w3.org/TR/webcodecs: // AudioData containing 24-bit samples SHOULD store those samples in s32 or f32. When samples are // stored in s32, each sample MUST be left-shifted by 8 bits. this.writeOutputValue = (view, byteOffset, value) => view.setInt32(byteOffset, value << 8, true); }; break; case 4: { this.outputSampleSize = 4; if (dataType === 'float') { this.outputFormat = 'f32'; this.writeOutputValue = (view, byteOffset, value) => view.setFloat32(byteOffset, value, true); } else { this.outputFormat = 's32'; this.writeOutputValue = (view, byteOffset, value) => view.setInt32(byteOffset, value, true); } }; break; case 8: { this.outputSampleSize = 4; this.outputFormat = 'f32'; this.writeOutputValue = (view, byteOffset, value) => view.setFloat32(byteOffset, value, true); }; break; default: { assertNever(sampleSize); assert(false); }; }; } getDecodeQueueSize() { return 0; } decode(packet: EncodedPacket) { const inputView = toDataView(packet.data); const numberOfFrames = packet.byteLength / this.decoderConfig.numberOfChannels / this.inputSampleSize; const outputBufferSize = numberOfFrames * this.decoderConfig.numberOfChannels * this.outputSampleSize; const outputBuffer = new ArrayBuffer(outputBufferSize); const outputView = new DataView(outputBuffer); for (let i = 0; i < numberOfFrames * this.decoderConfig.numberOfChannels; i++) { const inputIndex = i * this.inputSampleSize; const outputIndex = i * this.outputSampleSize; const value = this.readInputValue(inputView, inputIndex); this.writeOutputValue(outputView, outputIndex, value); } const preciseDuration = numberOfFrames / this.decoderConfig.sampleRate; if (this.currentTimestamp === null || Math.abs(packet.timestamp - this.currentTimestamp) >= preciseDuration) { // We need to sync with the packet timestamp again this.currentTimestamp = packet.timestamp; } const preciseTimestamp = this.currentTimestamp; this.currentTimestamp += preciseDuration; const audioSample = new AudioSample({ format: this.outputFormat, data: outputBuffer, numberOfChannels: this.decoderConfig.numberOfChannels, sampleRate: this.decoderConfig.sampleRate, numberOfFrames, timestamp: preciseTimestamp, }); this.onSample(audioSample); } async flush() { // Do nothing } close() { // Do nothing } } /** * Sink for retrieving decoded audio samples from an audio track. * @group Media sinks * @public */ export class AudioSampleSink extends BaseMediaSampleSink { /** @internal */ _track: InputAudioTrack; /** Creates a new {@link AudioSampleSink} for the given {@link InputAudioTrack}. */ constructor(audioTrack: InputAudioTrack) { if (!(audioTrack instanceof InputAudioTrack)) { throw new TypeError('audioTrack must be an InputAudioTrack.'); } super(); this._track = audioTrack; } /** @internal */ async _createDecoder( onSample: (sample: AudioSample) => unknown, onError: (error: Error) => unknown, ) { if (!(await this._track.canDecode())) { throw new Error( 'This audio track cannot be decoded by this browser. Make sure to check decodability before using' + ' a track.', ); } const codec = await this._track.getCodec(); const decoderConfig = await this._track.getDecoderConfig(); assert(codec && decoderConfig); if ((PCM_AUDIO_CODECS as readonly string[]).includes(decoderConfig.codec)) { return new PcmAudioDecoderWrapper(onSample, onError, decoderConfig); } else { return new AudioDecoderWrapper(onSample, onError, codec, decoderConfig); } } /** @internal */ _createPacketSink() { return new EncodedPacketSink(this._track); } /** * Retrieves the audio sample corresponding to the given timestamp, in seconds. More specifically, returns * the last audio sample (in presentation order) with a start timestamp less than or equal to the given timestamp. * Returns null if the timestamp is before the track's first timestamp. * * @param timestamp - The timestamp used for retrieval, in seconds. * @param options - Options used for the underlying packet retrieval. */ async getSample(timestamp: number, options: PacketRetrievalOptions = {}) { validateTimestamp(timestamp); for await (const sample of this.mediaSamplesAtTimestamps([timestamp], options)) { return sample; } throw new Error('Internal error: Iterator returned nothing.'); } /** * Creates an async iterator that yields the audio samples of this track in presentation order. This method * will intelligently pre-decode a few samples ahead to enable fast iteration. * * @param startTimestamp - The timestamp in seconds at which to start yielding samples (inclusive). * @param endTimestamp - The timestamp in seconds at which to stop yielding samples (exclusive). * @param options - Options used for the underlying packet retrieval. */ samples(startTimestamp?: number, endTimestamp?: number, options: PacketRetrievalOptions = {}) { return this.mediaSamplesInRange(startTimestamp, endTimestamp, options); } /** * Creates an async iterator that yields an audio sample for each timestamp in the argument. This method * uses an optimized decoding pipeline if these timestamps are monotonically sorted, decoding each packet at most * once, and is therefore more efficient than manually getting the sample for every timestamp. The iterator may * yield null if no sample is available for a given timestamp. * * This method is good for sparse access of media data. If you want primarily sequential media access, prefer * {@link AudioSampleSink.samples} instead. * * @param timestamps - An iterable or async iterable of timestamps in seconds. * @param options - Options used for the underlying packet retrieval. */ samplesAtTimestamps(timestamps: AnyIterable, options: PacketRetrievalOptions = {}) { return this.mediaSamplesAtTimestamps(timestamps, options); } } /** * An AudioBuffer with additional timing information (timestamp & duration). * @group Media sinks * @public */ export type WrappedAudioBuffer = { /** An AudioBuffer. */ buffer: AudioBuffer; /** The timestamp of the corresponding audio sample, in seconds. */ timestamp: number; /** The duration of the corresponding audio sample, in seconds. */ duration: number; }; /** * A sink that retrieves decoded audio samples from an audio track and converts them to `AudioBuffer` instances. This is * often more useful than directly retrieving audio samples, as audio buffers can be directly used with the * Web Audio API. * @group Media sinks * @public */ export class AudioBufferSink { /** @internal */ _audioSampleSink: AudioSampleSink; /** Creates a new {@link AudioBufferSink} for the given {@link InputAudioTrack}. */ constructor(audioTrack: InputAudioTrack) { if (!(audioTrack instanceof InputAudioTrack)) { throw new TypeError('audioTrack must be an InputAudioTrack.'); } this._audioSampleSink = new AudioSampleSink(audioTrack); } /** @internal */ _audioSampleToWrappedArrayBuffer(sample: AudioSample): WrappedAudioBuffer { const result: WrappedAudioBuffer = { buffer: sample.toAudioBuffer(), timestamp: sample.timestamp, duration: sample.duration, }; sample.close(); return result; } /** * Retrieves the audio buffer corresponding to the given timestamp, in seconds. More specifically, returns * the last audio buffer (in presentation order) with a start timestamp less than or equal to the given timestamp. * Returns null if the timestamp is before the track's first timestamp. * * @param timestamp - The timestamp used for retrieval, in seconds. * @param options - Options used for the underlying packet retrieval. */ async getBuffer(timestamp: number, options?: PacketRetrievalOptions) { validateTimestamp(timestamp); const data = await this._audioSampleSink.getSample(timestamp, options); return data && this._audioSampleToWrappedArrayBuffer(data); } /** * Creates an async iterator that yields audio buffers of this track in presentation order. This method * will intelligently pre-decode a few buffers ahead to enable fast iteration. * * @param startTimestamp - The timestamp in seconds at which to start yielding buffers (inclusive). * @param endTimestamp - The timestamp in seconds at which to stop yielding buffers (exclusive). * @param options - Options used for the underlying packet retrieval. */ buffers(startTimestamp?: number, endTimestamp?: number, options?: PacketRetrievalOptions) { return mapAsyncGenerator( this._audioSampleSink.samples(startTimestamp, endTimestamp, options), data => this._audioSampleToWrappedArrayBuffer(data), ); } /** * Creates an async iterator that yields an audio buffer for each timestamp in the argument. This method * uses an optimized decoding pipeline if these timestamps are monotonically sorted, decoding each packet at most * once, and is therefore more efficient than manually getting the buffer for every timestamp. The iterator may * yield null if no buffer is available for a given timestamp. * * @param timestamps - An iterable or async iterable of timestamps in seconds. * @param options - Options used for the underlying packet retrieval. */ buffersAtTimestamps(timestamps: AnyIterable, options?: PacketRetrievalOptions) { return mapAsyncGenerator( this._audioSampleSink.samplesAtTimestamps(timestamps, options), data => data && this._audioSampleToWrappedArrayBuffer(data), ); } } ===== src/input-track.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AudioCodec, MediaCodec, VideoCodec } from './codec'; import { determineVideoPacketType } from './codec-data'; import { customAudioDecoders, customVideoDecoders } from './custom-coder'; import { Input } from './input'; import { Logging } from './logging'; import { EncodedPacketSink, PacketRetrievalOptions } from './media-sink'; import { assert, MaybePromise, Rational, Rotation, roundToDivisor, simplifyRational } from './misc'; import { TrackType } from './output'; import { EncodedPacket, PacketType } from './packet'; import { TrackDisposition } from './metadata'; import { DurationMetadataRequestOptions } from './demuxer'; /** * Contains aggregate statistics about the encoded packets of a track. * @group Input files & tracks * @public */ export type PacketStats = { /** The total number of packets. */ packetCount: number; /** The average number of packets per second. For video tracks, this will equal the average frame rate (FPS). */ averagePacketRate: number; /** The average number of bits per second. */ averageBitrate: number; }; export interface InputTrackBacking { getType(): TrackType; getId(): number; getNumber(): number; getCodec(): MaybePromise; getInternalCodecId(): MaybePromise; getName(): MaybePromise; getLanguageCode(): MaybePromise; getTimeResolution(): MaybePromise; isRelativeToUnixEpoch(): MaybePromise; getUnixTimeForTimestamp(timestamp: number): MaybePromise; getDisposition(): MaybePromise; getPairingMask(): bigint; getBitrate(): MaybePromise; getAverageBitrate(): MaybePromise; getDurationFromMetadata(options: DurationMetadataRequestOptions): Promise; getLiveRefreshInterval(): Promise; getHasOnlyKeyPackets?(): MaybePromise; getDecoderConfig(): Promise; getMetadataCodecParameterString?(): MaybePromise; getFirstPacket(options: PacketRetrievalOptions): Promise; getPacket(timestamp: number, options: PacketRetrievalOptions): Promise; getNextPacket(packet: EncodedPacket, options: PacketRetrievalOptions): Promise; getKeyPacket(timestamp: number, options: PacketRetrievalOptions): Promise; getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions): Promise; } /** * Represents a media track in an input file. * @group Input files & tracks * @public */ export abstract class InputTrack { /** The input file this track belongs to. */ readonly input: Input; /** @internal */ _backing: InputTrackBacking; /** @internal */ constructor(input: Input, backing: InputTrackBacking) { this.input = input; this._backing = backing; } /** The type of the track. */ abstract get type(): TrackType; /** Returns the codec of the track's packets. */ abstract getCodec(): Promise; /** * The codec of the track's packets. * @deprecated Use {@link InputTrack.getCodec} instead. */ // eslint-disable-next-line @typescript-eslint/no-deprecated abstract get codec(): MediaCodec | null; /** Returns the full codec parameter string for this track. */ abstract getCodecParameterString(): Promise; /** Checks if this track's packets can be decoded by the browser. */ abstract canDecode(): Promise; /** * For a given packet of this track, this method determines the actual type of this packet (key/delta) by looking * into its bitstream. Returns null if the type couldn't be determined. */ abstract determinePacketType(packet: EncodedPacket): Promise; /** * Returns whether the track metadata says that this track only contains key packets. The actual packets may * differ. */ abstract hasOnlyKeyPackets(): Promise; /** Returns true if and only if this track is a video track. */ isVideoTrack(): this is InputVideoTrack { return this instanceof InputVideoTrack; } /** Returns true if and only if this track is an audio track. */ isAudioTrack(): this is InputAudioTrack { return this instanceof InputAudioTrack; } /** The unique ID of this track in the input file. */ get id() { return this._backing.getId(); } /** * The 1-based index of this track among all tracks of the same type in the input file. For example, the first * video track has number 1, the second video track has number 2, and so on. The index refers to the order in * which the tracks are returned by {@link Input.getTracks}. */ get number() { return this._backing.getNumber(); } /** * Returns the identifier of the codec used internally by the container. It is not homogenized by Mediabunny * and depends entirely on the container format. * * This method can be used to determine the codec of a track in case Mediabunny doesn't know that codec. * * - For ISOBMFF files, this resolves to the name of the Sample Description Box (e.g. `'avc1'`). * - For Matroska files, this resolves to the value of the `CodecID` element. * - For WAVE files, this resolves to the value of the format tag in the `'fmt '` chunk. * - For ADTS files, this resolves to the `MPEG-4 Audio Object Type`. * - For MPEG-TS files, this resolves to the `streamType` value from the Program Map Table. * - In all other cases, this resolves to `null`. */ async getInternalCodecId() { return this._backing.getInternalCodecId(); } /** * See {@link InputTrack.getInternalCodecId}. * @deprecated Use {@link InputTrack.getInternalCodecId} instead. */ get internalCodecId() { return requireSync(this._backing.getInternalCodecId(), 'internalCodecId', 'getInternalCodecId'); } /** * Returns the ISO 639-2/T language code for this track. If the language is unknown, this resolves to `'und'` * (undetermined). */ async getLanguageCode() { return this._backing.getLanguageCode(); } /** * The ISO 639-2/T language code for this track. If the language is unknown, this field is `'und'` (undetermined). * @deprecated Use {@link InputTrack.getLanguageCode} instead. */ get languageCode() { return requireSync(this._backing.getLanguageCode(), 'languageCode', 'getLanguageCode'); } /** Returns the user-defined name for this track. */ async getName() { return this._backing.getName(); } /** * A user-defined name for this track. * @deprecated Use {@link InputTrack.getName} instead. */ get name() { return requireSync(this._backing.getName(), 'name', 'getName'); } /** * Returns a positive number x such that all timestamps and durations of all packets of this track are * integer multiples of 1/x. */ async getTimeResolution() { return this._backing.getTimeResolution(); } /** * A positive number x such that all timestamps and durations of all packets of this track are * integer multiples of 1/x. * @deprecated Use {@link InputTrack.getTimeResolution} instead. */ get timeResolution() { return requireSync(this._backing.getTimeResolution(), 'timeResolution', 'getTimeResolution'); } /** * Returns whether the timestamps of this track are relative to the Unix epoch (January 1, 1970 00:00:00 UTC). * When `true`, each timestamp maps to a definitive point in time. */ async isRelativeToUnixEpoch() { return this._backing.isRelativeToUnixEpoch(); } /** * Returns the Unix time (in seconds since January 1, 1970 00:00:00 UTC) that the given track timestamp (in seconds) * maps to, or `null` if there is no such mapping. This provides a piecewise-continuous mapping from this track's * timestamp space into wall-clock time. Such mapping exists, for example, for HLS playlists with * `#EXT-X-PROGRAM-DATE-TIME` tags present. * * This mapping can be available even when {@link InputTrack.isRelativeToUnixEpoch} is `false`, for example for HLS * streams with program date time information but with {@link HlsInputFormatOptions.offsetTimestampsByDateTime} * set to `false`. */ async getUnixTimeForTimestamp(timestamp: number): Promise { return this._backing.getUnixTimeForTimestamp(timestamp); } /** * Whether the track's timestamps can be mapped to Unix wall clock time via * {@link InputTrack.getUnixTimeForTimestamp}. */ async hasUnixTimeMapping(): Promise { return (await this._backing.getUnixTimeForTimestamp(await this.getFirstTimestamp())) !== null; } /** Returns the track's disposition, i.e. information about its intended usage. */ async getDisposition() { return this._backing.getDisposition(); } /** * The track's disposition, i.e. information about its intended usage. * @deprecated Use {@link InputTrack.getDisposition} instead. */ get disposition() { return requireSync(this._backing.getDisposition(), 'disposition', 'getDisposition'); } /** * Returns the peak bitrate of the track in bits per second, as specified in the track's metadata. This might not * match the actual media data's bitrate. */ async getBitrate() { return this._backing.getBitrate(); } /** * Returns the average bitrate of the track in bits per second, as specified in the track's metadata. This might * not match the actual media data's bitrate. */ async getAverageBitrate() { return this._backing.getAverageBitrate(); } /** * Returns the start timestamp of the first packet of this track, in seconds. While often near zero, this value * may be positive or even negative. A negative starting timestamp means the track's timing has been offset. Samples * with a negative timestamp should not be presented. */ async getFirstTimestamp() { const firstPacket = await this._backing.getFirstPacket({ metadataOnly: true }); return firstPacket?.timestamp ?? 0; } /** * Returns the end timestamp of the last packet of this track, in seconds. * * By default, when the underlying media is live, this method will only resolve once the live stream ends. If you * want to query the current end timestamp of the stream, set {@link PacketRetrievalOptions.skipLiveWait} to `true` * in the options. */ async computeDuration(options?: PacketRetrievalOptions) { const lastPacket = await this._backing.getPacket(Infinity, { metadataOnly: true, ...options }); const result = (lastPacket?.timestamp ?? 0) + (lastPacket?.duration ?? 0); return roundToDivisor(result, await this.getTimeResolution()); } /** * Gets the duration (end timestamp) in seconds of this track from metadata stored in the file. This value may be * approximate or diverge from the actual, precise duration returned by `.computeDuration()`, but compared to that * method, this method is cheaper. When the duration cannot be determined from the file metadata, `null` * is returned. * * By default, when the underlying media is live, this method will only resolve once the live stream * ends. If you want to query the current duration of the media, set * {@link DurationMetadataRequestOptions.skipLiveWait} to `true` in the options. */ async getDurationFromMetadata(options: DurationMetadataRequestOptions = {}) { return this._backing.getDurationFromMetadata(options); } /** * Computes aggregate packet statistics for this track, such as average packet rate or bitrate. * * @param targetPacketCount - This optional parameter sets a target for how many packets this method must have * looked at before it can return early; this means, you can use it to aggregate only a subset (prefix) of all * packets. This is very useful for getting a great estimate of video frame rate without having to scan through the * entire file. * * By default, when the underlying media is live and `targetPacketCount` is not set, this method will only resolve * once the live stream ends. If you want to query the current packet statistics of the stream, set * {@link PacketRetrievalOptions.skipLiveWait} to `true` in the options. */ async computePacketStats(targetPacketCount = Infinity, options?: PacketRetrievalOptions): Promise { const sink = new EncodedPacketSink(this); let startTimestamp = Infinity; let endTimestamp = -Infinity; let packetCount = 0; let totalPacketBytes = 0; for await (const packet of sink.packets(undefined, undefined, { metadataOnly: true, ...options })) { if ( packetCount >= targetPacketCount // This additional condition is needed to produce correct results with out-of-presentation-order packets && packet.timestamp >= endTimestamp ) { break; } startTimestamp = Math.min(startTimestamp, packet.timestamp); endTimestamp = Math.max(endTimestamp, packet.timestamp + packet.duration); packetCount++; totalPacketBytes += packet.byteLength; } return { packetCount, averagePacketRate: packetCount ? Number((packetCount / (endTimestamp - startTimestamp)).toPrecision(16)) : 0, averageBitrate: packetCount ? Number((8 * totalPacketBytes / (endTimestamp - startTimestamp)).toPrecision(16)) : 0, }; } /** * Whether or not this track is currently live, meaning the media's end is still unknown. * * The value returned by this method may change over time as the track stops being live. To keep track of the * track's live status, poll this method at the track's refresh interval * via {@link InputTrack.getLiveRefreshInterval}. */ async isLive() { return (await this._backing.getLiveRefreshInterval()) !== null; } /** * Returns the track's live refresh interval in seconds, or `null` if the track is not live. This interval describes * the time it takes, on average, for new live media data to become available. */ async getLiveRefreshInterval() { return this._backing.getLiveRefreshInterval(); } /** * Returns `true` if this track can be paired with the given track. Two tracks being pairable means they can be * presented (displayed) together. * * Returns `false` if `other` equals `this`. */ canBePairedWith(other: InputTrack) { if (!(other instanceof InputTrack)) { throw new TypeError('other must be an InputTrack.'); } if (this.input !== other.input || this === other) { return false; } return (this._backing.getPairingMask() & other._backing.getPairingMask()) !== 0n; } /** * Gets the list of other tracks that can be paired with this track. An optional query can be provided to narrow * down the results. */ async getPairableTracks(query?: InputTrackQuery) { return this.input.getTracks(mergeInputTrackQueries({ filter: t => t.canBePairedWith(this), }, query)); } /** * Gets the list of other video tracks that can be paired with this track. An optional query can be provided to * narrow down the results. */ async getPairableVideoTracks(query?: InputTrackQuery) { return this.input.getVideoTracks(mergeInputTrackQueries({ filter: t => t.canBePairedWith(this), }, query)); } /** * Gets the list of other audio tracks that can be paired with this track. An optional query can be provided to * narrow down the results. */ async getPairableAudioTracks(query?: InputTrackQuery) { return this.input.getAudioTracks(mergeInputTrackQueries({ filter: t => t.canBePairedWith(this), }, query)); } /** Returns the primary track that can be paired with this track, optionally steered by the provided query. */ async getPrimaryPairableVideoTrack(query?: InputTrackQuery) { return this.input.getPrimaryVideoTrack(mergeInputTrackQueries({ filter: t => t.canBePairedWith(this), }, query)); } /** Returns the primary track that can be paired with this track, optionally steered by the provided query. */ async getPrimaryPairableAudioTrack(query?: InputTrackQuery) { return this.input.getPrimaryAudioTrack(mergeInputTrackQueries({ filter: t => t.canBePairedWith(this), }, query)); } /** Returns `true` if there is another track that can be paired with this track. */ async hasPairableTrack(predicate?: (track: InputTrack) => MaybePromise): Promise { predicate &&= toValidatedPredicate(predicate); const tracks = await this.input.getTracks(); for (const track of tracks) { if (!this.canBePairedWith(track)) { continue; } if (!predicate || await predicate(track)) { return true; } } return false; } /** Returns `true` if there is a video track that can be paired with this track. */ hasPairableVideoTrack(predicate?: (track: InputVideoTrack) => MaybePromise): Promise { predicate &&= toValidatedPredicate(predicate); return this.hasPairableTrack(async x => x.isVideoTrack() && (!predicate || await predicate(x)), ); } /** Returns `true` if there is an audio track that can be paired with this track. */ hasPairableAudioTrack(predicate?: (track: InputAudioTrack) => MaybePromise): Promise { predicate &&= toValidatedPredicate(predicate); return this.hasPairableTrack(async x => x.isAudioTrack() && (!predicate || await predicate(x)), ); } } const requireSync = (value: MaybePromise, getterName: string, asyncName: string): T => { if (value instanceof Promise) { throw new Error( `'${getterName}' is deprecated and not available synchronously for this track. Use the preferred` + ` '${asyncName}()' instead.`, ); } return value; }; const toValidatedPredicate = ( predicate?: (track: T) => MaybePromise, ) => { if (predicate !== undefined && typeof predicate !== 'function') { throw new TypeError('predicate, when provided, must be a function.'); } return predicate ? (track: T) => { const handle = (result: boolean) => { if (typeof result !== 'boolean') { throw new TypeError('predicate must return or resolve to a boolean value.'); } return result; }; const result = predicate(track); if (result instanceof Promise) { return result.then(handle); } return handle(result); } : undefined; }; export interface InputVideoTrackBacking extends InputTrackBacking { getType(): 'video'; getCodec(): MaybePromise; getCodedWidth(): MaybePromise; getCodedHeight(): MaybePromise; getSquarePixelWidth(): MaybePromise; getSquarePixelHeight(): MaybePromise; getMetadataDisplayWidth?(): MaybePromise; getMetadataDisplayHeight?(): MaybePromise; getRotation(): MaybePromise; getColorSpace(): Promise; canBeTransparent(): Promise; getDecoderConfig(): Promise; } /** * Represents a video track in an input file. * @group Input files & tracks * @public */ export class InputVideoTrack extends InputTrack { /** @internal */ override _backing: InputVideoTrackBacking; /** @internal */ _pixelAspectRatioCache: Rational | null = null; /** @internal */ constructor(input: Input, backing: InputVideoTrackBacking) { super(input, backing); this._backing = backing; } get type(): TrackType { return 'video'; } /** The codec of the track's packets. */ async getCodec(): Promise { return this._backing.getCodec(); } /** * The codec of the track's packets. * @deprecated Use {@link InputVideoTrack.getCodec} instead. */ get codec(): VideoCodec | null { return requireSync(this._backing.getCodec(), 'codec', 'getCodec'); } async hasOnlyKeyPackets() { return (await this._backing.getHasOnlyKeyPackets?.()) ?? false; } /** Returns the width in pixels of the track's coded samples, before any transformations or rotations. */ async getCodedWidth() { return this._backing.getCodedWidth(); } /** * The width in pixels of the track's coded samples, before any transformations or rotations. * @deprecated Use {@link InputVideoTrack.getCodedWidth} instead. */ get codedWidth() { return requireSync(this._backing.getCodedWidth(), 'codedWidth', 'getCodedWidth'); } /** Returns the height in pixels of the track's coded samples, before any transformations or rotations. */ async getCodedHeight() { return this._backing.getCodedHeight(); } /** * The height in pixels of the track's coded samples, before any transformations or rotations. * @deprecated Use {@link InputVideoTrack.getCodedHeight} instead. */ get codedHeight() { return requireSync(this._backing.getCodedHeight(), 'codedHeight', 'getCodedHeight'); } /** Returns the angle in degrees by which the track's frames should be rotated (clockwise). */ async getRotation() { return this._backing.getRotation(); } /** * The angle in degrees by which the track's frames should be rotated (clockwise). * @deprecated Use {@link InputVideoTrack.getRotation} instead. */ get rotation() { return requireSync(this._backing.getRotation(), 'rotation', 'getRotation'); } /** * Returns the width of the track's frames in square pixels, adjusted for pixel aspect ratio but before rotation. */ async getSquarePixelWidth() { return this._backing.getSquarePixelWidth(); } /** * The width of the track's frames in square pixels, adjusted for pixel aspect ratio but before rotation. * @deprecated Use {@link InputVideoTrack.getSquarePixelWidth} instead. */ get squarePixelWidth() { return requireSync(this._backing.getSquarePixelWidth(), 'squarePixelWidth', 'getSquarePixelWidth'); } /** * Returns the height of the track's frames in square pixels, adjusted for pixel aspect ratio but before rotation. */ async getSquarePixelHeight() { return this._backing.getSquarePixelHeight(); } /** * The height of the track's frames in square pixels, adjusted for pixel aspect ratio but before rotation. * @deprecated Use {@link InputVideoTrack.getSquarePixelHeight} instead. */ get squarePixelHeight() { return requireSync(this._backing.getSquarePixelHeight(), 'squarePixelHeight', 'getSquarePixelHeight'); } /** * Returns the pixel aspect ratio of the track's frames as a rational number in its reduced form. Most videos use * square pixels (1:1). */ async getPixelAspectRatio() { // Potential minor async race condition here if called twice, but doesn't matter since the computation is // so cheap return this._pixelAspectRatioCache ??= simplifyRational({ num: (await this.getSquarePixelWidth()) * (await this.getCodedHeight()), den: (await this.getSquarePixelHeight()) * (await this.getCodedWidth()), }); } /** * The pixel aspect ratio of the track's frames, as a rational number in its reduced form. Most videos use * square pixels (1:1). * @deprecated Use {@link InputVideoTrack.getPixelAspectRatio} instead. */ get pixelAspectRatio() { return this._pixelAspectRatioCache ??= simplifyRational({ num: requireSync(this._backing.getSquarePixelWidth(), 'pixelAspectRatio', 'getPixelAspectRatio') * requireSync(this._backing.getCodedHeight(), 'pixelAspectRatio', 'getPixelAspectRatio'), den: requireSync(this._backing.getSquarePixelHeight(), 'pixelAspectRatio', 'getPixelAspectRatio') * requireSync(this._backing.getCodedWidth(), 'pixelAspectRatio', 'getPixelAspectRatio'), }); } /** Returns the display width of the track's frames in pixels, after aspect ratio adjustment and rotation. */ async getDisplayWidth() { const metadata = await this._backing.getMetadataDisplayWidth?.(); if (metadata != null) { return metadata; } const rotation = await this.getRotation(); return rotation % 180 === 0 ? this.getSquarePixelWidth() : this.getSquarePixelHeight(); } /** * The display width of the track's frames in pixels, after aspect ratio adjustment and rotation. * @deprecated Use {@link InputVideoTrack.getDisplayWidth} instead. */ get displayWidth() { const metadataRaw = this._backing.getMetadataDisplayWidth?.(); if (metadataRaw !== undefined) { const metadata = requireSync(metadataRaw, 'displayWidth', 'getDisplayWidth'); if (metadata !== null) { return metadata; } } const rotation = requireSync(this._backing.getRotation(), 'displayWidth', 'getDisplayWidth'); const value = rotation % 180 === 0 ? this._backing.getSquarePixelWidth() : this._backing.getSquarePixelHeight(); return requireSync(value, 'displayWidth', 'getDisplayWidth'); } /** Returns the display height of the track's frames in pixels, after aspect ratio adjustment and rotation. */ async getDisplayHeight() { const metadata = await this._backing.getMetadataDisplayHeight?.(); if (metadata != null) { return metadata; } const rotation = await this.getRotation(); return rotation % 180 === 0 ? this.getSquarePixelHeight() : this.getSquarePixelWidth(); } /** * The display height of the track's frames in pixels, after aspect ratio adjustment and rotation. * @deprecated Use {@link InputVideoTrack.getDisplayHeight} instead. */ get displayHeight() { const metadataRaw = this._backing.getMetadataDisplayHeight?.(); if (metadataRaw !== undefined) { const metadata = requireSync(metadataRaw, 'displayHeight', 'getDisplayHeight'); if (metadata !== null) { return metadata; } } const rotation = requireSync(this._backing.getRotation(), 'displayHeight', 'getDisplayHeight'); const value = rotation % 180 === 0 ? this._backing.getSquarePixelHeight() : this._backing.getSquarePixelWidth(); return requireSync(value, 'displayHeight', 'getDisplayHeight'); } /** Returns the color space of the track's samples. */ async getColorSpace() { return this._backing.getColorSpace(); } /** If this method returns true, the track's samples use a high dynamic range (HDR). */ async hasHighDynamicRange() { const colorSpace = await this._backing.getColorSpace(); return (colorSpace.primaries as string) === 'bt2020' || (colorSpace.primaries as string) === 'smpte432' || (colorSpace.transfer as string) === 'pq' || (colorSpace.transfer as string) === 'hlg' || (colorSpace.matrix as string) === 'bt2020-ncl'; } /** Checks if this track may contain transparent samples with alpha data. */ async canBeTransparent() { return this._backing.canBeTransparent(); } /** * Returns the [decoder configuration](https://www.w3.org/TR/webcodecs/#video-decoder-config) for decoding the * track's packets using a [`VideoDecoder`](https://developer.mozilla.org/en-US/docs/Web/API/VideoDecoder). Returns * null if the track's codec is unknown. */ async getDecoderConfig() { return this._backing.getDecoderConfig(); } async getCodecParameterString() { const fromMetadata = await this._backing.getMetadataCodecParameterString?.(); if (fromMetadata != null) { return fromMetadata; } const decoderConfig = await this._backing.getDecoderConfig(); return decoderConfig?.codec ?? null; } async canDecode() { try { const decoderConfig = await this._backing.getDecoderConfig(); if (!decoderConfig) { return false; } const codec = await this._backing.getCodec(); assert(codec !== null); if (customVideoDecoders.some(x => x.supports(codec, decoderConfig))) { return true; } if (typeof VideoDecoder === 'undefined') { return false; } const support = await VideoDecoder.isConfigSupported(decoderConfig); return support.supported === true; } catch (error) { Logging._error('Error during decodability check:', error); return false; } } async determinePacketType(packet: EncodedPacket): Promise { if (!(packet instanceof EncodedPacket)) { throw new TypeError('packet must be an EncodedPacket.'); } if (packet.isMetadataOnly) { throw new TypeError('packet must not be metadata-only to determine its type.'); } const codec = await this.getCodec(); if (codec === null) { return null; } const decoderConfig = await this.getDecoderConfig(); assert(decoderConfig); return determineVideoPacketType(codec, decoderConfig, packet.data); } } export interface InputAudioTrackBacking extends InputTrackBacking { getType(): 'audio'; getCodec(): MaybePromise; getNumberOfChannels(): MaybePromise; getSampleRate(): MaybePromise; getDecoderConfig(): Promise; } /** * Represents an audio track in an input file. * @group Input files & tracks * @public */ export class InputAudioTrack extends InputTrack { /** @internal */ override _backing: InputAudioTrackBacking; /** @internal */ constructor(input: Input, backing: InputAudioTrackBacking) { super(input, backing); this._backing = backing; } get type(): TrackType { return 'audio'; } /** The codec of the track's packets. */ async getCodec(): Promise { return this._backing.getCodec(); } /** * The codec of the track's packets. * @deprecated Use {@link InputAudioTrack.getCodec} instead. */ get codec(): AudioCodec | null { return requireSync(this._backing.getCodec(), 'codec', 'getCodec'); } async hasOnlyKeyPackets() { return (await this._backing.getHasOnlyKeyPackets?.()) ?? true; } /** Returns the number of audio channels in the track. */ async getNumberOfChannels() { return this._backing.getNumberOfChannels(); } /** * The number of audio channels in the track. * @deprecated Use {@link InputAudioTrack.getNumberOfChannels} instead. */ get numberOfChannels() { return requireSync(this._backing.getNumberOfChannels(), 'numberOfChannels', 'getNumberOfChannels'); } /** Returns the track's audio sample rate in hertz. */ async getSampleRate() { return this._backing.getSampleRate(); } /** * The track's audio sample rate in hertz. * @deprecated Use {@link InputAudioTrack.getSampleRate} instead. */ get sampleRate() { return requireSync(this._backing.getSampleRate(), 'sampleRate', 'getSampleRate'); } /** * Returns the [decoder configuration](https://www.w3.org/TR/webcodecs/#audio-decoder-config) for decoding the * track's packets using an [`AudioDecoder`](https://developer.mozilla.org/en-US/docs/Web/API/AudioDecoder). Returns * null if the track's codec is unknown. */ async getDecoderConfig() { return this._backing.getDecoderConfig(); } async getCodecParameterString() { const fromMetadata = await this._backing.getMetadataCodecParameterString?.(); if (fromMetadata != null) { return fromMetadata; } const decoderConfig = await this._backing.getDecoderConfig(); return decoderConfig?.codec ?? null; } async canDecode() { try { const decoderConfig = await this._backing.getDecoderConfig(); if (!decoderConfig) { return false; } const codec = await this._backing.getCodec(); assert(codec !== null); if (customAudioDecoders.some(x => x.supports(codec, decoderConfig))) { return true; } if (decoderConfig.codec.startsWith('pcm-')) { return true; // Since we decode it ourselves } else { if (typeof AudioDecoder === 'undefined') { return false; } const support = await AudioDecoder.isConfigSupported(decoderConfig); return support.supported === true; } } catch (error) { Logging._error('Error during decodability check:', error); return false; } } async determinePacketType(packet: EncodedPacket): Promise { if (!(packet instanceof EncodedPacket)) { throw new TypeError('packet must be an EncodedPacket.'); } if ((await this.getCodec()) === null) { return null; } return 'key'; // No audio codec with delta packets } } /** * Defines a query for input tracks. Can be used to query tracks tersely and expressively, which is especially useful * for media inputs with many tracks, such as HLS manifests. * * @group Input files & tracks * @public */ export type InputTrackQuery = { /** * A filter predicate function called for every track. Returning or resolving to `false` excludes the track from * the result. */ filter?: (track: T) => MaybePromise; /** * A function called for every track, used to define a track ordering. Tracks are ordered in ascending order using * the value returned by this function. When the function returns an array of numbers `arr`, tracks will be sorted * by `arr[0]` unless they have the same value, in which case they will be sorted by `arr[1]`, and so on. This * allows you to construct a list of ordering criteria, sorted by importance. * * To help construct complex ordering criteria, the {@link asc}, {@link desc}, and {@link prefer} helper functions * can be used. */ sortBy?: (track: T) => MaybePromise; }; /** * Helper function for use in {@link InputTrackQuery.sortBy}, used to describe sorting tracks by a numeric property in * ascending order. `null` and `undefined` are accepted too and are last in the order (sorted to the end). * * @group Input files & tracks * @public */ export const asc = (value: number | null | undefined) => { return value ?? Infinity; // nulls and undefined last }; /** * Helper function for use in {@link InputTrackQuery.sortBy}, used to describe sorting tracks by a numeric property in * descending order. `null` and `undefined` are accepted too and are last in the order (sorted to the end). * * @group Input files & tracks * @public */ export const desc = (value: number | null | undefined) => { return -(value ?? -Infinity); // nulls and undefined last }; /** * Helper function for use in {@link InputTrackQuery.sortBy}, used to sort tracks by boolean properties. `true` is * sorted to the start, `false` to the end. Useful for expressing soft preferences (e.g., "I'd prefer 1080p, but other * resolutions are fine too") as opposed to {@link InputTrackQuery.filter} which expresses hard requirements for * tracks. * * @group Input files & tracks * @public */ export const prefer = (value: boolean) => { return -value; }; export const toValidatedInputTrackQuery = ( query: InputTrackQuery, ): InputTrackQuery => { if (typeof query !== 'object' || !query) { throw new TypeError('query must be an object.'); } if (query.filter !== undefined && typeof query.filter !== 'function') { throw new TypeError('query.filter, when provided, must be a function.'); } if (query.sortBy !== undefined && typeof query.sortBy !== 'function') { throw new TypeError('query.sortBy, when provided, must be a function.'); } // Instead of validating the return types of the functions everywhere the query is used, simply return a new query // which wraps the old one while validating it. return { filter: query.filter ? (track) => { const handle = (bool: boolean) => { if (typeof bool !== 'boolean') { throw new TypeError('query.filter must return or resolve to a boolean.'); } return bool; }; const result = query.filter!(track); if (result instanceof Promise) { return result.then(handle); } else { return handle(result); } } : undefined, sortBy: query.sortBy ? (track) => { const handle = (value: number | number[]) => { if ( typeof value !== 'number' && (!Array.isArray(value) || !value.every(x => typeof x === 'number')) ) { throw new TypeError( 'query.sortBy must return or resolve to a number or an array of numbers.', ); } return value; }; const result = query.sortBy!(track); if (result instanceof Promise) { return result.then(handle); } else { return handle(result); } } : undefined, }; }; export const mergeInputTrackQueries = ( queryA: InputTrackQuery | undefined, queryB: InputTrackQuery | undefined, ): InputTrackQuery => { return { filter: queryA?.filter || queryB?.filter ? (track) => { const resultA = queryA?.filter?.(track) ?? true; const handleResultA = (resultA: boolean) => { if (resultA === false) { return false; } return queryB?.filter?.(track) ?? true; }; if (resultA instanceof Promise) { return resultA.then(handleResultA); } else { return handleResultA(resultA); } } : undefined, sortBy: queryA?.sortBy || queryB?.sortBy ? (track) => { const resultA = queryA?.sortBy?.(track) ?? []; const resultB = queryB?.sortBy?.(track) ?? []; type Result = Awaited; const join = (resultA: Result, resultB: Result) => { return [ ...(Array.isArray(resultA) ? resultA : [resultA]), ...(Array.isArray(resultB) ? resultB : [resultB]), ]; }; if (resultA instanceof Promise || resultB instanceof Promise) { return Promise.all([resultA, resultB]).then(([resultA, resultB]) => { return join(resultA, resultB); }); } else { return join(resultA, resultB); } } : undefined, }; }; export const queryInputTracks = async ( tracks: T[], query?: InputTrackQuery, ): Promise => { let matched = tracks; if (query?.filter) { const filterMatches = tracks.map(t => query.filter!(t)); const hasAsyncFilter = filterMatches.some(x => x instanceof Promise); if (hasAsyncFilter) { // eslint-disable-next-line @typescript-eslint/await-thenable const resolvedFilterMatches = await Promise.all(filterMatches); matched = tracks.filter((_, i) => resolvedFilterMatches[i]); } else { matched = tracks.filter((_, i) => filterMatches[i] as boolean); } } if (!query?.sortBy) { return matched; } const sortValues = matched.map(t => query.sortBy!(t)); const hasAsyncSort = sortValues.some(x => x instanceof Promise); const resolvedSortValues = hasAsyncSort // eslint-disable-next-line @typescript-eslint/await-thenable ? await Promise.all(sortValues) : sortValues as (number | number[])[]; return matched .map((track, i) => ({ track, sortValue: resolvedSortValues[i] })) .sort((a, b) => { const aValues = Array.isArray(a.sortValue) ? a.sortValue : [a.sortValue]; const bValues = Array.isArray(b.sortValue) ? b.sortValue : [b.sortValue]; const maxLength = Math.max(aValues.length, bValues.length); for (let i = 0; i < maxLength; i++) { const aValue = aValues[i] ?? 0; const bValue = bValues[i] ?? 0; if (aValue === bValue) { continue; } return aValue - bValue; } return 0; }) .map(x => x.track); }; ===== src/ogg/ogg-muxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { OPUS_SAMPLE_RATE, validateAudioChunkMetadata } from '../codec'; import { createVorbisComments, parseModesFromVorbisSetupPacket, parseOpusIdentificationHeader } from '../codec-data'; import { assert, promiseWithResolvers, setInt64, toDataView, toUint8Array, } from '../misc'; import { Muxer } from '../muxer'; import { Output, OutputAudioTrack, OutputTrack } from '../output'; import { OggOutputFormat } from '../output-format'; import { EncodedPacket } from '../packet'; import { Writer } from '../writer'; import { buildOggMimeType, computeOggPageCrc, extractSampleMetadata, OggCodecInfo, OGGS, } from './ogg-misc'; import { MAX_PAGE_SIZE } from './ogg-reader'; const PAGE_SIZE_TARGET = 8192; type OggTrackData = { track: OutputAudioTrack; serialNumber: number; internalSampleRate: number; codecInfo: OggCodecInfo; vorbisLastBlocksize: number | null; packetQueue: Packet[]; currentTimestampInSamples: number; pagesWritten: number; currentGranulePosition: number; currentLacingValues: number[]; currentPageData: Uint8Array[]; currentPageSize: number; currentPageStartsWithFreshPacket: boolean; currentPageStartTimestampInSamples: number; closed: boolean; }; type Packet = { data: Uint8Array; timestampInSamples: number; durationInSamples: number; forcePageFlush: boolean; }; export class OggMuxer extends Muxer { private format: OggOutputFormat; private writer!: Writer; private trackDatas: OggTrackData[] = []; private bosPagesWritten = false; private allTracksKnown = promiseWithResolvers(); private pageBytes = new Uint8Array(MAX_PAGE_SIZE); private pageView = new DataView(this.pageBytes.buffer); constructor(output: Output, format: OggOutputFormat) { super(output); this.format = format; } async start() { const release = await this.mutex.acquire(); this.writer = await this.output._getRootWriter(true); // Ogg is always monotonically written! release(); } async getMimeType() { await this.allTracksKnown.promise; return buildOggMimeType({ codecStrings: this.trackDatas.map(x => x.codecInfo.codec!), }); } addEncodedVideoPacket(): never { throw new Error('Video tracks are not supported.'); } private getTrackData(track: OutputAudioTrack, meta?: EncodedAudioChunkMetadata) { const existingTrackData = this.trackDatas.find(td => td.track === track); if (existingTrackData) { return existingTrackData; } // Give the track a unique random serial number let serialNumber: number; do { serialNumber = Math.floor(2 ** 32 * Math.random()); } while (this.trackDatas.some(td => td.serialNumber === serialNumber)); assert(track.source._codec === 'vorbis' || track.source._codec === 'opus'); validateAudioChunkMetadata(meta); assert(meta); assert(meta.decoderConfig); const newTrackData: OggTrackData = { track, serialNumber, internalSampleRate: track.source._codec === 'opus' ? OPUS_SAMPLE_RATE : meta.decoderConfig.sampleRate, codecInfo: { codec: track.source._codec, vorbisInfo: null, opusInfo: null, }, vorbisLastBlocksize: null, packetQueue: [], currentTimestampInSamples: 0, pagesWritten: 0, currentGranulePosition: 0, currentLacingValues: [], currentPageData: [], currentPageSize: 27, currentPageStartsWithFreshPacket: true, currentPageStartTimestampInSamples: 0, closed: false, }; this.queueHeaderPackets(newTrackData, meta); this.trackDatas.push(newTrackData); if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } return newTrackData; } private queueHeaderPackets(trackData: OggTrackData, meta: EncodedAudioChunkMetadata) { assert(meta.decoderConfig); if (trackData.track.source._codec === 'vorbis') { assert(meta.decoderConfig.description); const bytes = toUint8Array(meta.decoderConfig.description); if (bytes[0] !== 2) { throw new TypeError('First byte of Vorbis decoder description must be 2.'); } let pos = 1; const readPacketLength = () => { let length = 0; while (true) { const value = bytes[pos++]; if (value === undefined) { throw new TypeError('Vorbis decoder description is too short.'); } length += value; if (value < 255) { return length; } } }; const identificationHeaderLength = readPacketLength(); const commentHeaderLength = readPacketLength(); const setupHeaderLength = bytes.length - pos; // Setup header fills the remaining bytes if (setupHeaderLength <= 0) { throw new TypeError('Vorbis decoder description is too short.'); } const identificationHeader = bytes.subarray(pos, pos += identificationHeaderLength); pos += commentHeaderLength; // Skip the comment header, we'll build our own const setupHeader = bytes.subarray(pos); const commentHeaderHeader = new Uint8Array(7); commentHeaderHeader[0] = 3; // Packet type commentHeaderHeader[1] = 0x76; // 'v' commentHeaderHeader[2] = 0x6f; // 'o' commentHeaderHeader[3] = 0x72; // 'r' commentHeaderHeader[4] = 0x62; // 'b' commentHeaderHeader[5] = 0x69; // 'i' commentHeaderHeader[6] = 0x73; // 's' const commentHeader = createVorbisComments(commentHeaderHeader, this.output._metadataTags, true); trackData.packetQueue.push({ data: identificationHeader, timestampInSamples: 0, durationInSamples: 0, forcePageFlush: true, }, { data: commentHeader, timestampInSamples: 0, durationInSamples: 0, forcePageFlush: false, }, { data: setupHeader, timestampInSamples: 0, durationInSamples: 0, forcePageFlush: true, // The last header packet must flush the page }); const view = toDataView(identificationHeader); const blockSizeByte = view.getUint8(28); trackData.codecInfo.vorbisInfo = { blocksizes: [ 1 << (blockSizeByte & 0xf), 1 << (blockSizeByte >> 4), ], modeBlockflags: parseModesFromVorbisSetupPacket(setupHeader).modeBlockflags, }; } else if (trackData.track.source._codec === 'opus') { if (!meta.decoderConfig.description) { throw new TypeError('For Ogg, Opus decoder description is required.'); } const identificationHeader = toUint8Array(meta.decoderConfig.description); const commentHeaderHeader = new Uint8Array(8); const commentHeaderHeaderView = toDataView(commentHeaderHeader); commentHeaderHeaderView.setUint32(0, 0x4f707573, false); // 'Opus' commentHeaderHeaderView.setUint32(4, 0x54616773, false); // 'Tags' const commentHeader = createVorbisComments(commentHeaderHeader, this.output._metadataTags, true); trackData.packetQueue.push({ data: identificationHeader, timestampInSamples: 0, durationInSamples: 0, forcePageFlush: true, }, { data: commentHeader, timestampInSamples: 0, durationInSamples: 0, forcePageFlush: true, // The last header packet must flush the page }); trackData.codecInfo.opusInfo = { preSkip: parseOpusIdentificationHeader(identificationHeader).preSkip, }; } } async addEncodedAudioPacket(track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata) { const release = await this.mutex.acquire(); try { const trackData = this.getTrackData(track, meta); this.validateTimestamp(trackData.track, packet.timestamp, packet.type === 'key'); const currentTimestampInSamples = trackData.currentTimestampInSamples; const { durationInSamples, vorbisBlockSize } = extractSampleMetadata( packet.data, trackData.codecInfo, trackData.vorbisLastBlocksize, ); trackData.currentTimestampInSamples += durationInSamples; trackData.vorbisLastBlocksize = vorbisBlockSize; trackData.packetQueue.push({ data: packet.data, timestampInSamples: currentTimestampInSamples, durationInSamples, forcePageFlush: false, }); await this.interleavePages(); } finally { release(); } } addSubtitleCue(): never { throw new Error('Subtitle tracks are not supported.'); } allTracksAreKnown() { for (const track of this.output._tracks) { if (!track.source._closed && !this.trackDatas.some(x => x.track === track)) { return false; // We haven't seen a sample from this open track yet } } return true; } async interleavePages(isFinalCall = false) { if (!this.bosPagesWritten) { if (!this.allTracksAreKnown() && !isFinalCall) { return; // We can't interleave yet as we don't yet know how many tracks we'll truly have } // Write the header page for all bitstreams for (const trackData of this.trackDatas) { while (trackData.packetQueue.length > 0) { const packet = trackData.packetQueue.shift()!; this.writePacket(trackData, packet, false); if (packet.forcePageFlush) { // We say the header page ends once the first packet is encountered that forces a page flush break; } } } this.bosPagesWritten = true; } outer: while (true) { let trackWithMinTimestamp: OggTrackData | null = null; let minTimestamp = Infinity; for (const trackData of this.trackDatas) { if ( !isFinalCall && trackData.packetQueue.length <= 1 // Limit is 1, not 0, for correct EOS flag logic && !trackData.closed ) { break outer; } if ( trackData.packetQueue.length > 0 && trackData.packetQueue[0]!.timestampInSamples < minTimestamp ) { trackWithMinTimestamp = trackData; minTimestamp = trackData.packetQueue[0]!.timestampInSamples; } } if (!trackWithMinTimestamp) { break; } const packet = trackWithMinTimestamp.packetQueue.shift()!; const isFinalPacket = trackWithMinTimestamp.packetQueue.length === 0; this.writePacket(trackWithMinTimestamp, packet, isFinalPacket); } if (!isFinalCall) { await this.writer.flush(); } } writePacket(trackData: OggTrackData, packet: Packet, isFinalPacket: boolean) { const packetEndTimestampInSamples = packet.timestampInSamples + packet.durationInSamples; if (this.format._options.maximumPageDuration !== undefined) { const maxDurationInSamples = this.format._options.maximumPageDuration * trackData.internalSampleRate; if ( trackData.currentLacingValues.length > 0 && packetEndTimestampInSamples - trackData.currentPageStartTimestampInSamples > maxDurationInSamples ) { // Flush the current page early to avoid exceeding the maximum page duration this.writePage(trackData, false); } } let remainingLength = packet.data.length; let dataStartOffset = 0; let dataOffset = 0; while (true) { if (trackData.currentLacingValues.length === 0 && dataStartOffset > 0) { // This is a packet spanning multiple pages trackData.currentPageStartsWithFreshPacket = false; } const segmentSize = Math.min(255, remainingLength); trackData.currentLacingValues.push(segmentSize); trackData.currentPageSize++; dataOffset += segmentSize; const segmentIsLastOfPacket = remainingLength < 255; if (trackData.currentLacingValues.length === 255) { // The page is full, we need to add part of the packet data and then flush the page const slice = packet.data.subarray(dataStartOffset, dataOffset); dataStartOffset = dataOffset; trackData.currentPageData.push(slice); trackData.currentPageSize += slice.length; this.writePage(trackData, isFinalPacket && segmentIsLastOfPacket); if (segmentIsLastOfPacket) { return; } } if (segmentIsLastOfPacket) { break; } remainingLength -= 255; } const slice = packet.data.subarray(dataStartOffset); trackData.currentPageData.push(slice); trackData.currentPageSize += slice.length; trackData.currentGranulePosition = packetEndTimestampInSamples; if (trackData.currentPageSize >= PAGE_SIZE_TARGET || packet.forcePageFlush) { this.writePage(trackData, isFinalPacket); } } writePage(trackData: OggTrackData, isEos: boolean) { this.pageView.setUint32(0, OGGS, true); // Capture pattern this.pageView.setUint8(4, 0); // Version let headerType = 0; if (!trackData.currentPageStartsWithFreshPacket) { headerType |= 1; } if (trackData.pagesWritten === 0) { headerType |= 2; // Beginning of stream } if (isEos) { headerType |= 4; // End of stream } this.pageView.setUint8(5, headerType); // Header type const granulePosition = trackData.currentLacingValues.every(x => x === 255) ? -1 // No packets end on this page : trackData.currentGranulePosition; setInt64(this.pageView, 6, granulePosition, true); // Granule position this.pageView.setUint32(14, trackData.serialNumber, true); // Serial number this.pageView.setUint32(18, trackData.pagesWritten, true); // Page sequence number this.pageView.setUint32(22, 0, true); // Checksum placeholder this.pageView.setUint8(26, trackData.currentLacingValues.length); // Number of page segments this.pageBytes.set(trackData.currentLacingValues, 27); let pos = 27 + trackData.currentLacingValues.length; for (const data of trackData.currentPageData) { this.pageBytes.set(data, pos); pos += data.length; } const slice = this.pageBytes.subarray(0, pos); const crc = computeOggPageCrc(slice); this.pageView.setUint32(22, crc, true); // Checksum trackData.pagesWritten++; trackData.currentLacingValues.length = 0; trackData.currentPageData.length = 0; trackData.currentPageSize = 27; trackData.currentPageStartsWithFreshPacket = true; trackData.currentPageStartTimestampInSamples = trackData.currentGranulePosition; if (this.format._options.onPage) { this.writer.startTrackingWrites(); } this.writer.write(slice); if (this.format._options.onPage) { const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onPage(data, start, trackData.track.source); } } // eslint-disable-next-line @typescript-eslint/no-misused-promises override async onTrackClose(track: OutputTrack) { const release = await this.mutex.acquire(); const trackData = this.trackDatas.find(x => x.track === track); if (trackData) { trackData.closed = true; } if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } // Since a track is now closed, we may be able to write out chunks that were previously waiting await this.interleavePages(); release(); } async finalize() { const release = await this.mutex.acquire(); this.allTracksKnown.resolve(); for (const trackData of this.trackDatas) { trackData.closed = true; } await this.interleavePages(true); for (const trackData of this.trackDatas) { if (trackData.currentLacingValues.length > 0) { this.writePage(trackData, true); } } release(); } } ===== src/ogg/ogg-reader.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { FileSlice, readI64Le, readU32Le, readU8 } from '../reader'; import { OGGS } from './ogg-misc'; export const MIN_PAGE_HEADER_SIZE = 27; export const MAX_PAGE_HEADER_SIZE = 27 + 255; export const MAX_PAGE_SIZE = MAX_PAGE_HEADER_SIZE + 255 * 255; export type Page = { headerStartPos: number; totalSize: number; dataStartPos: number; dataSize: number; headerType: number; granulePosition: number; serialNumber: number; sequenceNumber: number; checksum: number; lacingValues: Uint8Array; }; export const readPageHeader = (slice: FileSlice): Page | null => { const startPos = slice.filePos; const capturePattern = readU32Le(slice); if (capturePattern !== OGGS) { return null; } slice.skip(1); // Version const headerType = readU8(slice); const granulePosition = readI64Le(slice); const serialNumber = readU32Le(slice); const sequenceNumber = readU32Le(slice); const checksum = readU32Le(slice); const numberPageSegments = readU8(slice); const lacingValues = new Uint8Array(numberPageSegments); for (let i = 0; i < numberPageSegments; i++) { lacingValues[i] = readU8(slice); } const headerSize = 27 + numberPageSegments; const dataSize = lacingValues.reduce((a, b) => a + b, 0); const totalSize = headerSize + dataSize; return { headerStartPos: startPos, totalSize, dataStartPos: startPos + headerSize, dataSize, headerType, granulePosition, serialNumber, sequenceNumber, checksum, lacingValues, }; }; export const findNextPageHeader = (slice: FileSlice, until: number) => { while (slice.filePos < until - (4 - 1)) { // Size of word minus 1 const word = readU32Le(slice); const firstByte = word & 0xff; const secondByte = (word >>> 8) & 0xff; const thirdByte = (word >>> 16) & 0xff; const fourthByte = (word >>> 24) & 0xff; const O = 0x4f; // 'O' if (firstByte !== O && secondByte !== O && thirdByte !== O && fourthByte !== O) { continue; } slice.skip(-4); if (word === OGGS) { // We have found the capture pattern return true; } slice.skip(1); } return false; }; ===== src/ogg/ogg-demuxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { OPUS_SAMPLE_RATE } from '../codec'; import { parseModesFromVorbisSetupPacket, parseOpusIdentificationHeader, readVorbisComments } from '../codec-data'; import { Demuxer } from '../demuxer'; import { Input } from '../input'; import { InputAudioTrackBacking } from '../input-track'; import { PacketRetrievalOptions } from '../media-sink'; import { DEFAULT_TRACK_DISPOSITION, MetadataTags, TrackDisposition } from '../metadata'; import { assert, AsyncMutex, binarySearchLessOrEqual, findLast, last, roundIfAlmostInteger, toDataView, UNDETERMINED_LANGUAGE, } from '../misc'; import { EncodedPacket, PLACEHOLDER_DATA } from '../packet'; import { readBytes, Reader } from '../reader'; import { buildOggMimeType, computeOggPageCrc, extractSampleMetadata, OggCodecInfo } from './ogg-misc'; import { findNextPageHeader, MAX_PAGE_HEADER_SIZE, MAX_PAGE_SIZE, MIN_PAGE_HEADER_SIZE, Page, readPageHeader, } from './ogg-reader'; type LogicalBitstream = { serialNumber: number; bosPage: Page; description: Uint8Array | null; numberOfChannels: number; sampleRate: number; codecInfo: OggCodecInfo; lastMetadataPacket: Packet | null; }; type Packet = { data: Uint8Array; endPage: Page; endSegmentIndex: number; }; export class OggDemuxer extends Demuxer { reader: Reader; metadataPromise: Promise | null = null; bitstreams: LogicalBitstream[] = []; trackBackings: OggAudioTrackBacking[] = []; metadataTags: MetadataTags = {}; constructor(input: Input) { super(input); this.reader = input._reader; } async readMetadata() { return this.metadataPromise ??= (async () => { let currentPos = 0; while (true) { let slice = this.reader.requestSliceRange(currentPos, MIN_PAGE_HEADER_SIZE, MAX_PAGE_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const page = readPageHeader(slice); if (!page) { break; } const isBos = !!(page.headerType & 0x02); if (!isBos) { // All bos pages for all bitstreams are required to be at the start, so if the page is not bos then // we know we've seen all bitstreams (minus chaining) break; } this.bitstreams.push({ serialNumber: page.serialNumber, bosPage: page, description: null, numberOfChannels: -1, sampleRate: -1, codecInfo: { codec: null, vorbisInfo: null, opusInfo: null, }, lastMetadataPacket: null, }); currentPos = page.headerStartPos + page.totalSize; } for (const bitstream of this.bitstreams) { const firstPacket = await this.readPacket(bitstream.bosPage, 0); if (!firstPacket) { continue; } if ( // Check for Vorbis firstPacket.data.byteLength >= 7 && firstPacket.data[0] === 0x01 // Packet type 1 = identification header && firstPacket.data[1] === 0x76 // 'v' && firstPacket.data[2] === 0x6f // 'o' && firstPacket.data[3] === 0x72 // 'r' && firstPacket.data[4] === 0x62 // 'b' && firstPacket.data[5] === 0x69 // 'i' && firstPacket.data[6] === 0x73 // 's' ) { await this.readVorbisMetadata(firstPacket, bitstream); } else if ( // Check for Opus firstPacket.data.byteLength >= 8 && firstPacket.data[0] === 0x4f // 'O' && firstPacket.data[1] === 0x70 // 'p' && firstPacket.data[2] === 0x75 // 'u' && firstPacket.data[3] === 0x73 // 's' && firstPacket.data[4] === 0x48 // 'H' && firstPacket.data[5] === 0x65 // 'e' && firstPacket.data[6] === 0x61 // 'a' && firstPacket.data[7] === 0x64 // 'd' ) { await this.readOpusMetadata(firstPacket, bitstream); } if (bitstream.codecInfo.codec !== null) { this.trackBackings.push(new OggAudioTrackBacking(bitstream, this)); } } })(); } async readVorbisMetadata(firstPacket: Packet, bitstream: LogicalBitstream) { let nextPacketPosition = await this.findNextPacketStart(firstPacket); if (!nextPacketPosition) { return; } const secondPacket = await this.readPacket(nextPacketPosition.startPage, nextPacketPosition.startSegmentIndex); if (!secondPacket) { return; } nextPacketPosition = await this.findNextPacketStart(secondPacket); if (!nextPacketPosition) { return; } const thirdPacket = await this.readPacket(nextPacketPosition.startPage, nextPacketPosition.startSegmentIndex); if (!thirdPacket) { return; } if (secondPacket.data[0] !== 0x03 || thirdPacket.data[0] !== 0x05) { return; } const lacingValues: number[] = []; const addBytesToSegmentTable = (bytes: number) => { while (true) { lacingValues.push(Math.min(255, bytes)); if (bytes < 255) { break; } bytes -= 255; } }; addBytesToSegmentTable(firstPacket.data.length); addBytesToSegmentTable(secondPacket.data.length); // We don't add the last packet to the segment table, as it is assumed to be whatever bytes remain const description = new Uint8Array( 1 + lacingValues.length + firstPacket.data.length + secondPacket.data.length + thirdPacket.data.length, ); description[0] = 2; // Num entries in the segment table description.set( lacingValues, 1, ); description.set( firstPacket.data, 1 + lacingValues.length, ); description.set( secondPacket.data, 1 + lacingValues.length + firstPacket.data.length, ); description.set( thirdPacket.data, 1 + lacingValues.length + firstPacket.data.length + secondPacket.data.length, ); bitstream.codecInfo.codec = 'vorbis'; bitstream.description = description; bitstream.lastMetadataPacket = thirdPacket; const view = toDataView(firstPacket.data); bitstream.numberOfChannels = view.getUint8(11); bitstream.sampleRate = view.getUint32(12, true); const blockSizeByte = view.getUint8(28); bitstream.codecInfo.vorbisInfo = { blocksizes: [ 1 << (blockSizeByte & 0xf), 1 << (blockSizeByte >> 4), ], modeBlockflags: parseModesFromVorbisSetupPacket(thirdPacket.data).modeBlockflags, }; readVorbisComments(secondPacket.data.subarray(7), this.metadataTags); // Skip header type and 'vorbis' } async readOpusMetadata(firstPacket: Packet, bitstream: LogicalBitstream) { // From https://datatracker.ietf.org/doc/html/rfc7845#section-5: // "An Ogg Opus logical stream contains exactly two mandatory header packets: an identification header and a // comment header." const nextPacketPosition = await this.findNextPacketStart(firstPacket); if (!nextPacketPosition) { return; } const secondPacket = await this.readPacket( nextPacketPosition.startPage, nextPacketPosition.startSegmentIndex, ); if (!secondPacket) { return; } bitstream.codecInfo.codec = 'opus'; bitstream.description = firstPacket.data; bitstream.lastMetadataPacket = secondPacket; const header = parseOpusIdentificationHeader(firstPacket.data); bitstream.numberOfChannels = header.outputChannelCount; bitstream.sampleRate = OPUS_SAMPLE_RATE; // Always the same bitstream.codecInfo.opusInfo = { preSkip: header.preSkip, }; readVorbisComments(secondPacket.data.subarray(8), this.metadataTags); // Skip 'OpusTags' } async readPacket(startPage: Page, startSegmentIndex: number): Promise { assert(startSegmentIndex < startPage.lacingValues.length); let startDataOffset = 0; for (let i = 0; i < startSegmentIndex; i++) { startDataOffset += startPage.lacingValues[i]!; } let currentPage: Page = startPage; let currentDataOffset = startDataOffset; let currentSegmentIndex = startSegmentIndex; const chunks: Uint8Array[] = []; outer: while (true) { // Load the entire page data let pageSlice = this.reader.requestSlice(currentPage.dataStartPos, currentPage.dataSize); if (pageSlice instanceof Promise) pageSlice = await pageSlice; assert(pageSlice); const pageData = readBytes(pageSlice, currentPage.dataSize); while (true) { if (currentSegmentIndex === currentPage.lacingValues.length) { chunks.push(pageData.subarray(startDataOffset, currentDataOffset)); break; } const lacingValue = currentPage.lacingValues[currentSegmentIndex]!; currentDataOffset += lacingValue; if (lacingValue < 255) { chunks.push(pageData.subarray(startDataOffset, currentDataOffset)); break outer; } currentSegmentIndex++; } // The packet extends to the next page; let's find it let currentPos = currentPage.headerStartPos + currentPage.totalSize; while (true) { let headerSlice = this.reader.requestSliceRange(currentPos, MIN_PAGE_HEADER_SIZE, MAX_PAGE_HEADER_SIZE); if (headerSlice instanceof Promise) headerSlice = await headerSlice; if (!headerSlice) { return null; } const nextPage = readPageHeader(headerSlice); if (!nextPage) { return null; } currentPage = nextPage; if (currentPage.serialNumber === startPage.serialNumber) { break; } currentPos = currentPage.headerStartPos + currentPage.totalSize; } startDataOffset = 0; currentDataOffset = 0; currentSegmentIndex = 0; } const totalPacketSize = chunks.reduce((sum, chunk) => sum + chunk.length, 0); if (totalPacketSize === 0) { return null; // Invalid packet, treat it as end of stream } const packetData = new Uint8Array(totalPacketSize); let offset = 0; for (let i = 0; i < chunks.length; i++) { const chunk = chunks[i]!; packetData.set(chunk, offset); offset += chunk.length; } return { data: packetData, endPage: currentPage, endSegmentIndex: currentSegmentIndex, }; } async findNextPacketStart(lastPacket: Packet) { // If there's another segment in the same page, return it if (lastPacket.endSegmentIndex < lastPacket.endPage.lacingValues.length - 1) { return { startPage: lastPacket.endPage, startSegmentIndex: lastPacket.endSegmentIndex + 1 }; } const isEos = !!(lastPacket.endPage.headerType & 0x04); if (isEos) { // The page is marked as the last page of the logical bitstream, so we won't find anything beyond it return null; } // Otherwise, search for the next page belonging to the same bitstream let currentPos = lastPacket.endPage.headerStartPos + lastPacket.endPage.totalSize; while (true) { let slice = this.reader.requestSliceRange(currentPos, MIN_PAGE_HEADER_SIZE, MAX_PAGE_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) { return null; } const nextPage = readPageHeader(slice); if (!nextPage) { return null; } if (nextPage.serialNumber === lastPacket.endPage.serialNumber) { return { startPage: nextPage, startSegmentIndex: 0 }; } currentPos = nextPage.headerStartPos + nextPage.totalSize; } } async getMimeType() { await this.readMetadata(); const codecStrings = await Promise.all(this.trackBackings.map( x => x.getDecoderConfig().then(c => c?.codec ?? null), )); return buildOggMimeType({ codecStrings: codecStrings.filter(Boolean) as string[], }); } async getTrackBackings() { await this.readMetadata(); return this.trackBackings; } async getMetadataTags() { await this.readMetadata(); return this.metadataTags; } } type EncodedPacketMetadata = { packet: Packet; timestampInSamples: number; durationInSamples: number; vorbisLastBlockSize: number | null; vorbisBlockSize: number | null; }; class OggAudioTrackBacking implements InputAudioTrackBacking { internalSampleRate: number; encodedPacketToMetadata = new WeakMap(); sequentialScanCache: EncodedPacketMetadata[] = []; sequentialScanMutex = new AsyncMutex(); constructor(public bitstream: LogicalBitstream, public demuxer: OggDemuxer) { // Opus always uses a fixed sample rate for its internal calculations, even if the actual rate is different this.internalSampleRate = bitstream.codecInfo.codec === 'opus' ? OPUS_SAMPLE_RATE : bitstream.sampleRate; } getType() { return 'audio' as const; } getId() { return this.bitstream.serialNumber; } getNumber() { // All Ogg tracks are audio, so the track's index + 1 is its number const index = this.demuxer.trackBackings.findIndex( x => x.bitstream === this.bitstream, ); assert(index !== -1); return index + 1; } getNumberOfChannels() { return this.bitstream.numberOfChannels; } getSampleRate() { return this.bitstream.sampleRate; } getTimeResolution() { return this.bitstream.sampleRate; } isRelativeToUnixEpoch() { return false; } getUnixTimeForTimestamp() { return null; } getPairingMask() { return 1n; } getBitrate() { return null; } getAverageBitrate() { return null; } async getDurationFromMetadata() { return null; // Not stored anywhere } async getLiveRefreshInterval() { return null; } getCodec() { return this.bitstream.codecInfo.codec; } getInternalCodecId() { return null; } async getDecoderConfig(): Promise { assert(this.bitstream.codecInfo.codec); return { codec: this.bitstream.codecInfo.codec, numberOfChannels: this.bitstream.numberOfChannels, sampleRate: this.bitstream.sampleRate, description: this.bitstream.description ?? undefined, }; } getName() { return null; } getLanguageCode() { return UNDETERMINED_LANGUAGE; } getDisposition(): TrackDisposition { return { ...DEFAULT_TRACK_DISPOSITION, primary: false, }; } granulePositionToTimestampInSamples(granulePosition: number) { if (this.bitstream.codecInfo.codec === 'opus') { assert(this.bitstream.codecInfo.opusInfo); return granulePosition - this.bitstream.codecInfo.opusInfo.preSkip; } return granulePosition; } createEncodedPacketFromOggPacket( packet: Packet | null, additional: { timestampInSamples: number; vorbisLastBlocksize: number | null; }, options: PacketRetrievalOptions, ) { if (!packet) { return null; } const { durationInSamples, vorbisBlockSize } = extractSampleMetadata( packet.data, this.bitstream.codecInfo, additional.vorbisLastBlocksize, ); const encodedPacket = new EncodedPacket( options.metadataOnly ? PLACEHOLDER_DATA : packet.data, 'key', Math.max(0, additional.timestampInSamples) / this.internalSampleRate, durationInSamples / this.internalSampleRate, packet.endPage.headerStartPos + packet.endSegmentIndex, packet.data.byteLength, ); this.encodedPacketToMetadata.set(encodedPacket, { packet, timestampInSamples: additional.timestampInSamples, durationInSamples, vorbisLastBlockSize: additional.vorbisLastBlocksize, vorbisBlockSize, }); return encodedPacket; } async getFirstPacket(options: PacketRetrievalOptions) { assert(this.bitstream.lastMetadataPacket); const packetPosition = await this.demuxer.findNextPacketStart(this.bitstream.lastMetadataPacket); if (!packetPosition) { return null; } let timestampInSamples = 0; if (this.bitstream.codecInfo.codec === 'opus') { assert(this.bitstream.codecInfo.opusInfo); timestampInSamples -= this.bitstream.codecInfo.opusInfo.preSkip; } const packet = await this.demuxer.readPacket(packetPosition.startPage, packetPosition.startSegmentIndex); return this.createEncodedPacketFromOggPacket( packet, { timestampInSamples, vorbisLastBlocksize: null, }, options, ); } async getNextPacket(prevPacket: EncodedPacket, options: PacketRetrievalOptions) { const prevMetadata = this.encodedPacketToMetadata.get(prevPacket); if (!prevMetadata) { throw new Error('Packet was not created from this track.'); } const packetPosition = await this.demuxer.findNextPacketStart(prevMetadata.packet); if (!packetPosition) { return null; } const timestampInSamples = prevMetadata.timestampInSamples + prevMetadata.durationInSamples; const packet = await this.demuxer.readPacket( packetPosition.startPage, packetPosition.startSegmentIndex, ); return this.createEncodedPacketFromOggPacket( packet, { timestampInSamples, vorbisLastBlocksize: prevMetadata.vorbisBlockSize, }, options, ); } async getPacket(timestamp: number, options: PacketRetrievalOptions) { if (this.demuxer.reader.fileSize === null) { // No file size known, can't do binary search, but fall back to sequential algo instead return this.getPacketSequential(timestamp, options); } const timestampInSamples = roundIfAlmostInteger(timestamp * this.internalSampleRate); if (timestampInSamples === 0) { // Fast path for timestamp 0 - avoids binary search when playing back from the start return this.getFirstPacket(options); } if (timestampInSamples < 0) { // There's nothing here return null; } assert(this.bitstream.lastMetadataPacket); const startPosition = await this.demuxer.findNextPacketStart(this.bitstream.lastMetadataPacket); if (!startPosition) { return null; } let lowPage = startPosition.startPage; let high = this.demuxer.reader.fileSize; const lowPages: Page[] = [lowPage]; // First, let's perform a binary serach (bisection search) on the file to find the approximate page where // we'll find the packet. We want to find a page whose end packet position is less than or equal to the // packet position we're searching for. // Outer loop: Does the binary serach outer: while (lowPage.headerStartPos + lowPage.totalSize < high) { const low = lowPage.headerStartPos; const mid = Math.floor((low + high) / 2); let searchStartPos = mid; // Inner loop: Does a linear forward scan if the page cannot be found immediately while (true) { const until = Math.min( searchStartPos + MAX_PAGE_SIZE, high - MIN_PAGE_HEADER_SIZE, ); let searchSlice = this.demuxer.reader.requestSlice(searchStartPos, until - searchStartPos); if (searchSlice instanceof Promise) searchSlice = await searchSlice; assert(searchSlice); const found = findNextPageHeader(searchSlice, until); if (!found) { high = mid + MIN_PAGE_HEADER_SIZE; continue outer; } let headerSlice = this.demuxer.reader.requestSliceRange( searchSlice.filePos, MIN_PAGE_HEADER_SIZE, MAX_PAGE_HEADER_SIZE, ); if (headerSlice instanceof Promise) headerSlice = await headerSlice; assert(headerSlice); const page = readPageHeader(headerSlice); assert(page); let pageValid = false; if (page.serialNumber === this.bitstream.serialNumber) { // Serial numbers are basically random numbers, and the chance of finding a fake page with // matching serial number is astronomically low, so we can be pretty sure this page is legit. pageValid = true; } else { let pageSlice = this.demuxer.reader.requestSlice(page.headerStartPos, page.totalSize); if (pageSlice instanceof Promise) pageSlice = await pageSlice; assert(pageSlice); // Validate the page by checking checksum const bytes = readBytes(pageSlice, page.totalSize); const crc = computeOggPageCrc(bytes); pageValid = crc === page.checksum; } if (!pageValid) { // Keep searching for a valid page searchStartPos = page.headerStartPos + 4; // 'OggS' is 4 bytes continue; } if (pageValid && page.serialNumber !== this.bitstream.serialNumber) { // Page is valid but from a different bitstream, so keep searching forward until we find one // belonging to the our bitstream searchStartPos = page.headerStartPos + page.totalSize; continue; } const isContinuationPage = page.granulePosition === -1; if (isContinuationPage) { // No packet ends on this page - keep looking searchStartPos = page.headerStartPos + page.totalSize; continue; } // The page is valid and belongs to our bitstream; let's check its granule position to see where we // need to take the bisection search. if (this.granulePositionToTimestampInSamples(page.granulePosition) > timestampInSamples) { high = page.headerStartPos; } else { lowPage = page; lowPages.push(page); } continue outer; } } // Now we have the last page with a packet position <= the packet position we're looking for, but there // might be multiple pages with the packet position, in which case we actually need to find the first of // such pages. We'll do this in two steps: First, let's find the latest page we know with an earlier packet // position, and then linear scan ourselves forward until we find the correct page. let lowerPage = startPosition.startPage; for (const otherLowPage of lowPages) { if (otherLowPage.granulePosition === lowPage.granulePosition) { break; } if (!lowerPage || otherLowPage.headerStartPos > lowerPage.headerStartPos) { lowerPage = otherLowPage; } } let currentPage = lowerPage; // Keep track of the pages we traversed, we need these later for backwards seeking const previousPages: Page[] = [currentPage]; while (true) { // This loop must terminate as we'll eventually reach lowPage if ( currentPage.serialNumber === this.bitstream.serialNumber && currentPage.granulePosition === lowPage.granulePosition ) { break; } const nextPos = currentPage.headerStartPos + currentPage.totalSize; let slice = this.demuxer.reader.requestSliceRange(nextPos, MIN_PAGE_HEADER_SIZE, MAX_PAGE_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; assert(slice); const nextPage = readPageHeader(slice); assert(nextPage); currentPage = nextPage; if (currentPage.serialNumber === this.bitstream.serialNumber) { previousPages.push(currentPage); } } assert(currentPage.granulePosition !== -1); let currentSegmentIndex: number | null = null; let currentTimestampInSamples: number; let currentTimestampIsCorrect: boolean; // These indicate the end position of the packet that the granule position belongs to let endPage = currentPage; let endSegmentIndex = 0; if (currentPage.headerStartPos === startPosition.startPage.headerStartPos) { currentTimestampInSamples = this.granulePositionToTimestampInSamples(0); currentTimestampIsCorrect = true; currentSegmentIndex = 0; } else { currentTimestampInSamples = 0; // Placeholder value! We'll refine it once we can currentTimestampIsCorrect = false; // Find the segment index of the next packet for (let i = currentPage.lacingValues.length - 1; i >= 0; i--) { const value = currentPage.lacingValues[i]!; if (value < 255) { // We know the last packet ended at i, so the next one starts at i + 1 currentSegmentIndex = i + 1; break; } } // This must hold: Since this page has a granule position set, that means there must be a packet that // ends in this page. if (currentSegmentIndex === null) { throw new Error('Invalid page with granule position: no packets end on this page.'); } endSegmentIndex = currentSegmentIndex - 1; const pseudopacket: Packet = { data: PLACEHOLDER_DATA, endPage, endSegmentIndex, }; const nextPosition = await this.demuxer.findNextPacketStart(pseudopacket); if (nextPosition) { // Let's rewind a single step (packet) - this previous packet ensures that we'll correctly compute // the duration for the packet we're looking for. const endPosition = findPreviousPacketEndPosition(previousPages, currentPage, currentSegmentIndex); assert(endPosition); const startPosition = findPacketStartPosition( previousPages, endPosition.page, endPosition.segmentIndex, ); if (startPosition) { currentPage = startPosition.page; currentSegmentIndex = startPosition.segmentIndex; } } else { // There is no next position, which means we're looking for the last packet in the bitstream. The // granule position on the last page tends to be fucky, so let's instead start the search on the // page before that. So let's loop until we find a packet that ends in a previous page. while (true) { const endPosition = findPreviousPacketEndPosition( previousPages, currentPage, currentSegmentIndex, ); if (!endPosition) { break; } const startPosition = findPacketStartPosition( previousPages, endPosition.page, endPosition.segmentIndex, ); if (!startPosition) { break; } currentPage = startPosition.page; currentSegmentIndex = startPosition.segmentIndex; if (endPosition.page.headerStartPos !== endPage.headerStartPos) { endPage = endPosition.page; endSegmentIndex = endPosition.segmentIndex; break; } } } } let lastEncodedPacket: EncodedPacket | null = null; let lastEncodedPacketMetadata: EncodedPacketMetadata | null = null; // Alright, now it's time for the final, granular seek: We keep iterating over packets until we've found the // one with the correct timestamp - i.e., the last one with a timestamp <= the timestamp we're looking for. while (currentPage !== null) { assert(currentSegmentIndex !== null); const packet = await this.demuxer.readPacket(currentPage, currentSegmentIndex); if (!packet) { break; } // We might need to skip the packet if it's a metadata one const skipPacket = currentPage.headerStartPos === startPosition.startPage.headerStartPos && currentSegmentIndex < startPosition.startSegmentIndex; if (!skipPacket) { let encodedPacket = this.createEncodedPacketFromOggPacket( packet, { timestampInSamples: currentTimestampInSamples, vorbisLastBlocksize: lastEncodedPacketMetadata?.vorbisBlockSize ?? null, }, options, ); assert(encodedPacket); let encodedPacketMetadata = this.encodedPacketToMetadata.get(encodedPacket); assert(encodedPacketMetadata); if ( !currentTimestampIsCorrect && packet.endPage.headerStartPos === endPage.headerStartPos && packet.endSegmentIndex === endSegmentIndex ) { // We know this packet end timestamp can be derived from the page's granule position currentTimestampInSamples = this.granulePositionToTimestampInSamples( currentPage.granulePosition, ); currentTimestampIsCorrect = true; // Let's backpatch the packet we just created with the correct timestamp encodedPacket = this.createEncodedPacketFromOggPacket( packet, { timestampInSamples: currentTimestampInSamples - encodedPacketMetadata.durationInSamples, vorbisLastBlocksize: lastEncodedPacketMetadata?.vorbisBlockSize ?? null, }, options, ); assert(encodedPacket); encodedPacketMetadata = this.encodedPacketToMetadata.get(encodedPacket); assert(encodedPacketMetadata); } else { currentTimestampInSamples += encodedPacketMetadata.durationInSamples; } lastEncodedPacket = encodedPacket; lastEncodedPacketMetadata = encodedPacketMetadata; if ( currentTimestampIsCorrect && ( // Next timestamp will be too late Math.max(currentTimestampInSamples, 0) > timestampInSamples // This timestamp already matches || Math.max(encodedPacketMetadata.timestampInSamples, 0) === timestampInSamples ) ) { break; } } const nextPosition = await this.demuxer.findNextPacketStart(packet); if (!nextPosition) { break; } currentPage = nextPosition.startPage; currentSegmentIndex = nextPosition.startSegmentIndex; } return lastEncodedPacket; } // A slower but simpler and sequential algorithm for finding a packet in a file async getPacketSequential(timestamp: number, options: PacketRetrievalOptions) { const release = await this.sequentialScanMutex.acquire(); // Requires exclusivity because we write to a cache try { const timestampInSamples = roundIfAlmostInteger(timestamp * this.internalSampleRate); timestamp = timestampInSamples / this.internalSampleRate; const index = binarySearchLessOrEqual( this.sequentialScanCache, timestampInSamples, x => x.timestampInSamples, ); let currentPacket: EncodedPacket | null; if (index !== -1) { // We don't need to start from the beginning, we can start at a previous scan point const cacheEntry = this.sequentialScanCache[index]!; currentPacket = this.createEncodedPacketFromOggPacket( cacheEntry.packet, { timestampInSamples: cacheEntry.timestampInSamples, vorbisLastBlocksize: cacheEntry.vorbisLastBlockSize, }, options, ); } else { currentPacket = await this.getFirstPacket(options); } let i = 0; while (currentPacket && currentPacket.timestamp < timestamp) { const nextPacket = await this.getNextPacket(currentPacket, options); if (!nextPacket || nextPacket.timestamp > timestamp) { break; } currentPacket = nextPacket; i++; if (i === 100) { // Add "checkpoints" every once in a while to speed up subsequent random accesses i = 0; const metadata = this.encodedPacketToMetadata.get(currentPacket); assert(metadata); if (this.sequentialScanCache.length > 0) { // If we reach this case, we must be at the end of the cache assert(last(this.sequentialScanCache)!.timestampInSamples <= metadata.timestampInSamples); } this.sequentialScanCache.push(metadata); } } return currentPacket; } finally { release(); } } getKeyPacket(timestamp: number, options: PacketRetrievalOptions) { return this.getPacket(timestamp, options); } getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { return this.getNextPacket(packet, options); } } /** Finds the start position of a packet given its end position. */ const findPacketStartPosition = (pageList: Page[], endPage: Page, endSegmentIndex: number) => { let page = endPage; let segmentIndex = endSegmentIndex; outer: while (true) { segmentIndex--; for (segmentIndex; segmentIndex >= 0; segmentIndex--) { const lacingValue = page.lacingValues[segmentIndex]!; if (lacingValue < 255) { segmentIndex++; // We know the last packet starts here break outer; } } assert(segmentIndex === -1); const pageStartsWithFreshPacket = !(page.headerType & 0x01); if (pageStartsWithFreshPacket) { // Fast exit: We know we don't need to look in the previous page segmentIndex = 0; break; } const previousPage = findLast( pageList, x => x.headerStartPos < page.headerStartPos, ); if (!previousPage) { return null; } page = previousPage; segmentIndex = page.lacingValues.length; } assert(segmentIndex !== -1); if (segmentIndex === page.lacingValues.length) { // Wrap back around to the first segment of the next page const nextPage = pageList[pageList.indexOf(page) + 1]; assert(nextPage); page = nextPage; segmentIndex = 0; } return { page, segmentIndex }; }; /** Finds the end position of a packet given the start position of the following packet. */ const findPreviousPacketEndPosition = (pageList: Page[], startPage: Page, startSegmentIndex: number) => { if (startSegmentIndex > 0) { // Easy return { page: startPage, segmentIndex: startSegmentIndex - 1 }; } const previousPage = findLast( pageList, x => x.headerStartPos < startPage.headerStartPos, ); if (!previousPage) { return null; } return { page: previousPage, segmentIndex: previousPage.lacingValues.length - 1 }; }; ===== src/ogg/ogg-misc.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { parseOpusTocByte } from '../codec-data'; import { assert, ilog, toDataView } from '../misc'; export const OGGS = 0x5367674f; // 'OggS' const OGG_CRC_POLYNOMIAL = 0x04c11db7; const OGG_CRC_TABLE = new Uint32Array(256); for (let n = 0; n < 256; n++) { let crc = n << 24; for (let k = 0; k < 8; k++) { crc = (crc & 0x80000000) ? ((crc << 1) ^ OGG_CRC_POLYNOMIAL) : (crc << 1); } OGG_CRC_TABLE[n] = (crc >>> 0) & 0xffffffff; } export const computeOggPageCrc = (bytes: Uint8Array) => { const view = toDataView(bytes); const originalChecksum = view.getUint32(22, true); view.setUint32(22, 0, true); // Zero out checksum field let crc = 0; for (let i = 0; i < bytes.length; i++) { const byte = bytes[i]!; crc = ((crc << 8) ^ OGG_CRC_TABLE[(crc >>> 24) ^ byte]!) >>> 0; } view.setUint32(22, originalChecksum, true); // Restore checksum field return crc; }; export type OggCodecInfo = { codec: 'vorbis' | 'opus' | null; vorbisInfo: { blocksizes: number[]; modeBlockflags: number[]; } | null; opusInfo: { preSkip: number; } | null; }; export const extractSampleMetadata = ( data: Uint8Array, codecInfo: OggCodecInfo, vorbisLastBlocksize: number | null, ) => { let durationInSamples = 0; let currentBlocksize: number | null = null; if (data.length > 0) { // To know sample duration, we'll need to peak inside the packet if (codecInfo.codec === 'vorbis') { assert(codecInfo.vorbisInfo); const vorbisModeCount = codecInfo.vorbisInfo.modeBlockflags.length; const bitCount = ilog(vorbisModeCount - 1); const modeMask = ((1 << bitCount) - 1) << 1; const modeNumber = (data[0]! & modeMask) >> 1; if (modeNumber >= codecInfo.vorbisInfo.modeBlockflags.length) { throw new Error('Invalid mode number.'); } // In Vorbis, packet duration also depends on the blocksize of the previous packet let prevBlocksize = vorbisLastBlocksize; const blockflag = codecInfo.vorbisInfo.modeBlockflags[modeNumber]!; currentBlocksize = codecInfo.vorbisInfo.blocksizes[blockflag]!; if (blockflag === 1) { const prevMask = (modeMask | 0x1) + 1; const flag = data[0]! & prevMask ? 1 : 0; prevBlocksize = codecInfo.vorbisInfo.blocksizes[flag]!; } durationInSamples = prevBlocksize !== null ? (prevBlocksize + currentBlocksize) >> 2 : 0; // The first sample outputs no audio data and therefore has a duration of 0 } else if (codecInfo.codec === 'opus') { const toc = parseOpusTocByte(data); durationInSamples = toc.durationInSamples; } } return { durationInSamples, vorbisBlockSize: currentBlocksize, }; }; export const buildOggMimeType = (info: { codecStrings: string[]; }) => { let string = 'audio/ogg'; if (info.codecStrings) { const uniqueCodecMimeTypes = [...new Set(info.codecStrings)]; string += `; codecs="${uniqueCodecMimeTypes.join(', ')}"`; } return string; }; ===== src/source.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import type { FileHandle } from 'node:fs/promises'; import { assert, binarySearchLessOrEqual, clamp, closedIntervalsOverlap, FilePath, isNumber, isWebKit, MaybePromise, mergeRequestInit, normalizeHeaders, polyfillSymbolDispose, promiseWithResolvers, retriedFetch, toDataView, toUint8Array, wait, EventEmitter, } from './misc'; import * as nodeAlias from './node'; import { InputDisposedError } from './input'; import { Logging } from './logging'; polyfillSymbolDispose(); const node = typeof nodeAlias !== 'undefined' ? nodeAlias // Aliasing it prevents some bundler warnings : undefined!; export type ReadResult = { bytes: Uint8Array; view: DataView; /** The offset of the bytes in the file. */ offset: number; }; export const DEFAULT_MIN_READ_POSITION = 0; export const DEFAULT_MAX_READ_POSITION = Infinity; /** * The events emitted by a {@link Source}, with each key being an event name and its value being the event data. * @group Input sources * @public */ export type SourceEvents = { /** Emitted each time data is retrieved from the source. */ read: { /** The start of the retrieved range, inclusive. */ start: number; /** The end of the retrieved range, exclusive. */ end: number; }; }; let sourceFinalizationRegistry: FinalizationRegistry<() => unknown> | null = null; if (typeof FinalizationRegistry !== 'undefined') { sourceFinalizationRegistry = new FinalizationRegistry((cleanup) => { cleanup(); }); } /** * The source base class, representing a resource from which bytes can be read. * @group Input sources * @public */ export abstract class Source extends EventEmitter { /** @internal */ abstract _getFileSize(): number | null | undefined; /** @internal */ abstract _read( start: number, end: number, minReadPosition: number, maxReadPosition: number, ): MaybePromise; /** @internal */ abstract _dispose(): void; /** @internal */ _disposed = false; /** @internal */ _refCount = 0; /** * Used internally to mark if a source stems from an HLS reading operation. Used to suppress certain warnings. * @internal */ _usedForHls = false; /** * FinalizationRegistry for rogue refs to this source that didn't get freed. It lives on the Source itself so that * in case the Source transitively points back to itself and forms a cycle (for example through a custom * CustomSource callback) that we're not leaking memory. * @internal */ _refFinalizationRegistry: FinalizationRegistry | null = null; /** @internal */ private _sizePromise: Promise | null = null; constructor() { super(); if (typeof FinalizationRegistry !== 'undefined') { this._refFinalizationRegistry = new FinalizationRegistry((source) => { source._decrementRefCount(); }); } } /** * Resolves with the total size of the file in bytes. This function is memoized, meaning only the first call * will retrieve the size. * * Returns null if the source is unsized. */ async getSizeOrNull() { if (this._disposed) { throw new InputDisposedError(); } return this._sizePromise ??= (async () => { let size = this._getFileSize(); if (size !== undefined) { return size; } await this._read(0, 1, DEFAULT_MIN_READ_POSITION, DEFAULT_MAX_READ_POSITION); size = this._getFileSize(); assert(size !== undefined); return size; })(); } /** * Resolves with the total size of the file in bytes. This function is memoized, meaning only the first call * will retrieve the size. * * Throws an error if the source is unsized. */ async getSize() { if (this._disposed) { throw new InputDisposedError(); } const result = await this.getSizeOrNull(); if (result === null) { throw new Error('Cannot determine the size of an unsized source.'); } return result; } /** * Returns a new {@link RangedSource} that maps data onto this source using the given offset and length. If a length * is not provided, the ranged source spans until the end of this source's data. * * Useful for reading files that are embedded within larger files. */ slice(offset: number, length?: number) { if (!Number.isInteger(offset) || offset < 0) { throw new TypeError('offset must be a non-negative integer.'); } if (length !== undefined && (!Number.isInteger(length) || length < 0)) { throw new TypeError('length, when provided, must be a non-negative integer.'); } return new RangedSource(this, offset, length); } /** * Called each time data is retrieved from the source. Will be called with the retrieved range (end exclusive). * * @deprecated Use `source.on('read', ({ start, end }) => ...)` instead. */ onread: ((start: number, end: number) => unknown) | null = null; /** @internal */ _dispatchRead(start: number, end: number) { // eslint-disable-next-line @typescript-eslint/no-deprecated this.onread?.(start, end); this._emit('read', { start, end }); } /** * Creates a new `SourceRef` pointing to this source. You are expected to call `.free()` on said `SourceRef` when * you're done with it. */ ref() { return new SourceRef(this); } /** @internal */ _incrementRefCount() { this._refCount++; } /** @internal */ _decrementRefCount() { this._refCount--; if (this._refCount === 0) { this._dispose(); this._disposed = true; } } } /** * A reference to a {@link Source}, used to manage a source's lifecycle. Creating a `SourceRef` via {@link Source.ref} * increases that source's internal reference count. As long as a source has a non-zero reference count, it is assumed * to still be in use. Once all references are freed via {@link SourceRef.free}, the source gets disposed. * * @group Input sources * @public */ export class SourceRef implements Disposable { /** @internal */ private _source: S | null; /** @internal */ private _freed = false; /** @internal */ constructor(source: S) { if (source._disposed) { throw new Error('Cannot ref a disposed source.'); } source._incrementRefCount(); source._refFinalizationRegistry?.register(this, source, this); this._source = source; } /** The {@link Source} this ref references. Accessing this field throws an error after having freed the ref. */ get source() { if (!this._source) { throw new Error('Can\'t get source; ref has already been freed.'); } return this._source; } /** Whether or not this reference has been freed via {@link SourceRef.free}. */ get freed() { return this._freed; } /** * Frees the ref, decrementing the source's internal reference count. If the source's internal reference count * reaches zero, it gets disposed. To catch bugs, this method throws if the ref is already freed. */ free() { if (this._freed) { throw new Error('Illegal operation: double free on SourceRef.'); } const source = this.source; assert(source._refCount > 0); source._decrementRefCount(); source._refFinalizationRegistry?.unregister(this); this._freed = true; this._source = null; } /** * Calls {@link SourceRef.free}. */ [Symbol.dispose]() { if (!this.freed) { this.free(); } } } /** * A source which can create new sources from file paths. Required for multi-file inputs such as HLS playlists. * @public * @group Input sources */ export abstract class PathedSource extends Source { constructor( /** * The path that points to the root file; the entry file of the media. * * This path may be modified by the source to indicate a redirect: an updated path to perform new requests * relative to. */ public rootPath: FilePath, /** The callback that is called for each requested file; must return a {@link Source} or {@link SourceRef}. */ public readonly requestHandler: (request: SourceRequest) => MaybePromise, ) { if (typeof rootPath !== 'string') { throw new TypeError('rootPath must be a string.'); } if (typeof requestHandler !== 'function') { throw new TypeError('requestHandler must be a function.'); } super(); } /** @internal */ _resolveRequest(request: SourceRequest): MaybePromise { const result = this.requestHandler(request); const handle = (result: Source | SourceRef) => { if (!(result instanceof Source || result instanceof SourceRef)) { throw new TypeError('requestHandler must return or resolve to a Source or SourceRef.'); } const ref = result instanceof Source ? result.ref() : result; ref.source._usedForHls ||= this._usedForHls; return ref; }; if (result instanceof Promise) { return result.then(handle); } else { return handle(result); } } } /** * A request for a {@link Source} at the given path. * @group Input sources * @public */ export type SourceRequest = { /** The requested file path. */ path: FilePath; /** Whether the requested file is the root file. */ isRoot: boolean; }; export const sourceRequestsAreEqual = (a: SourceRequest, b: SourceRequest) => { return a.path === b.path; }; /** * A custom multi-file source where each file is uniquely identified by a {@link FilePath} and can be resolved to * an arbitrary {@link Source}. * * @public * @group Input sources */ export class CustomPathedSource extends PathedSource { /** @internal */ _root: SourceRef | null = null; /** @internal */ _rootRequest: Promise | null = null; /** @internal */ override _read( start: number, end: number, minReadPosition: number, maxReadPosition: number, ): MaybePromise { if (!this._root) { if (!this._rootRequest) { const result = this._resolveRequest({ path: this.rootPath, isRoot: true }); const handle = (result: Source | SourceRef) => { const ref = result instanceof Source ? result.ref() : result; this._root = ref; this._rootRequest = null; return ref; }; if (result instanceof Promise) { this._rootRequest = result.then(handle); } else { handle(result); assert(this._root); } } if (this._rootRequest) { return this._rootRequest.then(ref => ref.source._read(start, end, minReadPosition, maxReadPosition)); } } return this._root!.source._read(start, end, minReadPosition, maxReadPosition); } /** @internal */ override _getFileSize(): number | null | undefined { if (this._root) { return this._root.source._getFileSize(); } return undefined; } /** @internal */ override _dispose(): void { if (this._root) { this._root.free(); } else if (this._rootRequest) { void this._rootRequest .then(ref => ref.free()); } } } /** * A source backed by an ArrayBuffer or ArrayBufferView, with the entire file held in memory. * @group Input sources * @public */ export class BufferSource extends Source { /** @internal */ _bytes: Uint8Array; /** @internal */ _view: DataView; /** @internal */ _onreadCalled = false; /** * Creates a new {@link BufferSource} backed by the specified `ArrayBuffer`, `SharedArrayBuffer`, * or `ArrayBufferView`. */ constructor(buffer: AllowSharedBufferSource) { if ( !(buffer instanceof ArrayBuffer) && !(typeof SharedArrayBuffer !== 'undefined' && buffer instanceof SharedArrayBuffer) && !ArrayBuffer.isView(buffer) ) { throw new TypeError('buffer must be an ArrayBuffer, SharedArrayBuffer, or ArrayBufferView.'); } super(); this._bytes = toUint8Array(buffer); this._view = toDataView(buffer); } /** @internal */ _getFileSize(): number { return this._bytes.byteLength; } /** @internal */ _read(): ReadResult { if (!this._onreadCalled) { // We just say the first read retrieves all bytes from the source (which, I mean, it does) this._dispatchRead(0, this._bytes.byteLength); this._onreadCalled = true; } return { bytes: this._bytes, view: this._view, offset: 0, }; } /** @internal */ _dispose() {} } /** * Options for {@link BlobSource}. * @group Input sources * @public */ export type BlobSourceOptions = { /** The maximum number of bytes the cache is allowed to hold in memory. Defaults to 8 MiB. */ maxCacheSize?: number; /** * Defaults to `true`. When `true`, Mediabunny will acquire a `ReadableStream` reader internally to efficiently read * data from the blob. Since this can lead to errors in some (very) rare cases due to browser bugs, you can set this * field to `false` to try a slower but more stable reading method. */ useStreamReader?: boolean; }; /** * A source backed by a [`Blob`](https://developer.mozilla.org/en-US/docs/Web/API/Blob). Since a * [`File`](https://developer.mozilla.org/en-US/docs/Web/API/File) is also a `Blob`, this is the source to use when * reading files off the disk. * @group Input sources * @public */ export class BlobSource extends Source { /** @internal */ _blob: Blob; /** @internal */ _options: BlobSourceOptions; /** @internal */ _orchestrator: ReadOrchestrator; /** * Creates a new {@link BlobSource} backed by the specified * [`Blob`](https://developer.mozilla.org/en-US/docs/Web/API/Blob). */ constructor(blob: Blob, options: BlobSourceOptions = {}) { if (!(blob instanceof Blob)) { throw new TypeError('blob must be a Blob.'); } if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if ( options.maxCacheSize !== undefined && (!isNumber(options.maxCacheSize) || options.maxCacheSize < 0) ) { throw new TypeError('options.maxCacheSize, when provided, must be a non-negative number.'); } if (options.useStreamReader !== undefined && typeof options.useStreamReader !== 'boolean') { throw new TypeError('options.useStreamReader, when provided, must be a boolean.'); } super(); this._blob = blob; this._options = options; this._orchestrator = new ReadOrchestrator({ maxCacheSize: options.maxCacheSize ?? (8 * 2 ** 20 /* 8 MiB */), maxWorkerCount: 4, runWorker: this._runWorker.bind(this), prefetchProfile: PREFETCH_PROFILES.fileSystem, }); this._orchestrator.fileSize = blob.size; } /** @internal */ _getFileSize(): number { return this._orchestrator.fileSize!; // Faster than blob.size } /** @internal */ _read( start: number, end: number, minReadPosition: number, maxReadPosition: number, ): MaybePromise { return this._orchestrator.read(start, end, minReadPosition, maxReadPosition); } /** @internal */ _readers = new WeakMap | null>(); /** @internal */ private async _runWorker(worker: ReadWorker) { assert(worker.strictTarget); let reader = this._readers.get(worker); if (reader === undefined) { // https://github.com/Vanilagy/mediabunny/issues/184 // WebKit has critical bugs with blob.stream(): // - WebKitBlobResource error 1 when streaming large files // - Memory buildup and reload loops on iOS (network process crashes) // - ReadableStream stalls under backpressure (especially video) // Affects Safari and all iOS browsers (Chrome, Firefox, etc.). // Use arrayBuffer() fallback for WebKit browsers. if ('stream' in this._blob && !isWebKit() && this._options.useStreamReader !== false) { // Get a reader of the blob starting at the required offset, and then keep it around const slice = this._blob.slice(worker.currentPos); reader = slice.stream().getReader(); } else { // We'll need to use more primitive ways reader = null; } this._readers.set(worker, reader); } while (worker.currentPos < worker.targetPos && !worker.aborted) { if (reader) { const { done, value } = await reader.read(); if (done) { this._orchestrator.onWorkerFinished(worker); throw new Error('Blob reader stopped unexpectedly before all requested data was read.'); } if (worker.aborted) { break; } this._dispatchRead(worker.currentPos, worker.currentPos + value.length); this._orchestrator.supplyWorkerData(worker, value); } else { const data = await this._blob.slice(worker.currentPos, worker.targetPos).arrayBuffer(); if (worker.aborted) { break; } this._dispatchRead(worker.currentPos, worker.currentPos + data.byteLength); this._orchestrator.supplyWorkerData(worker, new Uint8Array(data)); } } this._orchestrator.signalWorkerStoppedRunning(worker); if (worker.aborted) { // MDN: "Calling this method signals a loss of interest in the stream by a consumer." await reader?.cancel(); } } /** @internal */ _dispose() { this._orchestrator.dispose(); } } const URL_SOURCE_MIN_LOAD_AMOUNT = 0.5 * 2 ** 20; // 0.5 MiB const DEFAULT_RETRY_DELAY = ((previousAttempts, error, src) => { // Check if this could be a CORS error. If so, we cannot recover from it and // should not attempt to retry. // CORS errors are intentionally not opaque, so we need to rely on heuristics. const couldBeCorsError = error instanceof Error && ( error.message.includes('Failed to fetch') // Chrome || error.message.includes('Load failed') // Safari || error.message.includes('NetworkError when attempting to fetch resource') // Firefox ) && typeof window !== 'undefined'; // CORS only happens in browser environments if (couldBeCorsError) { let originOfSrc: string | null = null; // Checking if the origin is different, because only then a CORS error could originate try { if (typeof window !== 'undefined' && typeof window.location !== 'undefined') { originOfSrc = new URL(src instanceof Request ? src.url : src, window.location.href).origin; } } catch { // URL parse failed } // If user is offline, it is probably not a CORS error. const isOnline = typeof navigator !== 'undefined' && typeof navigator.onLine === 'boolean' ? navigator.onLine : true; if (isOnline && originOfSrc !== null && originOfSrc !== window.location.origin) { Logging._warn( `Request will not be retried because a CORS error was suspected due to different origins. You can` + ` modify this behavior by providing your own function for the 'getRetryDelay' option.`, ); return null; } } return Math.min(2 ** (previousAttempts - 2), 16); }) satisfies UrlSourceOptions['getRetryDelay']; const warnedOrigins = new Set(); /** * Options for {@link UrlSource}. * @group Input sources * @public */ export type UrlSourceOptions = { /** * The [`RequestInit`](https://developer.mozilla.org/en-US/docs/Web/API/RequestInit) used by the Fetch API. Can be * used to further control the requests, such as setting custom headers. * * The `signal` field is not available, as Mediabunny controls request cancellation internally. If you want to * cancel ongoing requests, use {@link Input.dispose}. */ requestInit?: Omit; /** * A function that returns the delay (in seconds) before retrying a failed request. The function is called * with the number of previous, unsuccessful attempts, as well as with the error with which the previous request * failed. If the function returns `null`, no more retries will be made. * * By default, it uses an exponential backoff algorithm that never gives up unless * a CORS error is suspected (`fetch()` did reject, `navigator.onLine` is true and origin is different). */ getRetryDelay?: (previousAttempts: number, error: unknown, url: string | URL | Request) => number | null; /** The maximum number of bytes the cache is allowed to hold in memory. Defaults to 64 MiB. */ maxCacheSize?: number; /** The maximum number of parallel requests to use for fetching. Defaults to 2. */ parallelism?: number; /** * A WHATWG-compatible fetch function. You can use this field to polyfill the `fetch` function, add missing * features, or use a custom implementation. */ fetchFn?: typeof fetch; }; /** * A source backed by a URL. This is useful for reading data from the network. Requests will be made using an optimized * reading and prefetching pattern to minimize request count and latency. * @group Input sources * @public */ export class UrlSource extends PathedSource { /** @internal */ _url: string | URL | Request; /** @internal */ _getRetryDelay: (previousAttempts: number, error: unknown, url: string | URL | Request) => number | null; /** @internal */ _options: UrlSourceOptions; /** @internal */ _requestInit: RequestInit; /** @internal */ _offset = 0; /** @internal */ _length: number | null = null; /** @internal */ _orchestrator: ReadOrchestrator; /** * Note that this value being true does NOT mean the file size can't change anymore; it just signals that we have at * least checked if we know the file size or not. * @internal */ _fileSizeDetermined = false; /** * Creates a new {@link UrlSource} backed by the resource at the specified URL. * * When passing a `Request` instance, note that its `signal` will be overridden by Mediabunny; if you want to cancel * ongoing requests, use {@link Input.dispose}. */ constructor( url: string | URL | Request, options: UrlSourceOptions = {}, ) { if ( typeof url !== 'string' && !(url instanceof URL) && !(typeof Request !== 'undefined' && url instanceof Request) ) { throw new TypeError('url must be a string, URL or Request.'); } if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.requestInit !== undefined && (!options.requestInit || typeof options.requestInit !== 'object')) { throw new TypeError('options.requestInit, when provided, must be an object.'); } if (options.getRetryDelay !== undefined && typeof options.getRetryDelay !== 'function') { throw new TypeError('options.getRetryDelay, when provided, must be a function.'); } if ( options.maxCacheSize !== undefined && (!isNumber(options.maxCacheSize) || options.maxCacheSize < 0) ) { throw new TypeError('options.maxCacheSize, when provided, must be a non-negative number.'); } if (options.parallelism !== undefined && (!Number.isInteger(options.parallelism) || options.parallelism < 1)) { throw new TypeError('options.parallelism, when provided, must be a positive number.'); } if (options.fetchFn !== undefined && typeof options.fetchFn !== 'function') { throw new TypeError('options.fetchFn, when provided, must be a function.'); // Won't bother validating this function beyond this } const urlString = url instanceof Request ? url.url : url instanceof URL ? url.href : url; super( urlString, request => new UrlSource(request.path, this._options), ); this._url = url; this._options = options; this._getRetryDelay = options.getRetryDelay ?? DEFAULT_RETRY_DELAY; // A user-supplied Range header is interpreted as a byte offset (and optional length) into the resource. We // pull it out of the request and remember it for subsequent requests. this._requestInit = { ...options.requestInit }; let rangeHeaderValue: string | null = null; if (options.requestInit?.headers) { const headers = { ...normalizeHeaders(options.requestInit.headers) }; const rangeKey = Object.keys(headers).find(key => key.toLowerCase() === 'range'); if (rangeKey !== undefined) { rangeHeaderValue = headers[rangeKey]!; delete headers[rangeKey]; this._requestInit.headers = headers; } } if (url instanceof Request) { const requestRange = url.headers.get('Range'); if (requestRange !== null) { rangeHeaderValue ??= requestRange; // Clone the request so we don't mutate the user's object, then strip the Range header const strippedRequest = new Request(url); strippedRequest.headers.delete('Range'); this._url = strippedRequest; } } if (rangeHeaderValue !== null) { const parsed = parseByteRangeHeader(rangeHeaderValue); if (parsed) { this._offset = parsed.offset; this._length = parsed.length; } } // Most files in the real-world have a single sequential access pattern, but having two in parallel can // also happen const DEFAULT_PARALLELISM = 2; this._orchestrator = new ReadOrchestrator({ maxCacheSize: options.maxCacheSize ?? (64 * 2 ** 20 /* 64 MiB */), maxWorkerCount: options.parallelism ?? DEFAULT_PARALLELISM, runWorker: this._runWorker.bind(this), prefetchProfile: PREFETCH_PROFILES.network, }); } /** @internal */ _getFileSize(): number | null | undefined { if (!this._fileSizeDetermined) { return this._length !== null ? this._length : undefined; } const baseSize = this._orchestrator.fileSize; if (baseSize === null) { return this._length !== null ? this._length : null; } return clamp(baseSize - this._offset, 0, this._length ?? Infinity); } /** @internal */ _read( start: number, end: number, minReadPosition: number, maxReadPosition: number, ): MaybePromise { if (this._length !== null && end > this._length) { return null; } const offset = this._offset; const result = this._orchestrator.read( offset + start, offset + end, Math.max(offset + minReadPosition, offset), offset + Math.min(maxReadPosition, this._length ?? Infinity), ); const processResult = (result: ReadResult | null) => { if (!result) { return null; } result.offset -= this._offset; return result; }; if (result instanceof Promise) { return result.then(processResult); } else { return processResult(result); } } /** @internal */ private async _runWorker(worker: ReadWorker) { // The outer loop is for resuming a request if it dies mid-response while (true) { const abortController = new AbortController(); const response = await retriedFetch( this._options.fetchFn ?? fetch, this._url, mergeRequestInit(this._requestInit, { headers: { // Always sending a range request is a good way to probe if the server supports them Range: `bytes=${worker.currentPos}-`, }, signal: abortController.signal, }), this._getRetryDelay, () => this._disposed, ); if (!response.ok) { // eslint-disable-next-line @typescript-eslint/no-base-to-string throw new Error(`Error fetching ${String(this._url)}: ${response.status} ${response.statusText}`); } if (response.redirected) { // Modify our own root path so that future subrequests get made relative to the redirected URL this.rootPath = response.url; } outer: if (this._orchestrator.fileSize === null) { // See if we can deduce the file size from the response const contentRange = response.headers.get('Content-Range'); if (contentRange) { const match = /\/(\d+)/.exec(contentRange); if (match) { this._orchestrator.supplyFileSize(Number(match[1])); break outer; } } const contentLength = response.headers.get('Content-Length'); if (contentLength) { // Note: For range requests, this is _technically_ not correct, as the range response could contain // less data than was requested. In practice, it seems most servers don't do this though, and the // Content-Length header actually contains the length until the end of the file. this._orchestrator.supplyFileSize(worker.currentPos + Number(contentLength)); } } this._fileSizeDetermined = true; // Yes, this is correct even if file size is still null if (response.status !== 206) { if (!this._usedForHls) { const url = new URL( this._url instanceof Request ? this._url.url : this._url, typeof window !== 'undefined' ? window.location.href : undefined, ); if ( url.origin !== 'null' // Don't show the warning for M3U8 playlist files, it's irrelevant for those && !(url.pathname.endsWith('.m3u8') || url.pathname.endsWith('.m3u')) ) { if (!warnedOrigins.has(url.origin)) { Logging._warn( `HTTP server (origin ${url.origin}) did not respond to a range request with 206 Partial` + ' Content, meaning the entire resource will now be downloaded. To enable efficient' + ' media file streaming across a network, please make sure your server supports' + ' range requests.', ); warnedOrigins.add(url.origin); } } } worker.currentPos = 0; this._orchestrator.options.maxCacheSize = Infinity; // 🤷 if (this._orchestrator.fileSize !== null) { worker.targetPos = this._orchestrator.fileSize; } else { // The server is dumb, doesn't even surface the content length, but we'll work with it. worker.targetPos = Infinity; worker.strictTarget = false; } this._orchestrator.consolidateEverythingIntoOneWorker(worker); } if (!response.body) { throw new Error( 'Missing HTTP response body stream. The used fetch function must provide the response body as a' + ' ReadableStream.', ); } const reader = response.body.getReader(); while (true) { if (worker.currentPos >= worker.targetPos || worker.aborted) { abortController.abort(); this._orchestrator.signalWorkerStoppedRunning(worker); return; } let readResult: ReadableStreamReadResult; try { readResult = await reader.read(); } catch (error) { if (this._disposed) { // No need to try to retry throw error; } const retryDelayInSeconds = this._getRetryDelay(1, error, this._url); if (retryDelayInSeconds !== null) { Logging._error('Error while reading response stream. Attempting to resume.', error); await wait(1000 * retryDelayInSeconds); break; } else { throw error; } } if (worker.aborted) { continue; // Cleanup happens in next iteration } const { done, value } = readResult; if (done) { if (worker.currentPos >= worker.targetPos) { // All data was delivered, we're good this._orchestrator.onWorkerFinished(worker); return; } if (worker.strictTarget) { // The response stopped early, before the target. This can happen if server decides to cap range // requests arbitrarily, even if the request had an uncapped end. In this case, let's fetch the // rest of the data using a new request. break; } else { // Assume we have simply reached the end of the resource this._orchestrator.onWorkerFinished(worker); return; } } this._dispatchRead(worker.currentPos, worker.currentPos + value.length); this._orchestrator.supplyWorkerData(worker, value); } } // The previous UrlSource had logic for circumventing https://issues.chromium.org/issues/436025873; I haven't // been able to observe this bug with the new UrlSource (maybe because we're using response streaming), so the // logic for that has vanished for now. Leaving a comment here if this becomes relevant again. } /** @internal */ _dispose() { this._orchestrator.dispose(); } } const BYTE_RANGE_REGEX = /^bytes=(\d+)-(\d*)$/; const parseByteRangeHeader = (value: string) => { const match = BYTE_RANGE_REGEX.exec(value.trim()); if (!match) { return null; } const offset = Number(match[1]); const end = match[2] === '' ? null : Number(match[2]); if (end !== null && end < offset) { return null; } return { offset, length: end !== null ? end - offset + 1 : null, }; }; /** * Options for {@link FilePathSource}. * @group Input sources * @public */ export type FilePathSourceOptions = { /** The maximum number of bytes the cache is allowed to hold in memory. Defaults to 8 MiB. */ maxCacheSize?: number; }; /** * A source backed by a path to a file. Intended for server-side usage in Node, Bun, or Deno. * * Make sure to call `.dispose()` on the corresponding {@link Input} when done to explicitly free the internal file * handle acquired by this source. * @group Input sources * @public */ export class FilePathSource extends PathedSource { /** @internal */ _customSource: CustomSource; /** @internal */ _fileHandle: FileHandle | null = null; /** Creates a new {@link FilePathSource} backed by the file at the specified file path. */ constructor(filePath: string, options: FilePathSourceOptions = {}) { if (typeof filePath !== 'string') { throw new TypeError('filePath must be a string.'); } if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if ( options.maxCacheSize !== undefined && (!isNumber(options.maxCacheSize) || options.maxCacheSize < 0) ) { throw new TypeError('options.maxCacheSize, when provided, must be a non-negative number.'); } if (!node.fs) { throw new Error( 'FilePathSource is only available in server-side environments (Node.js, Bun, Deno).', ); } super(filePath, request => new FilePathSource(request.path, options)); // Let's back this source with a CustomSource, makes the implementation very simple this._customSource = new CustomSource({ getSize: async () => { const fileHandle = await node.fs.open(filePath, 'r'); this._fileHandle = fileHandle; sourceFinalizationRegistry?.register(this, () => { // If it's not closed, Node prints annoying warnings void fileHandle.close(); }, this); const stats = await fileHandle.stat(); return stats.size; }, read: async (start, end) => { assert(this._fileHandle); const buffer = new Uint8Array(end - start); await this._fileHandle.read(buffer, 0, end - start, start); return buffer; }, maxCacheSize: options.maxCacheSize, prefetchProfile: 'fileSystem', }); } /** @internal */ _read( start: number, end: number, minReadPosition: number, maxReadPosition: number, ): MaybePromise { return this._customSource._read(start, end, minReadPosition, maxReadPosition); } /** @internal */ _getFileSize(): number | null | undefined { return this._customSource._getFileSize(); } /** @internal */ _dispose() { this._customSource._dispose(); if (this._fileHandle) { void this._fileHandle.close(); this._fileHandle = null; sourceFinalizationRegistry?.unregister(this); } } } /** * Options for defining a {@link CustomSource}. * @group Input sources * @public */ export type CustomSourceOptions = { /** * Called when the size of the entire file is requested. Must return or resolve to the size in bytes. This function * is guaranteed to be called before `read`. */ getSize: () => MaybePromise; /** * Called when data is requested. Must return or resolve to the bytes from the specified byte range, or a stream * that yields these bytes. * * You are guaranteed that `0 <= start < end < fileSize`. */ read: (start: number, end: number) => MaybePromise>; /** * Called when the {@link Input} driven by this source is disposed. */ dispose?: () => unknown; /** The maximum number of bytes the cache is allowed to hold in memory. Defaults to 8 MiB. */ maxCacheSize?: number; /** * Specifies the prefetch profile that the reader should use with this source. A prefetch profile specifies the * pattern with which bytes outside of the requested range are preloaded to reduce latency for future reads. * * - `'none'` (default): No prefetching; only the data needed in the moment is requested. * - `'fileSystem'`: File system-optimized prefetching: a small amount of data is prefetched bidirectionally, * aligned with page boundaries. * - `'network'`: Network-optimized prefetching, or more generally, prefetching optimized for any high-latency * environment: tries to minimize the amount of read calls and aggressively prefetches data when sequential access * patterns are detected. */ prefetchProfile?: 'none' | 'fileSystem' | 'network'; }; /** * A general-purpose, callback-driven source that can get its data from anywhere. Use this source to implement your own * custom source if the other sources don't cover your case. * @group Input sources * @public */ export class CustomSource extends Source { /** @internal */ _options: CustomSourceOptions; /** @internal */ _orchestrator: ReadOrchestrator; /** Creates a new {@link CustomSource} whose behavior is specified by `options`. */ constructor(options: CustomSourceOptions) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (typeof options.getSize !== 'function') { throw new TypeError('options.getSize must be a function.'); } if (typeof options.read !== 'function') { throw new TypeError('options.read must be a function.'); } if (options.dispose !== undefined && typeof options.dispose !== 'function') { throw new TypeError('options.dispose, when provided, must be a function.'); } if ( options.maxCacheSize !== undefined && (!isNumber(options.maxCacheSize) || options.maxCacheSize < 0) ) { throw new TypeError('options.maxCacheSize, when provided, must be a non-negative number.'); } if (options.prefetchProfile && !['none', 'fileSystem', 'network'].includes(options.prefetchProfile)) { throw new TypeError( 'options.prefetchProfile, when provided, must be one of \'none\', \'fileSystem\' or \'network\'.', ); } super(); this._options = options; this._orchestrator = new ReadOrchestrator({ maxCacheSize: options.maxCacheSize ?? (8 * 2 ** 20 /* 8 MiB */), maxWorkerCount: 2, // Fixed for now, *should* be fine prefetchProfile: PREFETCH_PROFILES[options.prefetchProfile ?? 'none'], runWorker: this._runWorker.bind(this), }); } /** @internal */ _getFileSize(): number | null | undefined { return this._orchestrator.fileSize ?? undefined; } /** @internal */ _read( start: number, end: number, minReadPosition: number, maxReadPosition: number, ): MaybePromise { if (this._orchestrator.fileSize !== null) { return this._orchestrator.read(start, end, minReadPosition, maxReadPosition); } const result = this._options.getSize(); if (result instanceof Promise) { return result.then((size) => { if (!Number.isInteger(size) || size < 0) { throw new TypeError('options.getSize must return or resolve to a non-negative integer.'); } this._orchestrator.fileSize = size; return this._orchestrator.read(start, end, minReadPosition, maxReadPosition); }); } else { if (!Number.isInteger(result) || result < 0) { throw new TypeError('options.getSize must return or resolve to a non-negative integer.'); } this._orchestrator.fileSize = result; return this._orchestrator.read(start, end, minReadPosition, maxReadPosition); } } /** @internal */ private async _runWorker(worker: ReadWorker) { while (worker.currentPos < worker.targetPos && !worker.aborted) { const originalCurrentPos = worker.currentPos; const originalTargetPos = worker.targetPos; let data = this._options.read(worker.currentPos, originalTargetPos); if (data instanceof Promise) data = await data; if (worker.aborted) { break; } if (data instanceof Uint8Array) { data = toUint8Array(data); // Normalize things like Node.js Buffer to Uint8Array if (data.length !== originalTargetPos - worker.currentPos) { // Yes, we're that strict throw new Error( `options.read returned a Uint8Array with unexpected length: Requested ${ originalTargetPos - worker.currentPos } bytes, but got ${data.length}.`, ); } this._dispatchRead(worker.currentPos, worker.currentPos + data.length); this._orchestrator.supplyWorkerData(worker, data); } else if (data instanceof ReadableStream) { const reader = data.getReader(); while (worker.currentPos < originalTargetPos && !worker.aborted) { const { done, value } = await reader.read(); if (done) { if (worker.currentPos < originalTargetPos) { // Yes, we're *that* strict throw new Error( `ReadableStream returned by options.read ended before supplying enough data.` + ` Requested ${originalTargetPos - originalCurrentPos} bytes, but got ${ worker.currentPos - originalCurrentPos }`, ); } break; } if (!(value instanceof Uint8Array)) { throw new TypeError('ReadableStream returned by options.read must yield Uint8Array chunks.'); } if (worker.aborted) { break; } const data = toUint8Array(value); // Normalize things like Node.js Buffer to Uint8Array this._dispatchRead(worker.currentPos, worker.currentPos + data.length); this._orchestrator.supplyWorkerData(worker, data); } } else { throw new TypeError('options.read must return or resolve to a Uint8Array or a ReadableStream.'); } } this._orchestrator.signalWorkerStoppedRunning(worker); } /** @internal */ _dispose() { this._orchestrator.dispose(); this._options.dispose?.(); } } /** * An alias for {@link CustomSource}. * @deprecated This name is misleading and will be removed in a future release. Please use {@link CustomSource} instead. * * @group Input sources * @public */ export const StreamSource = CustomSource; /** * An alias for {@link CustomSourceOptions}. * @deprecated This name is misleading and will be removed in a future release. Please use * {@link CustomSourceOptions} instead. * * @group Input sources * @public */ export type StreamSourceOptions = CustomSourceOptions; type ReadableStreamSourcePendingSlice = { start: number; end: number; bytes: Uint8Array; resolve: (bytes: ReadResult | null) => void; reject: (error: unknown) => void; }; /** * Options for {@link ReadableStreamSource}. * @group Input sources * @public */ export type ReadableStreamSourceOptions = { /** The maximum number of bytes the cache is allowed to hold in memory. Defaults to 32 MiB. */ maxCacheSize?: number; }; /** * A source backed by a [`ReadableStream`](https://developer.mozilla.org/en-US/docs/Web/API/ReadableStream) of * `Uint8Array`, representing an append-only byte stream of unknown length. This is the source to use for incrementally * streaming in input files that are still being constructed and whose size we don't yet know, like for example the * output chunks of [MediaRecorder](https://developer.mozilla.org/en-US/docs/Web/API/MediaRecorder). * * This source is *unsized*, meaning calls to `.getSize()` will throw and readers are more limited due to the * lack of random file access. You should only use this source with sequential access patterns, such as reading all * packets from start to end. This source does not work well with random access patterns unless you increase its * max cache size. * * @group Input sources * @public */ export class ReadableStreamSource extends Source { /** @internal */ _stream: ReadableStream; /** @internal */ _reader: ReadableStreamDefaultReader | null = null; /** @internal */ _cache: CacheEntry[] = []; /** @internal */ _maxCacheSize: number; /** @internal */ _pendingSlices: ReadableStreamSourcePendingSlice[] = []; /** @internal */ _currentIndex = 0; /** @internal */ _targetIndex = 0; /** @internal */ _maxRequestedIndex = 0; /** @internal */ _endIndex: number | null = null; /** @internal */ _pulling = false; /** Creates a new {@link ReadableStreamSource} backed by the specified `ReadableStream`. */ constructor(stream: ReadableStream, options: ReadableStreamSourceOptions = {}) { if (!(stream instanceof ReadableStream)) { throw new TypeError('stream must be a ReadableStream.'); } if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if ( options.maxCacheSize !== undefined && (!isNumber(options.maxCacheSize) || options.maxCacheSize < 0) ) { throw new TypeError('options.maxCacheSize, when provided, must be a non-negative number.'); } super(); this._stream = stream; this._maxCacheSize = options.maxCacheSize ?? (32 * 2 ** 20 /* 32 MiB */); } /** @internal */ _getFileSize(): number | null { return this._endIndex; // Starts out as null, meaning this source is unsized } /** @internal */ _read(start: number, end: number): MaybePromise { if (this._endIndex !== null && end > this._endIndex) { return null; } this._maxRequestedIndex = Math.max(this._maxRequestedIndex, end); const cacheStartIndex = binarySearchLessOrEqual(this._cache, start, x => x.start); const cacheStartEntry = cacheStartIndex !== -1 ? this._cache[cacheStartIndex]! : null; if (cacheStartEntry && cacheStartEntry.start <= start && end <= cacheStartEntry.end) { // The request can be satisfied with a single cache entry return { bytes: cacheStartEntry.bytes, view: cacheStartEntry.view, offset: cacheStartEntry.start, }; } let lastEnd = start; const bytes = new Uint8Array(end - start); if (cacheStartIndex !== -1) { // Walk over the cache to see if we can satisfy the request using multiple cache entries for (let i = cacheStartIndex; i < this._cache.length; i++) { const cacheEntry = this._cache[i]!; if (cacheEntry.start >= end) { break; } const cappedStart = Math.max(start, cacheEntry.start); if (cappedStart > lastEnd) { // We're too far behind this._throwDueToCacheMiss(); } const cappedEnd = Math.min(end, cacheEntry.end); if (cappedStart < cappedEnd) { bytes.set( cacheEntry.bytes.subarray(cappedStart - cacheEntry.start, cappedEnd - cacheEntry.start), cappedStart - start, ); lastEnd = cappedEnd; } } } if (lastEnd === end) { return { bytes, view: toDataView(bytes), offset: start, }; } // We need to pull more data if (this._currentIndex > lastEnd) { // We're too far behind this._throwDueToCacheMiss(); } const { promise, resolve, reject } = promiseWithResolvers(); this._pendingSlices.push({ start, end, bytes, resolve, reject, }); this._targetIndex = Math.max(this._targetIndex, end); // Start pulling from the stream if we're not already doing it if (!this._pulling) { this._pulling = true; void this._pull() .catch((error) => { this._pulling = false; if (this._pendingSlices.length > 0) { this._pendingSlices.forEach(x => x.reject(error)); // Make sure to propagate any errors this._pendingSlices.length = 0; } else { throw error; // So it doesn't get swallowed } }); } return promise; } /** @internal */ _throwDueToCacheMiss() { throw new Error( 'Read is before the cached region. With ReadableStreamSource, you must access the data more' + ' sequentially or increase the size of its cache.', ); } /** @internal */ async _pull() { this._reader ??= this._stream.getReader(); // This is the loop that keeps pulling data from the stream until a target index is reached, filling requests // in the process while (this._currentIndex < this._targetIndex && !this._disposed) { const { done, value } = await this._reader.read(); if (done) { for (const pendingSlice of this._pendingSlices) { pendingSlice.resolve(null); } this._pendingSlices.length = 0; this._endIndex = this._currentIndex; // We know how long the file is now! break; } const startIndex = this._currentIndex; const endIndex = this._currentIndex + value.byteLength; this._dispatchRead(startIndex, endIndex); // Fill the pending slices with the data for (let i = 0; i < this._pendingSlices.length; i++) { const pendingSlice = this._pendingSlices[i]!; const cappedStart = Math.max(startIndex, pendingSlice.start); const cappedEnd = Math.min(endIndex, pendingSlice.end); if (cappedStart < cappedEnd) { pendingSlice.bytes.set( value.subarray(cappedStart - startIndex, cappedEnd - startIndex), cappedStart - pendingSlice.start, ); if (cappedEnd === pendingSlice.end) { // Pending slice fully filled pendingSlice.resolve({ bytes: pendingSlice.bytes, view: toDataView(pendingSlice.bytes), offset: pendingSlice.start, }); this._pendingSlices.splice(i, 1); i--; } } } this._cache.push({ start: startIndex, end: endIndex, bytes: value, view: toDataView(value), age: 0, // Unused }); // Do cache eviction, based on the distance from the last-requested index. It's important that we do it like // this and not based on where the reader is at, because if the reader is fast, we'll unnecessarily evict // data that we still might need. while (this._cache.length > 0) { const firstEntry = this._cache[0]!; const distance = this._maxRequestedIndex - firstEntry.end; if (distance <= this._maxCacheSize) { break; } this._cache.shift(); } this._currentIndex += value.byteLength; } this._pulling = false; } /** @internal */ _dispose() { this._pendingSlices.length = 0; this._cache.length = 0; void this._reader?.cancel(); } } type PrefetchProfile = (start: number, end: number, workers: ReadWorker[]) => { start: number; end: number; }; const PREFETCH_PROFILES = { none: (start, end) => ({ start, end }), fileSystem: (start, end) => { const padding = 2 ** 16; start = Math.floor((start - padding) / padding) * padding; end = Math.ceil((end + padding) / padding) * padding; return { start, end }; }, network: (start, end, workers) => { // Add a slight bit of start padding because backwards reading is painful const paddingStart = 2 ** 16; start = Math.max(0, Math.floor((start - paddingStart) / paddingStart) * paddingStart); // Remote resources have extreme latency (relatively speaking), so the benefit from intelligent // prefetching is great. The network prefetch strategy is as follows: When we notice // successive reads to a worker's read region, we prefetch more data at the end of that region, // growing exponentially (up to a cap). This performs well for real-world use cases: Either we read a // small part of the file once and then never need it again, in which case the requested about of data // is small. Or, we're repeatedly doing a sequential access pattern (common in media files), in which // case we can become more and more confident to prefetch more and more data. for (const worker of workers) { const maxExtensionAmount = 8 * 2 ** 20; // 8 MiB // When the read region cross the threshold point, we trigger a prefetch. This point is typically // in the middle of the worker's read region, or a fixed offset from the end if the region has grown // really large. const thresholdPoint = Math.max( (worker.startPos + worker.targetPos) / 2, worker.targetPos - maxExtensionAmount, ); if (closedIntervalsOverlap( start, end, thresholdPoint, worker.targetPos, )) { const size = worker.targetPos - worker.startPos; // If we extend by maxExtensionAmount const a = Math.ceil((size + 1) / maxExtensionAmount) * maxExtensionAmount; // If we extend to the next power of 2 const b = 2 ** Math.ceil(Math.log2(size + 1)); const extent = Math.min(b, a); end = Math.max(end, worker.startPos + extent); } } end = Math.max(end, start + URL_SOURCE_MIN_LOAD_AMOUNT); return { start, end, }; }, } satisfies Record; type PendingSlice = { start: number; bytes: Uint8Array; holes: Hole[]; resolve: (bytes: Uint8Array | null) => void; reject: (error: unknown) => void; }; type Hole = { start: number; end: number; }; type CacheEntry = { start: number; end: number; bytes: Uint8Array; view: DataView; age: number; }; type ReadWorker = { startPos: number; currentPos: number; targetPos: number; /** The target is considered _strict_ when it is an error for the worker to terminate before reaching the target. */ strictTarget: boolean; running: boolean; aborted: boolean; pendingSlices: PendingSlice[]; age: number; }; /** * Godclass for orchestrating complex, cached read operations. The reading model is as follows: Any reading task is * delegated to a *worker*, which is a sequential reader positioned somewhere along the file. All workers run in * parallel and can be stopped and resumed in their forward movement. When read requests come in, this orchestrator will * first try to satisfy the request with only the cached data. If this isn't possible, workers are spun up for all * missing parts (or existing workers are repurposed), and these workers will then fill the holes in the data as they * march along the file. */ class ReadOrchestrator { fileSize: number | null = null; nextAge = 0; // Used for multiple things workers: ReadWorker[] = []; cache: CacheEntry[] = []; currentCacheSize = 0; disposed = false; queuedReads: { hole: Hole; strictTarget: boolean; pendingSlices: PendingSlice[]; age: number; }[] = []; constructor(public options: { maxCacheSize: number; runWorker: (worker: ReadWorker) => Promise; prefetchProfile: PrefetchProfile; maxWorkerCount: number; }) {} read( innerStart: number, innerEnd: number, minReadPosition: number, maxReadPosition: number, ): MaybePromise { assert(!this.disposed); const prefetchRange = this.options.prefetchProfile(innerStart, innerEnd, this.workers); const outerStart = Math.max(prefetchRange.start, minReadPosition); const outerEnd = Math.min(prefetchRange.end, this.fileSize ?? Infinity, maxReadPosition); assert(outerStart <= innerStart && innerEnd <= outerEnd); let result: MaybePromise | null = null; const innerCacheStartIndex = binarySearchLessOrEqual(this.cache, innerStart, x => x.start); const innerStartEntry = innerCacheStartIndex !== -1 ? this.cache[innerCacheStartIndex] : null; // See if the read request can be satisfied by a single cache entry if (innerStartEntry && innerStartEntry.start <= innerStart && innerEnd <= innerStartEntry.end) { innerStartEntry.age = this.nextAge++; result = { bytes: innerStartEntry.bytes, view: innerStartEntry.view, offset: innerStartEntry.start, }; // Can't return yet though, still need to check if the prefetch range might lie outside the cached area } const outerCacheStartIndex = binarySearchLessOrEqual(this.cache, outerStart, x => x.start); const bytes = result ? null : new Uint8Array(innerEnd - innerStart); let contiguousBytesWriteEnd = 0; // Used to track if the cache is able to completely cover the bytes let lastEnd = outerStart; // The "holes" in the cache (the parts we need to load) const outerHoles: Hole[] = []; // Loop over the cache and build up the list of holes if (outerCacheStartIndex !== -1) { for (let i = outerCacheStartIndex; i < this.cache.length; i++) { const entry = this.cache[i]!; if (entry.start >= outerEnd) { break; } if (entry.end <= outerStart) { continue; } const cappedOuterStart = Math.max(outerStart, entry.start); const cappedOuterEnd = Math.min(outerEnd, entry.end); assert(cappedOuterStart <= cappedOuterEnd); if (lastEnd < cappedOuterStart) { outerHoles.push({ start: lastEnd, end: cappedOuterStart }); } lastEnd = cappedOuterEnd; if (bytes) { const cappedInnerStart = Math.max(innerStart, entry.start); const cappedInnerEnd = Math.min(innerEnd, entry.end); if (cappedInnerStart < cappedInnerEnd) { const relativeOffset = cappedInnerStart - innerStart; // Fill the relevant section of the bytes with the cached data bytes.set( entry.bytes.subarray(cappedInnerStart - entry.start, cappedInnerEnd - entry.start), relativeOffset, ); if (relativeOffset === contiguousBytesWriteEnd) { contiguousBytesWriteEnd = cappedInnerEnd - innerStart; } } } entry.age = this.nextAge++; } if (lastEnd < outerEnd) { outerHoles.push({ start: lastEnd, end: outerEnd }); } } else { outerHoles.push({ start: outerStart, end: outerEnd }); } if (bytes && contiguousBytesWriteEnd >= bytes.length) { // Multiple cache entries were able to completely cover the requested bytes! result = { bytes, view: toDataView(bytes), offset: innerStart, }; } if (outerHoles.length === 0) { assert(result); return result; } // We need to read more data, so now we're in async land const { promise, resolve, reject } = promiseWithResolvers(); const innerHoles: typeof outerHoles = []; for (const outerHole of outerHoles) { const cappedStart = Math.max(innerStart, outerHole.start); const cappedEnd = Math.min(innerEnd, outerHole.end); if (cappedStart === outerHole.start && cappedEnd === outerHole.end) { innerHoles.push(outerHole); // Can reuse without allocating a new object } else if (cappedStart < cappedEnd) { innerHoles.push({ start: cappedStart, end: cappedEnd }); } } const pendingSlice: PendingSlice | null = bytes && { start: innerStart, bytes, holes: innerHoles, resolve, reject, }; // Fire off workers to take care of patching the holes outer: for (const outerHole of outerHoles) { for (const worker of this.workers) { const addedToWorker = this.checkHoleAgainstWorker( worker, outerHole, pendingSlice ? [pendingSlice] : [], ); if (addedToWorker) { this.checkQueuedReadsAgainstWorker(worker); continue outer; } } // We need to spawn a new worker const strictTarget = outerHole.end < outerEnd || this.fileSize !== null; const newWorker = this.createWorker(outerHole.start, outerHole.end, strictTarget); if (newWorker) { if (pendingSlice) { newWorker.pendingSlices = [pendingSlice]; } this.runWorker(newWorker); } else { // Max worker count has been reached, let's queue a read for later let index = binarySearchLessOrEqual(this.queuedReads, outerHole.start, x => x.hole.start); let entry = index !== -1 ? this.queuedReads[index]! : null; if (entry && outerHole.start <= entry.hole.end) { entry.hole.end = Math.max(entry.hole.end, outerHole.end); entry.strictTarget &&= strictTarget; if (pendingSlice) { entry.pendingSlices.push(pendingSlice); } } else { index++; entry = { hole: { // Clone the hole because it might be mutated later start: outerHole.start, end: outerHole.end, }, strictTarget, pendingSlices: pendingSlice ? [pendingSlice] : [], age: this.nextAge++, }; this.queuedReads.splice(index, 0, entry); } // Merge with any subsequent entries that overlap while (index + 1 < this.queuedReads.length) { const nextEntry = this.queuedReads[index + 1]!; if (nextEntry.hole.start > entry.hole.end) { break; } entry.hole.end = Math.max(entry.hole.end, nextEntry.hole.end); entry.pendingSlices.push(...nextEntry.pendingSlices); entry.strictTarget &&= nextEntry.strictTarget; entry.age = Math.min(entry.age, nextEntry.age); this.queuedReads.splice(index + 1, 1); } } } if (!result) { assert(bytes); result = promise.then(bytes => bytes && ({ bytes, view: toDataView(bytes), offset: innerStart, } satisfies ReadResult)); } else { // The requested region was satisfied by the cache, but the entire prefetch region was not promise.catch((error) => { if (this.disposed) { return; // Swallow the error } // Nobody's awaiting this result but an errored read is still notable throw error; }); } return result; } checkHoleAgainstWorker(worker: ReadWorker, hole: Hole, pendingSlices: PendingSlice[]) { // A small tolerance in the case that the requested region is *just* after the target position of an // existing worker. In that case, it's probably more efficient to repurpose that worker than to spawn // another one so close to it const gapTolerance = 2 ** 17; // This check also implies worker.currentPos <= hole.start, a critical condition if (closedIntervalsOverlap( hole.start - gapTolerance, hole.start, worker.currentPos, worker.targetPos, )) { worker.targetPos = Math.max(worker.targetPos, hole.end); // Update the worker's target position for (let i = 0; i < pendingSlices.length; i++) { const pendingSlice = pendingSlices[i]!; if (!worker.pendingSlices.includes(pendingSlice)) { worker.pendingSlices.push(pendingSlice); } } if (!worker.running) { // Kick it off if it's idle this.runWorker(worker); } return true; } return false; } checkQueuedReadsAgainstWorker(worker: ReadWorker) { let wasTrueOnce = false; for (let i = 0; i < this.queuedReads.length; i++) { const queuedRead = this.queuedReads[i]!; const result = this.checkHoleAgainstWorker(worker, queuedRead.hole, queuedRead.pendingSlices); if (result) { this.queuedReads.splice(i, 1); i--; wasTrueOnce = true; } else if (wasTrueOnce) { // We can stop since the holes are sorted break; } } } createWorker(startPos: number, targetPos: number, strictTarget: boolean) { if (this.workers.length >= this.options.maxWorkerCount) { let oldestWorker: ReadWorker | null = null; let oldestIndex: number | null = null; for (let i = 0; i < this.workers.length; i++) { const worker = this.workers[i]!; if ( !worker.running && worker.pendingSlices.length === 0 && (!oldestWorker || worker.age < oldestWorker.age) ) { oldestIndex = i; oldestWorker = worker; } } if (oldestWorker) { // LRU eviction assert(oldestIndex !== null); assert(oldestWorker.pendingSlices.length === 0); this.workers.splice(oldestIndex, 1); } else { return null; // All workers are still running, we can't create a new one } } const worker: ReadWorker = { startPos, currentPos: startPos, targetPos, strictTarget, running: false, // Due to async shenanigans, it can happen that workers are started after disposal. In this case, instead of // simply not creating the worker, we allow it to run but immediately label it as aborted, so it can then // shut itself down. aborted: this.disposed, pendingSlices: [], age: this.nextAge++, }; this.workers.push(worker); return worker; } runWorker(worker: ReadWorker) { assert(!worker.running); assert(worker.currentPos < worker.targetPos); worker.running = true; worker.age = this.nextAge++; void this.options.runWorker(worker) .catch((error) => { worker.running = false; if (worker.pendingSlices.length > 0) { worker.pendingSlices.forEach(x => x.reject(error)); // Make sure to propagate any errors worker.pendingSlices.length = 0; } else if (!worker.aborted && !this.disposed) { throw error; // So it doesn't get swallowed } }) .finally(() => { if (worker.running || this.workers.length >= this.options.maxWorkerCount) { // Rare, but can happen with multiple concurrent reads. In this case, don't do anything. return; } if (this.queuedReads.length > 0) { let oldestIndex = 0; for (let i = 1; i < this.queuedReads.length; i++) { const queuedRead = this.queuedReads[i]!; if (queuedRead.age < this.queuedReads[oldestIndex]!.age) { oldestIndex = i; } } const queuedRead = this.queuedReads[oldestIndex]!; this.queuedReads.splice(oldestIndex, 1); const newWorker = this.createWorker( queuedRead.hole.start, queuedRead.hole.end, queuedRead.strictTarget, ); assert(newWorker); // We just freed up a worker, so this should never fail newWorker.pendingSlices = queuedRead.pendingSlices; this.runWorker(newWorker); } }); } consolidateEverythingIntoOneWorker(worker: ReadWorker) { // Here we merge everything into one "megaworker" that spans the entire file. We assume the passed-in worker // is already configured to be a megaworker. const uniqueSlices = new Set(worker.pendingSlices); for (let i = 0; i < this.workers.length; i++) { const otherWorker = this.workers[i]!; if (otherWorker === worker) { continue; } for (const slice of otherWorker.pendingSlices) { uniqueSlices.add(slice); } otherWorker.aborted = true; otherWorker.pendingSlices.length = 0; this.workers.splice(i, 1); i--; } for (let i = 0; i < this.queuedReads.length; i++) { const queuedRead = this.queuedReads[i]!; for (const slice of queuedRead.pendingSlices) { uniqueSlices.add(slice); } } worker.pendingSlices = [...uniqueSlices]; this.queuedReads.length = 0; } /** Called by a worker when it has read some data. */ supplyWorkerData(worker: ReadWorker, bytes: Uint8Array) { assert(!worker.aborted); const start = worker.currentPos; const end = start + bytes.length; this.insertIntoCache({ start, end, bytes, view: toDataView(bytes), age: this.nextAge++, }); worker.currentPos += bytes.length; if (worker.currentPos > worker.targetPos) { // In case it overshoots worker.targetPos = worker.currentPos; this.checkQueuedReadsAgainstWorker(worker); } // Now, let's see if we can use the read bytes to fill any pending slice for (let i = 0; i < worker.pendingSlices.length; i++) { const pendingSlice = worker.pendingSlices[i]!; const clampedStart = Math.max(start, pendingSlice.start); const clampedEnd = Math.min(end, pendingSlice.start + pendingSlice.bytes.length); if (clampedStart < clampedEnd) { pendingSlice.bytes.set( bytes.subarray(clampedStart - start, clampedEnd - start), clampedStart - pendingSlice.start, ); } for (let j = 0; j < pendingSlice.holes.length; j++) { // The hole is intentionally not modified here if the read section starts somewhere in the middle of // the hole. We don't need to do "hole splitting", since the workers are spawned *by* the holes, // meaning there's always a worker which will consume the hole left to right. const hole = pendingSlice.holes[j]!; if (start <= hole.start && end > hole.start) { hole.start = end; } if (hole.end <= hole.start) { pendingSlice.holes.splice(j, 1); j--; } } if (pendingSlice.holes.length === 0) { // The slice has been fulfilled, everything has been read. Let's resolve the promise pendingSlice.resolve(pendingSlice.bytes); worker.pendingSlices.splice(i, 1); i--; } } // Remove other idle workers if we "ate" into their territory for (let i = 0; i < this.workers.length; i++) { const otherWorker = this.workers[i]!; if (worker === otherWorker || otherWorker.running) { continue; } if (closedIntervalsOverlap( start, end, otherWorker.currentPos, otherWorker.targetPos, // These should typically be equal when the worker's idle )) { this.workers.splice(i, 1); i--; } } } supplyFileSize(size: number) { assert(this.fileSize === null); this.fileSize = size; // Trim the workers with this new information for (const worker of this.workers) { worker.targetPos = Math.min(worker.targetPos, size); worker.strictTarget = true; for (let i = 0; i < worker.pendingSlices.length; i++) { const pendingSlice = worker.pendingSlices[i]!; for (const hole of pendingSlice.holes) { if (hole.end > size) { // Can't satisfy this slice anymore pendingSlice.resolve(null); worker.pendingSlices.splice(i, 1); i--; break; } } } } // Trim the queued reads as well for (let i = 0; i < this.queuedReads.length; i++) { const queuedRead = this.queuedReads[i]!; if (queuedRead.hole.start >= size) { // Entirely out of bounds for (const slice of queuedRead.pendingSlices) slice.resolve(null); this.queuedReads.splice(i, 1); i--; } else if (queuedRead.hole.end > size) { // Partially out of bounds queuedRead.hole.end = size; queuedRead.strictTarget = true; for (let j = 0; j < queuedRead.pendingSlices.length; j++) { const slice = queuedRead.pendingSlices[j]!; // If the slice itself is out of bounds, resolve it if (slice.start >= size) { slice.resolve(null); queuedRead.pendingSlices.splice(j, 1); j--; } } } } } signalWorkerStoppedRunning(worker: ReadWorker) { worker.running = false; // When a worker stops running, that means it has hit its targetPos. It might still have pendingSlices assigned, // but this is because those pending slices cover data that other workers are assigned to fill. Since targetPos // has been reached, we can confidently say that this worker has completed its share of work on the pending // slices and must no longer care about them. worker.pendingSlices.length = 0; } /** Called when a worker reaches the end of the underlying data and must be cleaned up. */ onWorkerFinished(worker: ReadWorker) { const index = this.workers.indexOf(worker); assert(index !== -1); worker.running = false; this.workers.splice(index, 1); if (this.fileSize === null) { // We can now deduce the file size! this.supplyFileSize(worker.currentPos); } for (const pendingSlice of worker.pendingSlices) { pendingSlice.resolve(null); } } insertIntoCache(entry: CacheEntry) { if (this.options.maxCacheSize === 0) { return; // No caching } let insertionIndex = binarySearchLessOrEqual(this.cache, entry.start, x => x.start) + 1; if (insertionIndex > 0) { const previous = this.cache[insertionIndex - 1]!; if (previous.end >= entry.end) { // Previous entry swallows the one to be inserted; we don't need to do anything return; } if (previous.end > entry.start) { // Partial overlap with the previous entry, let's join const joined = new Uint8Array(entry.end - previous.start); joined.set(previous.bytes, 0); joined.set(entry.bytes, entry.start - previous.start); this.currentCacheSize += entry.end - previous.end; previous.bytes = joined; previous.view = toDataView(joined); previous.end = entry.end; // Do the rest of the logic with the previous entry instead insertionIndex--; entry = previous; } else { this.cache.splice(insertionIndex, 0, entry); this.currentCacheSize += entry.bytes.length; } } else { this.cache.splice(insertionIndex, 0, entry); this.currentCacheSize += entry.bytes.length; } for (let i = insertionIndex + 1; i < this.cache.length; i++) { const next = this.cache[i]!; if (entry.end <= next.start) { // Even if they touch, we don't wanna merge them, no need break; } if (entry.end >= next.end) { // The inserted entry completely swallows the next entry this.cache.splice(i, 1); this.currentCacheSize -= next.bytes.length; i--; continue; } // Partial overlap, let's join const joined = new Uint8Array(next.end - entry.start); joined.set(entry.bytes, 0); joined.set(next.bytes, next.start - entry.start); this.currentCacheSize -= entry.end - next.start; // Subtract the overlap entry.bytes = joined; entry.view = toDataView(joined); entry.end = next.end; this.cache.splice(i, 1); break; // After the join case, we're done: the next entry cannot possibly overlap with the inserted one. } // LRU eviction of cache entries while (this.currentCacheSize > this.options.maxCacheSize) { let oldestIndex = 0; let oldestEntry = this.cache[0]!; for (let i = 1; i < this.cache.length; i++) { const entry = this.cache[i]!; if (entry.age < oldestEntry.age) { oldestIndex = i; oldestEntry = entry; } } if (this.currentCacheSize - oldestEntry.bytes.length <= this.options.maxCacheSize) { // Don't evict if it would shrink the cache below the max size break; } this.cache.splice(oldestIndex, 1); this.currentCacheSize -= oldestEntry.bytes.length; } } dispose() { for (const worker of this.workers) { worker.aborted = true; } this.workers.length = 0; this.cache.length = 0; this.disposed = true; } } /** * A dummy source from which no data can be read. Can be used in conjunction with input formats that get their data * from another source. */ export class NullSource extends Source { override _getFileSize(): number | null { return null; } override _read(): MaybePromise { return null; } override _dispose(): void { // Do nothing } } /** * A source that covers a range (offset + length) of another source. Useful for reading files that are embedded within * larger files. * * @group Input sources * @public */ export class RangedSource extends Source { /** @internal */ _baseSource: Source; /** @internal */ _ref: SourceRef | null = null; /** @internal */ _offset: number; /** @internal */ _length: number | null; /** @internal */ constructor(baseSource: Source, offset: number, length?: number) { super(); if (baseSource._disposed) { throw new Error('Cannot create a slice of a disposed source.'); } this._baseSource = baseSource; this._offset = offset; this._length = length ?? null; } /** @internal */ override _getFileSize(): number | null | undefined { const baseSize = this._baseSource._getFileSize(); if (baseSize === undefined) { return this._length !== null ? this._length : undefined; } if (baseSize === null) { if (this._length !== null) { return this._length; } else { return null; } } return clamp(baseSize - this._offset, 0, this._length ?? Infinity); } /** @internal */ override _read( start: number, end: number, minReadPosition: number, maxReadPosition: number, ): MaybePromise { if (this._length !== null && end > this._length) { return null; } const result = this._baseSource._read( this._offset + start, this._offset + end, this._offset + minReadPosition, this._offset + maxReadPosition, ); const processResult = (result: ReadResult | null) => { if (!result) { return null; } result.offset -= this._offset; return result; }; if (result instanceof Promise) { return result.then(processResult); } else { return processResult(result); } } /** @internal */ override _dispose(): void { this._ref?.free(); } override ref() { this._ref ??= this._baseSource.ref(); return super.ref(); } } ===== src/decode.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AUDIO_CODECS, AudioCodec, buildAudioCodecString, buildVideoCodecString, guessDescriptionForAudio, guessDescriptionForVideo, inferCodecFromCodecString, MediaCodec, PCM_AUDIO_CODECS, VIDEO_CODECS, VideoCodec, } from './codec'; import { customAudioDecoders, customVideoDecoders } from './custom-coder'; import { isAllowSharedBufferSource, SetOptional } from './misc'; export const canDecodeVideoMemo = new Map>(); export const canDecodeAudioMemo = new Map>(); const validateVideoDecodingConfig = (codec: VideoCodec, options: SetOptional) => { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.codec !== undefined && typeof options.codec !== 'string') { throw new TypeError('options.codec, when provided, must be a string.'); } if (options.codec !== undefined && inferCodecFromCodecString(options.codec) !== codec) { throw new TypeError(`options.codec, when provided, must match the specified codec (${codec}).`); } if ( options.codedWidth !== undefined && (!Number.isInteger(options.codedWidth) || options.codedWidth <= 0) ) { throw new TypeError('options.codedWidth, when provided, must be a positive integer.'); } if ( options.codedHeight !== undefined && (!Number.isInteger(options.codedHeight) || options.codedHeight <= 0) ) { throw new TypeError('options.codedHeight, when provided, must be a positive integer.'); } if ( options.displayAspectWidth !== undefined && (!Number.isInteger(options.displayAspectWidth) || options.displayAspectWidth <= 0) ) { throw new TypeError('options.displayAspectWidth, when provided, must be a positive integer.'); } if ( options.displayAspectHeight !== undefined && (!Number.isInteger(options.displayAspectHeight) || options.displayAspectHeight <= 0) ) { throw new TypeError('options.displayAspectHeight, when provided, must be a positive integer.'); } if (options.description !== undefined && !isAllowSharedBufferSource(options.description)) { throw new TypeError('options.description, when provided, must be a buffer source.'); } if ( options.hardwareAcceleration !== undefined && !['no-preference', 'prefer-hardware', 'prefer-software'].includes(options.hardwareAcceleration) ) { throw new TypeError( 'options.hardwareAcceleration, when provided, must be \'no-preference\', \'prefer-hardware\' or' + ' \'prefer-software\'.', ); } if (options.optimizeForLatency !== undefined && typeof options.optimizeForLatency !== 'boolean') { throw new TypeError('options.optimizeForLatency, when provided, must be a boolean.'); } }; const validateAudioDecodingConfig = ( codec: AudioCodec, options: SetOptional, ) => { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.codec !== undefined && typeof options.codec !== 'string') { throw new TypeError('options.codec, when provided, must be a string.'); } if (options.codec !== undefined && inferCodecFromCodecString(options.codec) !== codec) { throw new TypeError(`options.codec, when provided, must match the specified codec (${codec}).`); } if ( options.numberOfChannels !== undefined && (!Number.isInteger(options.numberOfChannels) || options.numberOfChannels <= 0) ) { throw new TypeError('options.numberOfChannels, when provided, must be a positive integer.'); } if ( options.sampleRate !== undefined && (!Number.isInteger(options.sampleRate) || options.sampleRate <= 0) ) { throw new TypeError('options.sampleRate, when provided, must be a positive integer.'); } if (options.description !== undefined && !isAllowSharedBufferSource(options.description)) { throw new TypeError('options.description, when provided, must be a buffer source.'); } }; /** * Checks if the browser is able to decode the given codec. * @group Decoding * @public */ export const canDecode = (codec: MediaCodec) => { if ((VIDEO_CODECS as readonly string[]).includes(codec)) { return canDecodeVideo(codec as VideoCodec); } else if ((AUDIO_CODECS as readonly string[]).includes(codec)) { return canDecodeAudio(codec as AudioCodec); } return false; }; /** * Checks if the browser is able to decode the given video codec with the given parameters. * @group Decoding * @public */ export const canDecodeVideo = async ( codec: VideoCodec, options: SetOptional = {}, ) => { if (!VIDEO_CODECS.includes(codec)) { return false; } validateVideoDecodingConfig(codec, options); const resolvedOptions: VideoDecoderConfig = { ...options, codedWidth: options.codedWidth ?? 1280, codedHeight: options.codedHeight ?? 720, codec: options.codec ?? buildVideoCodecString(codec, 1280, 720, 1e6), }; resolvedOptions.description ??= guessDescriptionForVideo(resolvedOptions); const key = JSON.stringify(resolvedOptions); const memoized = canDecodeVideoMemo.get(key); if (memoized) { return memoized; } const promise = (async () => { if (customVideoDecoders.some(x => x.supports(codec, resolvedOptions))) { return true; } if (typeof VideoDecoder === 'undefined') { return false; } const support = await VideoDecoder.isConfigSupported(resolvedOptions); return support.supported === true; })(); canDecodeVideoMemo.set(key, promise); return promise; }; /** * Checks if the browser is able to decode the given audio codec with the given parameters. * @group Decoding * @public */ export const canDecodeAudio = async ( codec: AudioCodec, options: SetOptional = {}, ) => { if (!AUDIO_CODECS.includes(codec)) { return false; } validateAudioDecodingConfig(codec, options); const resolvedOptions: AudioDecoderConfig = { ...options, numberOfChannels: options.numberOfChannels ?? 2, sampleRate: options.sampleRate ?? 48000, codec: options.codec ?? buildAudioCodecString(codec, 2, 48000), }; if (resolvedOptions.description === undefined) { const generatedDescription = guessDescriptionForAudio(resolvedOptions); if (generatedDescription === false) { return false; } resolvedOptions.description = generatedDescription; } const key = JSON.stringify(resolvedOptions); const memoized = canDecodeAudioMemo.get(key); if (memoized) { return memoized; } const promise = (async () => { if (customAudioDecoders.some(x => x.supports(codec, resolvedOptions))) { return true; } if ((PCM_AUDIO_CODECS as readonly string[]).includes(codec)) { return true; } if (typeof AudioDecoder === 'undefined') { return false; } const support = await AudioDecoder.isConfigSupported(resolvedOptions); return support.supported === true; })(); canDecodeAudioMemo.set(key, promise); return promise; }; /** * Returns the list of all media codecs that can be decoded by the browser. * @group Decoding * @public */ export const getDecodableCodecs = async (): Promise => { const [videoCodecs, audioCodecs] = await Promise.all([ getDecodableVideoCodecs(), getDecodableAudioCodecs(), ]); return [...videoCodecs, ...audioCodecs]; }; /** * Returns the list of all video codecs that can be decoded by the browser. * @group Decoding * @public */ export const getDecodableVideoCodecs = async ( checkedCodecs: VideoCodec[] = VIDEO_CODECS as unknown as VideoCodec[], options?: SetOptional, ): Promise => { const bools = await Promise.all(checkedCodecs.map(codec => canDecodeVideo(codec, options))); return checkedCodecs.filter((_, i) => bools[i]); }; /** * Returns the list of all audio codecs that can be decoded by the browser. * @group Decoding * @public */ export const getDecodableAudioCodecs = async ( checkedCodecs: AudioCodec[] = AUDIO_CODECS as unknown as AudioCodec[], options?: SetOptional, ): Promise => { const bools = await Promise.all(checkedCodecs.map(codec => canDecodeAudio(codec, options))); return checkedCodecs.filter((_, i) => bools[i]); }; ===== src/demuxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Input } from './input'; import { InputTrackBacking } from './input-track'; import { MetadataTags } from './metadata'; /** * Options for retrieving media duration from metadata. * @group Input files & tracks * @public */ export type DurationMetadataRequestOptions = { /** * When the underlying media is live, querying the duration will, by default, wait until the live stream has ended. * Setting this field to `true` skips that wait and returns the current duration of the stream. When the media isn't * live, this field has no effect. * * See also {@link PacketRetrievalOptions.skipLiveWait}. */ skipLiveWait?: boolean; }; export abstract class Demuxer { input: Input; constructor(input: Input) { this.input = input; } abstract getTrackBackings(): Promise; abstract getMimeType(): Promise; abstract getMetadataTags(): Promise; dispose() { // Can be overridden } } ===== src/adts/adts-reader.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Bitstream } from '../../shared/bitstream'; import { FileSlice, readBytes } from '../reader'; export const MIN_ADTS_FRAME_HEADER_SIZE = 7; export const MAX_ADTS_FRAME_HEADER_SIZE = 9; export type AdtsFrameHeader = { objectType: number; samplingFrequencyIndex: number; channelConfiguration: number; frameLength: number; numberOfAacFrames: number; crcCheck: number | null; startPos: number; }; export const readAdtsFrameHeader = (slice: FileSlice): AdtsFrameHeader | null => { // https://wiki.multimedia.cx/index.php/ADTS (last visited: 2025/08/17) const startPos = slice.filePos; const bytes = readBytes(slice, 9); // 9 with CRC, 7 without CRC const bitstream = new Bitstream(bytes); const syncword = bitstream.readBits(12); if (syncword !== 0b1111_11111111) { return null; } bitstream.skipBits(1); // MPEG version const layer = bitstream.readBits(2); if (layer !== 0) { return null; } const protectionAbsence = bitstream.readBits(1); const objectType = bitstream.readBits(2) + 1; const samplingFrequencyIndex = bitstream.readBits(4); if (samplingFrequencyIndex === 15) { return null; } bitstream.skipBits(1); // Private bit const channelConfiguration = bitstream.readBits(3); if (channelConfiguration === 0) { throw new Error('ADTS frames with channel configuration 0 are not supported.'); } bitstream.skipBits(1); // Originality bitstream.skipBits(1); // Home bitstream.skipBits(1); // Copyright ID bit bitstream.skipBits(1); // Copyright ID start const frameLength = bitstream.readBits(13); bitstream.skipBits(11); // Buffer fullness const numberOfAacFrames = bitstream.readBits(2) + 1; if (numberOfAacFrames !== 1) { throw new Error('ADTS frames with more than one AAC frame are not supported.'); } let crcCheck: number | null = null; if (protectionAbsence === 1) { // No CRC slice.filePos -= 2; } else { // CRC crcCheck = bitstream.readBits(16); } return { objectType, samplingFrequencyIndex, channelConfiguration, frameLength, numberOfAacFrames, crcCheck, startPos, }; }; ===== src/adts/adts-muxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { buildAdtsHeaderTemplate, parseAacAudioSpecificConfig, writeAdtsFrameLength } from '../../shared/aac-misc'; import { Bitstream } from '../../shared/bitstream'; import { validateAudioChunkMetadata } from '../codec'; import { Id3V2Writer } from '../id3'; import { metadataTagsAreEmpty } from '../metadata'; import { assert, toUint8Array } from '../misc'; import { Muxer } from '../muxer'; import { Output, OutputAudioTrack } from '../output'; import { AdtsOutputFormat } from '../output-format'; import { EncodedPacket } from '../packet'; import { Writer } from '../writer'; export class AdtsMuxer extends Muxer { private format: AdtsOutputFormat; private writer!: Writer; private header: Uint8Array | null = null; private headerBitstream: Bitstream | null = null; private inputIsAdts: boolean | null = null; constructor(output: Output, format: AdtsOutputFormat) { super(output); this.format = format; } async start() { const release = await this.mutex.acquire(); this.writer = await this.output._getRootWriter(true); if (!metadataTagsAreEmpty(this.output._metadataTags)) { const id3Writer = new Id3V2Writer(this.writer); id3Writer.writeId3V2Tag(this.output._metadataTags); } release(); } async getMimeType() { return 'audio/aac'; } async addEncodedVideoPacket() { throw new Error('ADTS does not support video.'); } async addEncodedAudioPacket( track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata, ) { const release = await this.mutex.acquire(); try { this.validateTimestamp(track, packet.timestamp, packet.type === 'key'); // First packet - determine input format from metadata if (this.inputIsAdts === null) { validateAudioChunkMetadata(meta); const description = meta?.decoderConfig?.description; // Follows from the Mediabunny Codec Registry: this.inputIsAdts = !description; if (!this.inputIsAdts) { const config = parseAacAudioSpecificConfig(toUint8Array(description!)); const template = buildAdtsHeaderTemplate(config); this.header = template.header; this.headerBitstream = template.bitstream; } } if (this.inputIsAdts) { // Packets are already ADTS frames, write them directly const startPos = this.writer.getPos(); this.writer.write(packet.data); if (this.format._options.onFrame) { this.format._options.onFrame(packet.data, startPos); } } else { assert(this.header); // Packets are raw AAC, we gotta turn it into ADTS const frameLength = packet.data.byteLength + this.header.byteLength; writeAdtsFrameLength(this.headerBitstream!, frameLength); const startPos = this.writer.getPos(); this.writer.write(this.header); this.writer.write(packet.data); if (this.format._options.onFrame) { const frameBytes = new Uint8Array(frameLength); frameBytes.set(this.header, 0); frameBytes.set(packet.data, this.header.byteLength); this.format._options.onFrame(frameBytes, startPos); } } await this.writer.flush(); } finally { release(); } } async addSubtitleCue() { throw new Error('ADTS does not support subtitles.'); } async finalize() { const release = await this.mutex.acquire(); // Required so that finalize() can't resolve before other calls release(); } } ===== src/adts/adts-demuxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { aacChannelMap, aacFrequencyTable } from '../../shared/aac-misc'; import { AudioCodec } from '../codec'; import { Demuxer } from '../demuxer'; import { ID3_V2_HEADER_SIZE, parseId3V2Tag, readId3V2Header, } from '../id3'; import { Input } from '../input'; import { InputAudioTrackBacking } from '../input-track'; import { PacketRetrievalOptions } from '../media-sink'; import { DEFAULT_TRACK_DISPOSITION, MetadataTags } from '../metadata'; import { assert, AsyncMutex, binarySearchExact, binarySearchLessOrEqual, UNDETERMINED_LANGUAGE, } from '../misc'; import { EncodedPacket, PLACEHOLDER_DATA } from '../packet'; import { readBytes, Reader } from '../reader'; import { AdtsFrameHeader, MIN_ADTS_FRAME_HEADER_SIZE, MAX_ADTS_FRAME_HEADER_SIZE, readAdtsFrameHeader, } from './adts-reader'; export const SAMPLES_PER_AAC_FRAME = 1024; type Sample = { timestamp: number; duration: number; dataStart: number; dataSize: number; }; export class AdtsDemuxer extends Demuxer { reader: Reader; metadataPromise: Promise | null = null; firstFrameHeader: AdtsFrameHeader | null = null; loadedSamples: Sample[] = []; metadataTags: MetadataTags | null = null; trackBackings: AdtsAudioTrackBacking[] = []; readingMutex = new AsyncMutex(); lastSampleLoaded = false; lastLoadedPos = 0; nextTimestampInSamples = 0; constructor(input: Input) { super(input); this.reader = input._reader; } async readMetadata() { return this.metadataPromise ??= (async () => { // Keep loading until we find the first frame header while (!this.firstFrameHeader && !this.lastSampleLoaded) { await this.advanceReader(); } // There has to be a frame if this demuxer got selected assert(this.firstFrameHeader); // Create the single audio track this.trackBackings = [new AdtsAudioTrackBacking(this)]; })(); } async advanceReader() { if (this.lastLoadedPos === 0) { // Skip all ID3v2 tags at the start of the file while (true) { let slice = this.reader.requestSlice(this.lastLoadedPos, ID3_V2_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) { this.lastSampleLoaded = true; return; } const id3V2Header = readId3V2Header(slice); if (!id3V2Header) { break; } this.lastLoadedPos = slice.filePos + id3V2Header.size; } } let slice = this.reader.requestSliceRange( this.lastLoadedPos, MIN_ADTS_FRAME_HEADER_SIZE, MAX_ADTS_FRAME_HEADER_SIZE, ); if (slice instanceof Promise) slice = await slice; if (!slice) { this.lastSampleLoaded = true; return; } const header = readAdtsFrameHeader(slice); if (!header) { this.lastSampleLoaded = true; return; } if (this.reader.fileSize !== null && header.startPos + header.frameLength > this.reader.fileSize) { // Frame doesn't fit in the rest of the file this.lastSampleLoaded = true; return; } if (!this.firstFrameHeader) { this.firstFrameHeader = header; } const sampleRate = aacFrequencyTable[header.samplingFrequencyIndex]; assert(sampleRate !== undefined); const sampleDuration = SAMPLES_PER_AAC_FRAME / sampleRate; const sample: Sample = { timestamp: this.nextTimestampInSamples / sampleRate, duration: sampleDuration, dataStart: header.startPos, dataSize: header.frameLength, }; this.loadedSamples.push(sample); this.nextTimestampInSamples += SAMPLES_PER_AAC_FRAME; this.lastLoadedPos = header.startPos + header.frameLength; } async getMimeType() { return 'audio/aac'; } async getTrackBackings() { await this.readMetadata(); return this.trackBackings; } async getMetadataTags() { const release = await this.readingMutex.acquire(); try { await this.readMetadata(); if (this.metadataTags) { return this.metadataTags; } this.metadataTags = {}; let currentPos = 0; while (true) { let headerSlice = this.reader.requestSlice(currentPos, ID3_V2_HEADER_SIZE); if (headerSlice instanceof Promise) headerSlice = await headerSlice; if (!headerSlice) break; const id3V2Header = readId3V2Header(headerSlice); if (!id3V2Header) { break; } let contentSlice = this.reader.requestSlice(headerSlice.filePos, id3V2Header.size); if (contentSlice instanceof Promise) contentSlice = await contentSlice; if (!contentSlice) break; parseId3V2Tag(contentSlice, id3V2Header, this.metadataTags); currentPos = headerSlice.filePos + id3V2Header.size; } return this.metadataTags; } finally { release(); } } } class AdtsAudioTrackBacking implements InputAudioTrackBacking { constructor(public demuxer: AdtsDemuxer) {} getType() { return 'audio' as const; } getId() { return 1; } getNumber() { return 1; } getTimeResolution() { const sampleRate = this.getSampleRate(); return sampleRate / SAMPLES_PER_AAC_FRAME; } isRelativeToUnixEpoch() { return false; } getUnixTimeForTimestamp() { return null; } getPairingMask() { return 1n; } getBitrate() { return null; } getAverageBitrate() { return null; } async getDurationFromMetadata() { return null; // No way } async getLiveRefreshInterval() { return null; } getName() { return null; } getLanguageCode() { return UNDETERMINED_LANGUAGE; } getCodec(): AudioCodec { return 'aac'; } getInternalCodecId() { assert(this.demuxer.firstFrameHeader); return this.demuxer.firstFrameHeader.objectType; } getNumberOfChannels() { assert(this.demuxer.firstFrameHeader); const numberOfChannels = aacChannelMap[this.demuxer.firstFrameHeader.channelConfiguration]; assert(numberOfChannels !== undefined); return numberOfChannels; } getSampleRate() { assert(this.demuxer.firstFrameHeader); const sampleRate = aacFrequencyTable[this.demuxer.firstFrameHeader.samplingFrequencyIndex]; assert(sampleRate !== undefined); return sampleRate; } getDisposition() { return { ...DEFAULT_TRACK_DISPOSITION, }; } async getDecoderConfig(): Promise { assert(this.demuxer.firstFrameHeader); return { codec: `mp4a.40.${this.demuxer.firstFrameHeader.objectType}`, numberOfChannels: this.getNumberOfChannels(), sampleRate: this.getSampleRate(), }; } async getPacketAtIndex(sampleIndex: number, options: PacketRetrievalOptions) { if (sampleIndex === -1) { return null; } const rawSample = this.demuxer.loadedSamples[sampleIndex]; if (!rawSample) { return null; } let data: Uint8Array; if (options.metadataOnly) { data = PLACEHOLDER_DATA; } else { let slice = this.demuxer.reader.requestSlice(rawSample.dataStart, rawSample.dataSize); if (slice instanceof Promise) slice = await slice; if (!slice) { return null; // Data didn't fit into the rest of the file } data = readBytes(slice, rawSample.dataSize); } return new EncodedPacket( data, 'key', rawSample.timestamp, rawSample.duration, sampleIndex, rawSample.dataSize, ); } getFirstPacket(options: PacketRetrievalOptions) { return this.getPacketAtIndex(0, options); } async getNextPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { const release = await this.demuxer.readingMutex.acquire(); try { const sampleIndex = binarySearchExact( this.demuxer.loadedSamples, packet.timestamp, x => x.timestamp, ); if (sampleIndex === -1) { throw new Error('Packet was not created from this track.'); } const nextIndex = sampleIndex + 1; // Ensure the next sample exists while ( nextIndex >= this.demuxer.loadedSamples.length && !this.demuxer.lastSampleLoaded ) { await this.demuxer.advanceReader(); } return this.getPacketAtIndex(nextIndex, options); } finally { release(); } } async getPacket(timestamp: number, options: PacketRetrievalOptions) { const release = await this.demuxer.readingMutex.acquire(); try { while (true) { const index = binarySearchLessOrEqual( this.demuxer.loadedSamples, timestamp, x => x.timestamp, ); if (index === -1 && this.demuxer.loadedSamples.length > 0) { // We're before the first sample return null; } if (this.demuxer.lastSampleLoaded) { // All data is loaded, return what we found return this.getPacketAtIndex(index, options); } if (index >= 0 && index + 1 < this.demuxer.loadedSamples.length) { // The next packet also exists, we're done return this.getPacketAtIndex(index, options); } // Otherwise, keep loading data await this.demuxer.advanceReader(); } } finally { release(); } } getKeyPacket(timestamp: number, options: PacketRetrievalOptions) { return this.getPacket(timestamp, options); } getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { return this.getNextPacket(packet, options); } } ===== src/codec.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { parseAacAudioSpecificConfig } from '../shared/aac-misc'; import { Av1CodecInfo, AvcDecoderConfigurationRecord, HevcDecoderConfigurationRecord, Vp9CodecInfo, } from './codec-data'; import { COLOR_PRIMARIES_MAP, MATRIX_COEFFICIENTS_MAP, TRANSFER_CHARACTERISTICS_MAP, assert, base64ToBytes, bytesToHexString, isAllowSharedBufferSource, last, reverseBitsU32, toDataView, } from './misc'; import { SubtitleMetadata } from './subtitles'; /** * List of known video codecs, ordered by encoding preference. * @group Codecs * @public */ export const VIDEO_CODECS = [ 'avc', 'hevc', 'vp9', 'av1', 'vp8', ] as const; /** * List of known PCM (uncompressed) audio codecs, ordered by encoding preference. * @group Codecs * @public */ export const PCM_AUDIO_CODECS = [ 'pcm-s16', // We don't prefix 'le' so we're compatible with the WebCodecs-registered PCM codec strings 'pcm-s16be', 'pcm-s24', 'pcm-s24be', 'pcm-s32', 'pcm-s32be', 'pcm-f32', 'pcm-f32be', 'pcm-f64', 'pcm-f64be', 'pcm-u8', 'pcm-s8', 'ulaw', 'alaw', ] as const; /** * List of known compressed audio codecs, ordered by encoding preference. * @group Codecs * @public */ export const NON_PCM_AUDIO_CODECS = [ 'aac', 'opus', 'mp3', 'vorbis', 'flac', 'ac3', 'eac3', ] as const; /** * List of known audio codecs, ordered by encoding preference. * @group Codecs * @public */ export const AUDIO_CODECS = [ ...NON_PCM_AUDIO_CODECS, ...PCM_AUDIO_CODECS, ] as const; /** * List of known subtitle codecs, ordered by encoding preference. * @group Codecs * @public */ export const SUBTITLE_CODECS = [ 'webvtt', ] as const; // TODO add the rest /** * Union type of known video codecs. * @group Codecs * @public */ export type VideoCodec = typeof VIDEO_CODECS[number]; /** * Union type of known audio codecs. * @group Codecs * @public */ export type AudioCodec = typeof AUDIO_CODECS[number]; export type PcmAudioCodec = typeof PCM_AUDIO_CODECS[number]; /** * Union type of known subtitle codecs. * @group Codecs * @public */ export type SubtitleCodec = typeof SUBTITLE_CODECS[number]; /** * Union type of known media codecs. * @group Codecs * @public */ export type MediaCodec = VideoCodec | AudioCodec | SubtitleCodec; // https://en.wikipedia.org/wiki/Advanced_Video_Coding export const AVC_LEVEL_TABLE = [ { maxMacroblocks: 99, maxBitrate: 64000, maxDpbMbs: 396, level: 0x0A }, // Level 1 { maxMacroblocks: 396, maxBitrate: 192000, maxDpbMbs: 900, level: 0x0B }, // Level 1.1 { maxMacroblocks: 396, maxBitrate: 384000, maxDpbMbs: 2376, level: 0x0C }, // Level 1.2 { maxMacroblocks: 396, maxBitrate: 768000, maxDpbMbs: 2376, level: 0x0D }, // Level 1.3 { maxMacroblocks: 396, maxBitrate: 2000000, maxDpbMbs: 2376, level: 0x14 }, // Level 2 { maxMacroblocks: 792, maxBitrate: 4000000, maxDpbMbs: 4752, level: 0x15 }, // Level 2.1 { maxMacroblocks: 1620, maxBitrate: 4000000, maxDpbMbs: 8100, level: 0x16 }, // Level 2.2 { maxMacroblocks: 1620, maxBitrate: 10000000, maxDpbMbs: 8100, level: 0x1E }, // Level 3 { maxMacroblocks: 3600, maxBitrate: 14000000, maxDpbMbs: 18000, level: 0x1F }, // Level 3.1 { maxMacroblocks: 5120, maxBitrate: 20000000, maxDpbMbs: 20480, level: 0x20 }, // Level 3.2 { maxMacroblocks: 8192, maxBitrate: 20000000, maxDpbMbs: 32768, level: 0x28 }, // Level 4 { maxMacroblocks: 8192, maxBitrate: 50000000, maxDpbMbs: 32768, level: 0x29 }, // Level 4.1 { maxMacroblocks: 8704, maxBitrate: 50000000, maxDpbMbs: 34816, level: 0x2A }, // Level 4.2 { maxMacroblocks: 22080, maxBitrate: 135000000, maxDpbMbs: 110400, level: 0x32 }, // Level 5 { maxMacroblocks: 36864, maxBitrate: 240000000, maxDpbMbs: 184320, level: 0x33 }, // Level 5.1 { maxMacroblocks: 36864, maxBitrate: 240000000, maxDpbMbs: 184320, level: 0x34 }, // Level 5.2 { maxMacroblocks: 139264, maxBitrate: 240000000, maxDpbMbs: 696320, level: 0x3C }, // Level 6 { maxMacroblocks: 139264, maxBitrate: 480000000, maxDpbMbs: 696320, level: 0x3D }, // Level 6.1 { maxMacroblocks: 139264, maxBitrate: 800000000, maxDpbMbs: 696320, level: 0x3E }, // Level 6.2 ]; // https://en.wikipedia.org/wiki/High_Efficiency_Video_Coding const HEVC_LEVEL_TABLE = [ { maxPictureSize: 36864, maxBitrate: 128000, tier: 'L', level: 30 }, // Level 1 (Low Tier) { maxPictureSize: 122880, maxBitrate: 1500000, tier: 'L', level: 60 }, // Level 2 (Low Tier) { maxPictureSize: 245760, maxBitrate: 3000000, tier: 'L', level: 63 }, // Level 2.1 (Low Tier) { maxPictureSize: 552960, maxBitrate: 6000000, tier: 'L', level: 90 }, // Level 3 (Low Tier) { maxPictureSize: 983040, maxBitrate: 10000000, tier: 'L', level: 93 }, // Level 3.1 (Low Tier) { maxPictureSize: 2228224, maxBitrate: 12000000, tier: 'L', level: 120 }, // Level 4 (Low Tier) { maxPictureSize: 2228224, maxBitrate: 30000000, tier: 'H', level: 120 }, // Level 4 (High Tier) { maxPictureSize: 2228224, maxBitrate: 20000000, tier: 'L', level: 123 }, // Level 4.1 (Low Tier) { maxPictureSize: 2228224, maxBitrate: 50000000, tier: 'H', level: 123 }, // Level 4.1 (High Tier) { maxPictureSize: 8912896, maxBitrate: 25000000, tier: 'L', level: 150 }, // Level 5 (Low Tier) { maxPictureSize: 8912896, maxBitrate: 100000000, tier: 'H', level: 150 }, // Level 5 (High Tier) { maxPictureSize: 8912896, maxBitrate: 40000000, tier: 'L', level: 153 }, // Level 5.1 (Low Tier) { maxPictureSize: 8912896, maxBitrate: 160000000, tier: 'H', level: 153 }, // Level 5.1 (High Tier) { maxPictureSize: 8912896, maxBitrate: 60000000, tier: 'L', level: 156 }, // Level 5.2 (Low Tier) { maxPictureSize: 8912896, maxBitrate: 240000000, tier: 'H', level: 156 }, // Level 5.2 (High Tier) { maxPictureSize: 35651584, maxBitrate: 60000000, tier: 'L', level: 180 }, // Level 6 (Low Tier) { maxPictureSize: 35651584, maxBitrate: 240000000, tier: 'H', level: 180 }, // Level 6 (High Tier) { maxPictureSize: 35651584, maxBitrate: 120000000, tier: 'L', level: 183 }, // Level 6.1 (Low Tier) { maxPictureSize: 35651584, maxBitrate: 480000000, tier: 'H', level: 183 }, // Level 6.1 (High Tier) { maxPictureSize: 35651584, maxBitrate: 240000000, tier: 'L', level: 186 }, // Level 6.2 (Low Tier) { maxPictureSize: 35651584, maxBitrate: 800000000, tier: 'H', level: 186 }, // Level 6.2 (High Tier) ]; // https://en.wikipedia.org/wiki/VP9 export const VP9_LEVEL_TABLE = [ { maxPictureSize: 36864, maxBitrate: 200000, level: 10 }, // Level 1 { maxPictureSize: 73728, maxBitrate: 800000, level: 11 }, // Level 1.1 { maxPictureSize: 122880, maxBitrate: 1800000, level: 20 }, // Level 2 { maxPictureSize: 245760, maxBitrate: 3600000, level: 21 }, // Level 2.1 { maxPictureSize: 552960, maxBitrate: 7200000, level: 30 }, // Level 3 { maxPictureSize: 983040, maxBitrate: 12000000, level: 31 }, // Level 3.1 { maxPictureSize: 2228224, maxBitrate: 18000000, level: 40 }, // Level 4 { maxPictureSize: 2228224, maxBitrate: 30000000, level: 41 }, // Level 4.1 { maxPictureSize: 8912896, maxBitrate: 60000000, level: 50 }, // Level 5 { maxPictureSize: 8912896, maxBitrate: 120000000, level: 51 }, // Level 5.1 { maxPictureSize: 8912896, maxBitrate: 180000000, level: 52 }, // Level 5.2 { maxPictureSize: 35651584, maxBitrate: 180000000, level: 60 }, // Level 6 { maxPictureSize: 35651584, maxBitrate: 240000000, level: 61 }, // Level 6.1 { maxPictureSize: 35651584, maxBitrate: 480000000, level: 62 }, // Level 6.2 ]; // https://en.wikipedia.org/wiki/AV1 const AV1_LEVEL_TABLE = [ { maxPictureSize: 147456, maxBitrate: 1500000, tier: 'M', level: 0 }, // Level 2.0 (Main Tier) { maxPictureSize: 278784, maxBitrate: 3000000, tier: 'M', level: 1 }, // Level 2.1 (Main Tier) { maxPictureSize: 665856, maxBitrate: 6000000, tier: 'M', level: 4 }, // Level 3.0 (Main Tier) { maxPictureSize: 1065024, maxBitrate: 10000000, tier: 'M', level: 5 }, // Level 3.1 (Main Tier) { maxPictureSize: 2359296, maxBitrate: 12000000, tier: 'M', level: 8 }, // Level 4.0 (Main Tier) { maxPictureSize: 2359296, maxBitrate: 30000000, tier: 'H', level: 8 }, // Level 4.0 (High Tier) { maxPictureSize: 2359296, maxBitrate: 20000000, tier: 'M', level: 9 }, // Level 4.1 (Main Tier) { maxPictureSize: 2359296, maxBitrate: 50000000, tier: 'H', level: 9 }, // Level 4.1 (High Tier) { maxPictureSize: 8912896, maxBitrate: 30000000, tier: 'M', level: 12 }, // Level 5.0 (Main Tier) { maxPictureSize: 8912896, maxBitrate: 100000000, tier: 'H', level: 12 }, // Level 5.0 (High Tier) { maxPictureSize: 8912896, maxBitrate: 40000000, tier: 'M', level: 13 }, // Level 5.1 (Main Tier) { maxPictureSize: 8912896, maxBitrate: 160000000, tier: 'H', level: 13 }, // Level 5.1 (High Tier) { maxPictureSize: 8912896, maxBitrate: 60000000, tier: 'M', level: 14 }, // Level 5.2 (Main Tier) { maxPictureSize: 8912896, maxBitrate: 240000000, tier: 'H', level: 14 }, // Level 5.2 (High Tier) { maxPictureSize: 35651584, maxBitrate: 60000000, tier: 'M', level: 15 }, // Level 5.3 (Main Tier) { maxPictureSize: 35651584, maxBitrate: 240000000, tier: 'H', level: 15 }, // Level 5.3 (High Tier) { maxPictureSize: 35651584, maxBitrate: 60000000, tier: 'M', level: 16 }, // Level 6.0 (Main Tier) { maxPictureSize: 35651584, maxBitrate: 240000000, tier: 'H', level: 16 }, // Level 6.0 (High Tier) { maxPictureSize: 35651584, maxBitrate: 100000000, tier: 'M', level: 17 }, // Level 6.1 (Main Tier) { maxPictureSize: 35651584, maxBitrate: 480000000, tier: 'H', level: 17 }, // Level 6.1 (High Tier) { maxPictureSize: 35651584, maxBitrate: 160000000, tier: 'M', level: 18 }, // Level 6.2 (Main Tier) { maxPictureSize: 35651584, maxBitrate: 800000000, tier: 'H', level: 18 }, // Level 6.2 (High Tier) { maxPictureSize: 35651584, maxBitrate: 160000000, tier: 'M', level: 19 }, // Level 6.3 (Main Tier) { maxPictureSize: 35651584, maxBitrate: 800000000, tier: 'H', level: 19 }, // Level 6.3 (High Tier) ]; const VP9_DEFAULT_SUFFIX = '.01.01.01.01.00'; const AV1_DEFAULT_SUFFIX = '.0.110.01.01.01.0'; export const buildVideoCodecString = (codec: VideoCodec, width: number, height: number, bitrate: number) => { if (codec === 'avc') { const profileIndication = 0x64; // High Profile const totalMacroblocks = Math.ceil(width / 16) * Math.ceil(height / 16); // Determine the level based on the table const levelInfo = AVC_LEVEL_TABLE.find( level => totalMacroblocks <= level.maxMacroblocks && bitrate <= level.maxBitrate, ) ?? last(AVC_LEVEL_TABLE)!; const levelIndication = levelInfo ? levelInfo.level : 0; const hexProfileIndication = profileIndication.toString(16).padStart(2, '0'); const hexProfileCompatibility = '00'; const hexLevelIndication = levelIndication.toString(16).padStart(2, '0'); return `avc1.${hexProfileIndication}${hexProfileCompatibility}${hexLevelIndication}`; } else if (codec === 'hevc') { const profilePrefix = ''; // Profile space 0 const profileIdc = 1; // Main Profile const compatibilityFlags = '6'; // Taken from the example in ISO 14496-15 const pictureSize = width * height; const levelInfo = HEVC_LEVEL_TABLE.find( level => pictureSize <= level.maxPictureSize && bitrate <= level.maxBitrate, ) ?? last(HEVC_LEVEL_TABLE)!; const constraintFlags = 'B0'; // Progressive source flag return 'hev1.' + `${profilePrefix}${profileIdc}.` + `${compatibilityFlags}.` + `${levelInfo.tier}${levelInfo.level}.` + `${constraintFlags}`; } else if (codec === 'vp8') { return 'vp8'; // Easy, this one } else if (codec === 'vp9') { const profile = '00'; // Profile 0 const pictureSize = width * height; const levelInfo = VP9_LEVEL_TABLE.find( level => pictureSize <= level.maxPictureSize && bitrate <= level.maxBitrate, ) ?? last(VP9_LEVEL_TABLE)!; const bitDepth = '08'; // 8-bit return `vp09.${profile}.${levelInfo.level.toString().padStart(2, '0')}.${bitDepth}`; } else if (codec === 'av1') { const profile = 0; // Main Profile, single digit const pictureSize = width * height; const levelInfo = AV1_LEVEL_TABLE.find( level => pictureSize <= level.maxPictureSize && bitrate <= level.maxBitrate, ) ?? last(AV1_LEVEL_TABLE)!; const level = levelInfo.level.toString().padStart(2, '0'); const bitDepth = '08'; // 8-bit return `av01.${profile}.${level}${levelInfo.tier}.${bitDepth}`; } // eslint-disable-next-line @typescript-eslint/restrict-template-expressions throw new TypeError(`Unhandled codec '${codec}'.`); }; export const generateVp9CodecConfigurationFromCodecString = (codecString: string) => { // Reference: https://www.webmproject.org/docs/container/#vp9-codec-feature-metadata-codecprivate const parts = codecString.split('.'); // We can derive the required values from the codec string const profile = Number(parts[1]); const level = Number(parts[2]); const bitDepth = Number(parts[3]); const chromaSubsampling = parts[4] ? Number(parts[4]) : 1; return [ 1, 1, profile, 2, 1, level, 3, 1, bitDepth, 4, 1, chromaSubsampling, ]; }; export const generateAv1CodecConfigurationFromCodecString = (codecString: string) => { // Reference: https://aomediacodec.github.io/av1-isobmff/ const parts = codecString.split('.'); // We can derive the required values from the codec string const marker = 1; const version = 1; const firstByte = (marker << 7) + version; const profile = Number(parts[1]); const levelAndTier = parts[2]!; const level = Number(levelAndTier.slice(0, -1)); const secondByte = (profile << 5) + level; const tier = levelAndTier.slice(-1) === 'H' ? 1 : 0; const bitDepth = Number(parts[3]); const highBitDepth = bitDepth === 8 ? 0 : 1; const twelveBit = 0; const monochrome = parts[4] ? Number(parts[4]) : 0; const chromaSubsamplingX = parts[5] ? Number(parts[5][0]) : 1; const chromaSubsamplingY = parts[5] ? Number(parts[5][1]) : 1; const chromaSamplePosition = parts[5] ? Number(parts[5][2]) : 0; // CSP_UNKNOWN const thirdByte = (tier << 7) + (highBitDepth << 6) + (twelveBit << 5) + (monochrome << 4) + (chromaSubsamplingX << 3) + (chromaSubsamplingY << 2) + chromaSamplePosition; const initialPresentationDelayPresent = 0; // Should be fine const fourthByte = initialPresentationDelayPresent; return [firstByte, secondByte, thirdByte, fourthByte]; }; export const extractVideoCodecString = (trackInfo: { width: number; height: number; codec: VideoCodec | null; codecDescription: Uint8Array | null; colorSpace: VideoColorSpaceInit | null; avcType: 1 | 3 | null; avcCodecInfo: AvcDecoderConfigurationRecord | null; hevcCodecInfo: HevcDecoderConfigurationRecord | null; vp9CodecInfo: Vp9CodecInfo | null; av1CodecInfo: Av1CodecInfo | null; }) => { const { codec, codecDescription, colorSpace, avcCodecInfo, hevcCodecInfo, vp9CodecInfo, av1CodecInfo } = trackInfo; if (codec === 'avc') { assert(trackInfo.avcType !== null); if (avcCodecInfo) { const bytes = new Uint8Array([ avcCodecInfo.avcProfileIndication, avcCodecInfo.profileCompatibility, avcCodecInfo.avcLevelIndication, ]); return `avc${trackInfo.avcType}.${bytesToHexString(bytes)}`; } if (!codecDescription || codecDescription.byteLength < 4) { throw new TypeError('AVC decoder description is not provided or is not at least 4 bytes long.'); } return `avc${trackInfo.avcType}.${bytesToHexString(codecDescription.subarray(1, 4))}`; } else if (codec === 'hevc') { let generalProfileSpace: number; let generalProfileIdc: number; let compatibilityFlags: number; let generalTierFlag: number; let generalLevelIdc: number; let constraintFlags: number[]; if (hevcCodecInfo) { generalProfileSpace = hevcCodecInfo.generalProfileSpace; generalProfileIdc = hevcCodecInfo.generalProfileIdc; compatibilityFlags = reverseBitsU32(hevcCodecInfo.generalProfileCompatibilityFlags); generalTierFlag = hevcCodecInfo.generalTierFlag; generalLevelIdc = hevcCodecInfo.generalLevelIdc; constraintFlags = [...hevcCodecInfo.generalConstraintIndicatorFlags]; } else { if (!codecDescription || codecDescription.byteLength < 23) { throw new TypeError('HEVC decoder description is not provided or is not at least 23 bytes long.'); } const view = toDataView(codecDescription); const profileByte = view.getUint8(1); generalProfileSpace = (profileByte >> 6) & 0x03; generalProfileIdc = profileByte & 0x1F; compatibilityFlags = reverseBitsU32(view.getUint32(2)); generalTierFlag = (profileByte >> 5) & 0x01; generalLevelIdc = view.getUint8(12); constraintFlags = []; for (let i = 0; i < 6; i++) { constraintFlags.push(view.getUint8(6 + i)); } } let codecString = 'hev1.'; codecString += ['', 'A', 'B', 'C'][generalProfileSpace]! + generalProfileIdc; codecString += '.'; codecString += compatibilityFlags.toString(16).toUpperCase(); codecString += '.'; codecString += generalTierFlag === 0 ? 'L' : 'H'; codecString += generalLevelIdc; while (constraintFlags.length > 0 && constraintFlags[constraintFlags.length - 1] === 0) { constraintFlags.pop(); } if (constraintFlags.length > 0) { codecString += '.'; codecString += constraintFlags.map(x => x.toString(16).toUpperCase()).join('.'); } return codecString; } else if (codec === 'vp8') { return 'vp8'; // Easy, this one } else if (codec === 'vp9') { if (!vp9CodecInfo) { // Calculate level based on dimensions const pictureSize = trackInfo.width * trackInfo.height; let level = last(VP9_LEVEL_TABLE)!.level; // Default to highest level for (const entry of VP9_LEVEL_TABLE) { if (pictureSize <= entry.maxPictureSize) { level = entry.level; break; } } // We don't really know better, so let's return a general-purpose, common codec string and hope for the best return `vp09.00.${level.toString().padStart(2, '0')}.08`; } const profile = vp9CodecInfo.profile.toString().padStart(2, '0'); const level = vp9CodecInfo.level.toString().padStart(2, '0'); const bitDepth = vp9CodecInfo.bitDepth.toString().padStart(2, '0'); const chromaSubsampling = vp9CodecInfo.chromaSubsampling.toString().padStart(2, '0'); const colourPrimaries = vp9CodecInfo.colourPrimaries.toString().padStart(2, '0'); const transferCharacteristics = vp9CodecInfo.transferCharacteristics.toString().padStart(2, '0'); const matrixCoefficients = vp9CodecInfo.matrixCoefficients.toString().padStart(2, '0'); const videoFullRangeFlag = vp9CodecInfo.videoFullRangeFlag.toString().padStart(2, '0'); let string = `vp09.${profile}.${level}.${bitDepth}.${chromaSubsampling}`; string += `.${colourPrimaries}.${transferCharacteristics}.${matrixCoefficients}.${videoFullRangeFlag}`; if (string.endsWith(VP9_DEFAULT_SUFFIX)) { string = string.slice(0, -VP9_DEFAULT_SUFFIX.length); } return string; } else if (codec === 'av1') { if (!av1CodecInfo) { // Calculate level based on dimensions const pictureSize = trackInfo.width * trackInfo.height; let level = last(VP9_LEVEL_TABLE)!.level; // Default to highest level for (const entry of VP9_LEVEL_TABLE) { if (pictureSize <= entry.maxPictureSize) { level = entry.level; break; } } // We don't really know better, so let's return a general-purpose, common codec string and hope for the best return `av01.0.${level.toString().padStart(2, '0')}M.08`; } // https://aomediacodec.github.io/av1-isobmff/#codecsparam const profile = av1CodecInfo.profile; // Single digit const level = av1CodecInfo.level.toString().padStart(2, '0'); const tier = av1CodecInfo.tier ? 'H' : 'M'; const bitDepth = av1CodecInfo.bitDepth.toString().padStart(2, '0'); const monochrome = av1CodecInfo.monochrome ? '1' : '0'; const chromaSubsampling = 100 * av1CodecInfo.chromaSubsamplingX + 10 * av1CodecInfo.chromaSubsamplingY + 1 * ( av1CodecInfo.chromaSubsamplingX && av1CodecInfo.chromaSubsamplingY ? av1CodecInfo.chromaSamplePosition : 0 ); // The defaults are 1 (ITU-R BT.709) const colorPrimaries = colorSpace?.primaries ? COLOR_PRIMARIES_MAP[colorSpace.primaries] : 1; const transferCharacteristics = colorSpace?.transfer ? TRANSFER_CHARACTERISTICS_MAP[colorSpace.transfer] : 1; const matrixCoefficients = colorSpace?.matrix ? MATRIX_COEFFICIENTS_MAP[colorSpace.matrix] : 1; const videoFullRangeFlag = colorSpace?.fullRange ? 1 : 0; let string = `av01.${profile}.${level}${tier}.${bitDepth}`; string += `.${monochrome}.${chromaSubsampling.toString().padStart(3, '0')}`; string += `.${colorPrimaries.toString().padStart(2, '0')}`; string += `.${transferCharacteristics.toString().padStart(2, '0')}`; string += `.${matrixCoefficients.toString().padStart(2, '0')}`; string += `.${videoFullRangeFlag}`; if (string.endsWith(AV1_DEFAULT_SUFFIX)) { string = string.slice(0, -AV1_DEFAULT_SUFFIX.length); } return string; } throw new TypeError(`Unhandled codec '${codec}'.`); }; export const buildAudioCodecString = (codec: AudioCodec, numberOfChannels: number, sampleRate: number) => { if (codec === 'aac') { // If stereo or higher channels and lower sample rate, likely using HE-AAC v2 with PS if (numberOfChannels >= 2 && sampleRate <= 24000) { return 'mp4a.40.29'; // HE-AAC v2 (AAC LC + SBR + PS) } // If sample rate is low, likely using HE-AAC v1 with SBR if (sampleRate <= 24000) { return 'mp4a.40.5'; // HE-AAC v1 (AAC LC + SBR) } // Default to standard AAC-LC for higher sample rates return 'mp4a.40.2'; // AAC-LC } else if (codec === 'mp3') { return 'mp3'; } else if (codec === 'opus') { return 'opus'; } else if (codec === 'vorbis') { return 'vorbis'; } else if (codec === 'flac') { return 'flac'; } else if (codec === 'ac3') { return 'ac-3'; } else if (codec === 'eac3') { return 'ec-3'; } else if ((PCM_AUDIO_CODECS as readonly string[]).includes(codec)) { return codec; } throw new TypeError(`Unhandled codec '${codec}'.`); }; export type AacCodecInfo = { isMpeg2: boolean; objectType: number | null; }; export const extractAudioCodecString = (trackInfo: { codec: AudioCodec | null; codecDescription: Uint8Array | null; aacCodecInfo: AacCodecInfo | null; }) => { const { codec, codecDescription, aacCodecInfo } = trackInfo; if (codec === 'aac') { if (!aacCodecInfo) { throw new TypeError('AAC codec info must be provided.'); } if (aacCodecInfo.isMpeg2) { return 'mp4a.67'; } else { let objectType: number; if (aacCodecInfo.objectType !== null) { objectType = aacCodecInfo.objectType; } else { const audioSpecificConfig = parseAacAudioSpecificConfig(codecDescription); objectType = audioSpecificConfig.objectType; } return `mp4a.40.${objectType}`; } } else if (codec === 'mp3') { return 'mp3'; } else if (codec === 'opus') { return 'opus'; } else if (codec === 'vorbis') { return 'vorbis'; } else if (codec === 'flac') { return 'flac'; } else if (codec === 'ac3') { return 'ac-3'; } else if (codec === 'eac3') { return 'ec-3'; } else if (codec && (PCM_AUDIO_CODECS as readonly string[]).includes(codec)) { return codec; } throw new TypeError(`Unhandled codec '${codec}'.`); }; // eslint-disable-next-line @typescript-eslint/no-unused-vars export const guessDescriptionForVideo = (decoderConfig: VideoDecoderConfig): Uint8Array | undefined => { return undefined; // All codecs allow an undefined description }; export const guessDescriptionForAudio = (decoderConfig: AudioDecoderConfig): Uint8Array | undefined | false => { switch (decoderConfig.codec) { case 'flac': { const referenceDescription = base64ToBytes('ZkxhQ4AAACIQABAAAAYtACWtCsRC8AANRBhVFucAcYu5ASE2m1Dxv8tw'); if (decoderConfig.sampleRate >= (1 << 20) || decoderConfig.numberOfChannels > 8) { return false; } referenceDescription[18] = decoderConfig.sampleRate >>> 12; referenceDescription[19] = decoderConfig.sampleRate >>> 4; referenceDescription[20] = ((decoderConfig.sampleRate & 0x0f) << 4) | ((decoderConfig.numberOfChannels - 1) << 1); return referenceDescription; }; case 'vorbis': { // eslint-disable-next-line @stylistic/max-len const referenceDescription = base64ToBytes('Ah7/AgF2b3JiaXMAAAAAAoC7AAAAAAAAgLUBAAAAAAC4AQN2b3JiaXMNAAAATGF2ZjU4Ljc2LjEwMAgAAAAMAAAAbGFuZ3VhZ2U9dW5kGQAAAGhhbmRsZXJfbmFtZT1Tb3VuZEhhbmRsZXIWAAAAdmVuZG9yX2lkPVswXVswXVswXVswXSAAAABlbmNvZGVyPUxhdmM1OC4xMzQuMTAwIGxpYnZvcmJpcxAAAABtYWpvcl9icmFuZD1pc29tEQAAAG1pbm9yX3ZlcnNpb249NTEyIgAAAGNvbXBhdGlibGVfYnJhbmRzPWlzb21pc28yYXZjMW1wNDEmAAAAREVTQ1JJUFRJT049TWFkZSB3aXRoIFJlbW90aW9uIDQuMC4yNzgBBXZvcmJpcyVCQ1YBAEAAACRzGCpGpXMWhBAaQlAZ4xxCzmvsGUJMEYIcMkxbyyVzkCGkoEKIWyiB0JBVAABAAACHQXgUhIpBCCGEJT1YkoMnPQghhIg5eBSEaUEIIYQQQgghhBBCCCGERTlokoMnQQgdhOMwOAyD5Tj4HIRFOVgQgydB6CCED0K4moOsOQghhCQ1SFCDBjnoHITCLCiKgsQwuBaEBDUojILkMMjUgwtCiJqDSTX4GoRnQXgWhGlBCCGEJEFIkIMGQcgYhEZBWJKDBjm4FITLQagahCo5CB+EIDRkFQCQAACgoiiKoigKEBqyCgDIAAAQQFEUx3EcyZEcybEcCwgNWQUAAAEACAAAoEiKpEiO5EiSJFmSJVmSJVmS5omqLMuyLMuyLMsyEBqyCgBIAABQUQxFcRQHCA1ZBQBkAAAIoDiKpViKpWiK54iOCISGrAIAgAAABAAAEDRDUzxHlETPVFXXtm3btm3btm3btm3btm1blmUZCA1ZBQBAAAAQ0mlmqQaIMAMZBkJDVgEACAAAgBGKMMSA0JBVAABAAACAGEoOogmtOd+c46BZDppKsTkdnEi1eZKbirk555xzzsnmnDHOOeecopxZDJoJrTnnnMSgWQqaCa0555wnsXnQmiqtOeeccc7pYJwRxjnnnCateZCajbU555wFrWmOmkuxOeecSLl5UptLtTnnnHPOOeecc84555zqxekcnBPOOeecqL25lpvQxTnnnE/G6d6cEM4555xzzjnnnHPOOeecIDRkFQAABABAEIaNYdwpCNLnaCBGEWIaMulB9+gwCRqDnELq0ehopJQ6CCWVcVJKJwgNWQUAAAIAQAghhRRSSCGFFFJIIYUUYoghhhhyyimnoIJKKqmooowyyyyzzDLLLLPMOuyssw47DDHEEEMrrcRSU2011lhr7jnnmoO0VlprrbVSSimllFIKQkNWAQAgAAAEQgYZZJBRSCGFFGKIKaeccgoqqIDQkFUAACAAgAAAAABP8hzRER3RER3RER3RER3R8RzPESVREiVREi3TMjXTU0VVdWXXlnVZt31b2IVd933d933d+HVhWJZlWZZlWZZlWZZlWZZlWZYgNGQVAAACAAAghBBCSCGFFFJIKcYYc8w56CSUEAgNWQUAAAIACAAAAHAUR3EcyZEcSbIkS9IkzdIsT/M0TxM9URRF0zRV0RVdUTdtUTZl0zVdUzZdVVZtV5ZtW7Z125dl2/d93/d93/d93/d93/d9XQdCQ1YBABIAADqSIymSIimS4ziOJElAaMgqAEAGAEAAAIriKI7jOJIkSZIlaZJneZaomZrpmZ4qqkBoyCoAABAAQAAAAAAAAIqmeIqpeIqoeI7oiJJomZaoqZoryqbsuq7ruq7ruq7ruq7ruq7ruq7ruq7ruq7ruq7ruq7ruq7ruq4LhIasAgAkAAB0JEdyJEdSJEVSJEdygNCQVQCADACAAAAcwzEkRXIsy9I0T/M0TxM90RM901NFV3SB0JBVAAAgAIAAAAAAAAAMybAUy9EcTRIl1VItVVMt1VJF1VNVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVVN0zRNEwgNWQkAkAEAkBBTLS3GmgmLJGLSaqugYwxS7KWxSCpntbfKMYUYtV4ah5RREHupJGOKQcwtpNApJq3WVEKFFKSYYyoVUg5SIDRkhQAQmgHgcBxAsixAsiwAAAAAAAAAkDQN0DwPsDQPAAAAAAAAACRNAyxPAzTPAwAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAABA0jRA8zxA8zwAAAAAAAAA0DwP8DwR8EQRAAAAAAAAACzPAzTRAzxRBAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAABA0jRA8zxA8zwAAAAAAAAAsDwP8EQR0DwRAAAAAAAAACzPAzxRBDzRAwAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAEAAAEOAAABBgIRQasiIAiBMAcEgSJAmSBM0DSJYFTYOmwTQBkmVB06BpME0AAAAAAAAAAAAAJE2DpkHTIIoASdOgadA0iCIAAAAAAAAAAAAAkqZB06BpEEWApGnQNGgaRBEAAAAAAAAAAAAAzzQhihBFmCbAM02IIkQRpgkAAAAAAAAAAAAAAAAAAAAAAAAAAAAACAAAGHAAAAgwoQwUGrIiAIgTAHA4imUBAIDjOJYFAACO41gWAABYliWKAABgWZooAgAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAIAAAYcAAACDChDBQashIAiAIAcCiKZQHHsSzgOJYFJMmyAJYF0DyApgFEEQAIAAAocAAACLBBU2JxgEJDVgIAUQAABsWxLE0TRZKkaZoniiRJ0zxPFGma53meacLzPM80IYqiaJoQRVE0TZimaaoqME1VFQAAUOAAABBgg6bE4gCFhqwEAEICAByKYlma5nmeJ4qmqZokSdM8TxRF0TRNU1VJkqZ5niiKommapqqyLE3zPFEURdNUVVWFpnmeKIqiaaqq6sLzPE8URdE0VdV14XmeJ4qiaJqq6roQRVE0TdNUTVV1XSCKpmmaqqqqrgtETxRNU1Vd13WB54miaaqqq7ouEE3TVFVVdV1ZBpimaaqq68oyQFVV1XVdV5YBqqqqruu6sgxQVdd1XVmWZQCu67qyLMsCAAAOHAAAAoygk4wqi7DRhAsPQKEhKwKAKAAAwBimFFPKMCYhpBAaxiSEFEImJaXSUqogpFJSKRWEVEoqJaOUUmopVRBSKamUCkIqJZVSAADYgQMA2IGFUGjISgAgDwCAMEYpxhhzTiKkFGPOOScRUoox55yTSjHmnHPOSSkZc8w556SUzjnnnHNSSuacc845KaVzzjnnnJRSSuecc05KKSWEzkEnpZTSOeecEwAAVOAAABBgo8jmBCNBhYasBABSAQAMjmNZmuZ5omialiRpmud5niiapiZJmuZ5nieKqsnzPE8URdE0VZXneZ4oiqJpqirXFUXTNE1VVV2yLIqmaZqq6rowTdNUVdd1XZimaaqq67oubFtVVdV1ZRm2raqq6rqyDFzXdWXZloEsu67s2rIAAPAEBwCgAhtWRzgpGgssNGQlAJABAEAYg5BCCCFlEEIKIYSUUggJAAAYcAAACDChDBQashIASAUAAIyx1lprrbXWQGettdZaa62AzFprrbXWWmuttdZaa6211lJrrbXWWmuttdZaa6211lprrbXWWmuttdZaa6211lprrbXWWmuttdZaa6211lprrbXWWmstpZRSSimllFJKKaWUUkoppZRSSgUA+lU4APg/2LA6wknRWGChISsBgHAAAMAYpRhzDEIppVQIMeacdFRai7FCiDHnJKTUWmzFc85BKCGV1mIsnnMOQikpxVZjUSmEUlJKLbZYi0qho5JSSq3VWIwxqaTWWoutxmKMSSm01FqLMRYjbE2ptdhqq7EYY2sqLbQYY4zFCF9kbC2m2moNxggjWywt1VprMMYY3VuLpbaaizE++NpSLDHWXAAAd4MDAESCjTOsJJ0VjgYXGrISAAgJACAQUooxxhhzzjnnpFKMOeaccw5CCKFUijHGnHMOQgghlIwx5pxzEEIIIYRSSsaccxBCCCGEkFLqnHMQQgghhBBKKZ1zDkIIIYQQQimlgxBCCCGEEEoopaQUQgghhBBCCKmklEIIIYRSQighlZRSCCGEEEIpJaSUUgohhFJCCKGElFJKKYUQQgillJJSSimlEkoJJYQSUikppRRKCCGUUkpKKaVUSgmhhBJKKSWllFJKIYQQSikFAAAcOAAABBhBJxlVFmGjCRcegEJDVgIAZAAAkKKUUiktRYIipRikGEtGFXNQWoqocgxSzalSziDmJJaIMYSUk1Qy5hRCDELqHHVMKQYtlRhCxhik2HJLoXMOAAAAQQCAgJAAAAMEBTMAwOAA4XMQdAIERxsAgCBEZohEw0JweFAJEBFTAUBigkIuAFRYXKRdXECXAS7o4q4DIQQhCEEsDqCABByccMMTb3jCDU7QKSp1IAAAAAAADADwAACQXAAREdHMYWRobHB0eHyAhIiMkAgAAAAAABcAfAAAJCVAREQ0cxgZGhscHR4fICEiIyQBAIAAAgAAAAAggAAEBAQAAAAAAAIAAAAEBA=='); const view = toDataView(referenceDescription); view.setUint8(15, decoderConfig.numberOfChannels); view.setUint32(16, decoderConfig.sampleRate, true); return referenceDescription; }; default: return undefined; // All other codecs allow an undefined description } }; export const OPUS_SAMPLE_RATE = 48_000; const PCM_CODEC_REGEX = /^pcm-([usf])(\d+)(be)?$/; export const parsePcmCodec = (codec: PcmAudioCodec) => { assert(PCM_AUDIO_CODECS.includes(codec)); if (codec === 'ulaw') { return { dataType: 'ulaw' as const, sampleSize: 1 as const, littleEndian: true, silentValue: 255 }; } else if (codec === 'alaw') { return { dataType: 'alaw' as const, sampleSize: 1 as const, littleEndian: true, silentValue: 213 }; } const match = PCM_CODEC_REGEX.exec(codec); assert(match); let dataType: 'unsigned' | 'signed' | 'float' | 'ulaw' | 'alaw'; if (match[1] === 'u') { dataType = 'unsigned'; } else if (match[1] === 's') { dataType = 'signed'; } else { dataType = 'float'; } const sampleSize = (Number(match[2]) / 8) as 1 | 2 | 3 | 4 | 8; const littleEndian = match[3] !== 'be'; const silentValue = codec === 'pcm-u8' ? 2 ** 7 : 0; return { dataType, sampleSize, littleEndian, silentValue }; }; export const inferCodecFromCodecString = (codecString: string): MediaCodec | null => { // Video codecs if (codecString.startsWith('avc1') || codecString.startsWith('avc3')) { return 'avc'; } else if (codecString.startsWith('hev1') || codecString.startsWith('hvc1')) { return 'hevc'; } else if (codecString === 'vp8') { return 'vp8'; } else if (codecString.startsWith('vp09')) { return 'vp9'; } else if (codecString.startsWith('av01')) { return 'av1'; } // Audio codecs if ( codecString === 'mp3' || codecString === 'mp4a.69' || codecString === 'mp4a.6B' || codecString === 'mp4a.6b' || codecString === 'mp4a.40.34' ) { return 'mp3'; } else if (codecString.startsWith('mp4a.40.') || codecString === 'mp4a.67') { return 'aac'; } else if (codecString === 'opus') { return 'opus'; } else if (codecString === 'vorbis') { return 'vorbis'; } else if (codecString === 'flac') { return 'flac'; } else if (codecString === 'ac-3' || codecString === 'ac3') { return 'ac3'; } else if (codecString === 'ec-3' || codecString === 'eac3') { return 'eac3'; } else if (codecString === 'ulaw') { return 'ulaw'; } else if (codecString === 'alaw') { return 'alaw'; } else if (PCM_CODEC_REGEX.test(codecString)) { return codecString as PcmAudioCodec; } // Subtitle codecs if (codecString === 'webvtt') { return 'webvtt'; } return null; }; export const getVideoEncoderConfigExtension = (codec: VideoCodec) => { if (codec === 'avc') { return { avc: { format: 'avc' as const, // Ensure the format is not Annex B }, }; } else if (codec === 'hevc') { return { hevc: { format: 'hevc' as const, // Ensure the format is not Annex B }, }; } return {}; }; export const getAudioEncoderConfigExtension = (codec: AudioCodec) => { if (codec === 'aac') { return { aac: { format: 'aac' as const, // Ensure the format is not ADTS }, }; } else if (codec === 'opus') { return { opus: { format: 'opus' as const, }, }; } return {}; }; const VALID_VIDEO_CODEC_STRING_PREFIXES = ['avc1', 'avc3', 'hev1', 'hvc1', 'vp8', 'vp09', 'av01']; const AVC_CODEC_STRING_REGEX = /^(avc1|avc3)\.[0-9a-fA-F]{6}$/; const HEVC_CODEC_STRING_REGEX = /^(hev1|hvc1)\.(?:[ABC]?\d+)\.[0-9a-fA-F]{1,8}\.[LH]\d+(?:\.[0-9a-fA-F]{1,2}){0,6}$/; const VP9_CODEC_STRING_REGEX = /^vp09(?:\.\d{2}){3}(?:(?:\.\d{2}){5})?$/; const AV1_CODEC_STRING_REGEX = /^av01\.\d\.\d{2}[MH]\.\d{2}(?:\.\d\.\d{3}\.\d{2}\.\d{2}\.\d{2}\.\d)?$/; export const validateVideoChunkMetadata = (metadata: EncodedVideoChunkMetadata | undefined) => { if (!metadata) { throw new TypeError('Video chunk metadata must be provided.'); } if (typeof metadata !== 'object') { throw new TypeError('Video chunk metadata must be an object.'); } if (!metadata.decoderConfig) { throw new TypeError('Video chunk metadata must include a decoder configuration.'); } if (typeof metadata.decoderConfig !== 'object') { throw new TypeError('Video chunk metadata decoder configuration must be an object.'); } if (typeof metadata.decoderConfig.codec !== 'string') { throw new TypeError('Video chunk metadata decoder configuration must specify a codec string.'); } if (!VALID_VIDEO_CODEC_STRING_PREFIXES.some(prefix => metadata.decoderConfig!.codec.startsWith(prefix))) { throw new TypeError( 'Video chunk metadata decoder configuration codec string must be a valid video codec string as specified in' + ' the Mediabunny Codec Registry.', ); } if (!Number.isInteger(metadata.decoderConfig.codedWidth) || metadata.decoderConfig.codedWidth! <= 0) { throw new TypeError( 'Video chunk metadata decoder configuration must specify a valid codedWidth (positive integer).', ); } if (!Number.isInteger(metadata.decoderConfig.codedHeight) || metadata.decoderConfig.codedHeight! <= 0) { throw new TypeError( 'Video chunk metadata decoder configuration must specify a valid codedHeight (positive integer).', ); } if ( metadata.decoderConfig.displayAspectWidth !== undefined && ( !Number.isInteger(metadata.decoderConfig.displayAspectWidth) || metadata.decoderConfig.displayAspectWidth <= 0 ) ) { throw new TypeError( 'Video chunk metadata decoder configuration displayAspectWidth, when defined, must be a positive integer.', ); } if ( metadata.decoderConfig.displayAspectHeight !== undefined && ( !Number.isInteger(metadata.decoderConfig.displayAspectHeight) || metadata.decoderConfig.displayAspectHeight <= 0 ) ) { throw new TypeError( 'Video chunk metadata decoder configuration displayAspectHeight, when defined, must be a positive integer.', ); } if ( (metadata.decoderConfig.displayAspectWidth !== undefined) !== (metadata.decoderConfig.displayAspectHeight !== undefined) ) { throw new TypeError( 'Video chunk metadata decoder configuration must specify both displayAspectWidth and displayAspectHeight,' + ' or neither.', ); } if (metadata.decoderConfig.description !== undefined) { if (!isAllowSharedBufferSource(metadata.decoderConfig.description)) { throw new TypeError( 'Video chunk metadata decoder configuration description, when defined, must be an ArrayBuffer or an' + ' ArrayBuffer view.', ); } } if (metadata.decoderConfig.colorSpace !== undefined) { const { colorSpace } = metadata.decoderConfig; if (typeof colorSpace !== 'object') { throw new TypeError( 'Video chunk metadata decoder configuration colorSpace, when provided, must be an object.', ); } const primariesValues = Object.keys(COLOR_PRIMARIES_MAP); if (colorSpace.primaries != null && !primariesValues.includes(colorSpace.primaries)) { throw new TypeError( `Video chunk metadata decoder configuration colorSpace primaries, when defined, must be one of` + ` ${primariesValues.join(', ')}.`, ); } const transferValues = Object.keys(TRANSFER_CHARACTERISTICS_MAP); if (colorSpace.transfer != null && !transferValues.includes(colorSpace.transfer)) { throw new TypeError( `Video chunk metadata decoder configuration colorSpace transfer, when defined, must be one of` + ` ${transferValues.join(', ')}.`, ); } const matrixValues = Object.keys(MATRIX_COEFFICIENTS_MAP); if (colorSpace.matrix != null && !matrixValues.includes(colorSpace.matrix)) { throw new TypeError( `Video chunk metadata decoder configuration colorSpace matrix, when defined, must be one of` + ` ${matrixValues.join(', ')}.`, ); } if (colorSpace.fullRange != null && typeof colorSpace.fullRange !== 'boolean') { throw new TypeError( 'Video chunk metadata decoder configuration colorSpace fullRange, when defined, must be a boolean.', ); } } if (metadata.decoderConfig.codec.startsWith('avc1') || metadata.decoderConfig.codec.startsWith('avc3')) { // AVC-specific validation if (!AVC_CODEC_STRING_REGEX.test(metadata.decoderConfig.codec)) { throw new TypeError( 'Video chunk metadata decoder configuration codec string for AVC must be a valid AVC codec string as' + ' specified in Section 3.4 of RFC 6381.', ); } // `description` may or may not be set, depending on if the format is AVCC or Annex B, so don't perform any // validation for it. // https://www.w3.org/TR/webcodecs-avc-codec-registration } else if (metadata.decoderConfig.codec.startsWith('hev1') || metadata.decoderConfig.codec.startsWith('hvc1')) { // HEVC-specific validation if (!HEVC_CODEC_STRING_REGEX.test(metadata.decoderConfig.codec)) { throw new TypeError( 'Video chunk metadata decoder configuration codec string for HEVC must be a valid HEVC codec string as' + ' specified in Section E.3 of ISO 14496-15.', ); } // `description` may or may not be set, depending on if the format is HEVC or Annex B, so don't perform any // validation for it. // https://www.w3.org/TR/webcodecs-hevc-codec-registration } else if (metadata.decoderConfig.codec.startsWith('vp8')) { // VP8-specific validation if (metadata.decoderConfig.codec !== 'vp8') { throw new TypeError('Video chunk metadata decoder configuration codec string for VP8 must be "vp8".'); } } else if (metadata.decoderConfig.codec.startsWith('vp09')) { // VP9-specific validation if (!VP9_CODEC_STRING_REGEX.test(metadata.decoderConfig.codec)) { throw new TypeError( 'Video chunk metadata decoder configuration codec string for VP9 must be a valid VP9 codec string as' + ' specified in Section "Codecs Parameter String" of https://www.webmproject.org/vp9/mp4/.', ); } } else if (metadata.decoderConfig.codec.startsWith('av01')) { // AV1-specific validation if (!AV1_CODEC_STRING_REGEX.test(metadata.decoderConfig.codec)) { throw new TypeError( 'Video chunk metadata decoder configuration codec string for AV1 must be a valid AV1 codec string as' + ' specified in Section "Codecs Parameter String" of https://aomediacodec.github.io/av1-isobmff/.', ); } } }; const VALID_AUDIO_CODEC_STRING_PREFIXES = [ 'mp4a', 'mp3', 'opus', 'vorbis', 'flac', 'ulaw', 'alaw', 'pcm', 'ac-3', 'ec-3', ]; export const validateAudioChunkMetadata = (metadata: EncodedAudioChunkMetadata | undefined) => { if (!metadata) { throw new TypeError('Audio chunk metadata must be provided.'); } if (typeof metadata !== 'object') { throw new TypeError('Audio chunk metadata must be an object.'); } if (!metadata.decoderConfig) { throw new TypeError('Audio chunk metadata must include a decoder configuration.'); } if (typeof metadata.decoderConfig !== 'object') { throw new TypeError('Audio chunk metadata decoder configuration must be an object.'); } if (typeof metadata.decoderConfig.codec !== 'string') { throw new TypeError('Audio chunk metadata decoder configuration must specify a codec string.'); } if (!VALID_AUDIO_CODEC_STRING_PREFIXES.some(prefix => metadata.decoderConfig!.codec.startsWith(prefix))) { throw new TypeError( 'Audio chunk metadata decoder configuration codec string must be a valid audio codec string as specified in' + ' the Mediabunny Codec Registry.', ); } if (!Number.isInteger(metadata.decoderConfig.sampleRate) || metadata.decoderConfig.sampleRate <= 0) { throw new TypeError( 'Audio chunk metadata decoder configuration must specify a valid sampleRate (positive integer).', ); } if (!Number.isInteger(metadata.decoderConfig.numberOfChannels) || metadata.decoderConfig.numberOfChannels <= 0) { throw new TypeError( 'Audio chunk metadata decoder configuration must specify a valid numberOfChannels (positive integer).', ); } if (metadata.decoderConfig.description !== undefined) { if (!isAllowSharedBufferSource(metadata.decoderConfig.description)) { throw new TypeError( 'Audio chunk metadata decoder configuration description, when defined, must be an ArrayBuffer or an' + ' ArrayBuffer view.', ); } } if ( metadata.decoderConfig.codec.startsWith('mp4a') // These three refer to MP3: && metadata.decoderConfig.codec !== 'mp4a.69' && metadata.decoderConfig.codec !== 'mp4a.6B' && metadata.decoderConfig.codec !== 'mp4a.6b' ) { // AAC-specific validation const validStrings = ['mp4a.40.2', 'mp4a.40.02', 'mp4a.40.5', 'mp4a.40.05', 'mp4a.40.29', 'mp4a.67']; if (!validStrings.includes(metadata.decoderConfig.codec)) { throw new TypeError( 'Audio chunk metadata decoder configuration codec string for AAC must be a valid AAC codec string as' + ' specified in https://www.w3.org/TR/webcodecs-aac-codec-registration/.', ); } // `description` may or may not be set, depending on if the format is AAC or ADTS, so don't perform any // validation for it. // https://www.w3.org/TR/webcodecs-aac-codec-registration } else if (metadata.decoderConfig.codec.startsWith('mp3') || metadata.decoderConfig.codec.startsWith('mp4a')) { // MP3-specific validation if ( metadata.decoderConfig.codec !== 'mp3' && metadata.decoderConfig.codec !== 'mp4a.69' && metadata.decoderConfig.codec !== 'mp4a.6B' && metadata.decoderConfig.codec !== 'mp4a.6b' ) { throw new TypeError( 'Audio chunk metadata decoder configuration codec string for MP3 must be "mp3", "mp4a.69" or' + ' "mp4a.6B".', ); } } else if (metadata.decoderConfig.codec.startsWith('opus')) { // Opus-specific validation if (metadata.decoderConfig.codec !== 'opus') { throw new TypeError('Audio chunk metadata decoder configuration codec string for Opus must be "opus".'); } if (metadata.decoderConfig.description && metadata.decoderConfig.description.byteLength < 18) { // Description is optional for Opus per-spec, so we shouldn't enforce it throw new TypeError( 'Audio chunk metadata decoder configuration description, when specified, is expected to be an' + ' Identification Header as specified in Section 5.1 of RFC 7845.', ); } } else if (metadata.decoderConfig.codec.startsWith('vorbis')) { // Vorbis-specific validation if (metadata.decoderConfig.codec !== 'vorbis') { throw new TypeError('Audio chunk metadata decoder configuration codec string for Vorbis must be "vorbis".'); } if (!metadata.decoderConfig.description) { throw new TypeError( 'Audio chunk metadata decoder configuration for Vorbis must include a description, which is expected to' + ' adhere to the format described in https://www.w3.org/TR/webcodecs-vorbis-codec-registration/.', ); } } else if (metadata.decoderConfig.codec.startsWith('flac')) { // FLAC-specific validation if (metadata.decoderConfig.codec !== 'flac') { throw new TypeError('Audio chunk metadata decoder configuration codec string for FLAC must be "flac".'); } const minDescriptionSize = 4 + 4 + 34; // 'fLaC' + metadata block header + STREAMINFO block if (!metadata.decoderConfig.description || metadata.decoderConfig.description.byteLength < minDescriptionSize) { throw new TypeError( 'Audio chunk metadata decoder configuration for FLAC must include a description, which is expected to' + ' adhere to the format described in https://www.w3.org/TR/webcodecs-flac-codec-registration/.', ); } } else if (metadata.decoderConfig.codec.startsWith('ac-3') || metadata.decoderConfig.codec.startsWith('ac3')) { // AC3-specific validation if (metadata.decoderConfig.codec !== 'ac-3') { throw new TypeError('Audio chunk metadata decoder configuration codec string for AC-3 must be "ac-3".'); } } else if (metadata.decoderConfig.codec.startsWith('ec-3') || metadata.decoderConfig.codec.startsWith('eac3')) { // EAC3-specific validation if (metadata.decoderConfig.codec !== 'ec-3') { throw new TypeError('Audio chunk metadata decoder configuration codec string for EC-3 must be "ec-3".'); } } else if ( metadata.decoderConfig.codec.startsWith('pcm') || metadata.decoderConfig.codec.startsWith('ulaw') || metadata.decoderConfig.codec.startsWith('alaw') ) { // PCM-specific validation if (!(PCM_AUDIO_CODECS as readonly string[]).includes(metadata.decoderConfig.codec)) { throw new TypeError( 'Audio chunk metadata decoder configuration codec string for PCM must be one of the supported PCM' + ` codecs (${PCM_AUDIO_CODECS.join(', ')}).`, ); } } }; export const validateSubtitleMetadata = (metadata: SubtitleMetadata | undefined) => { if (!metadata) { throw new TypeError('Subtitle metadata must be provided.'); } if (typeof metadata !== 'object') { throw new TypeError('Subtitle metadata must be an object.'); } if (!metadata.config) { throw new TypeError('Subtitle metadata must include a config object.'); } if (typeof metadata.config !== 'object') { throw new TypeError('Subtitle metadata config must be an object.'); } if (typeof metadata.config.description !== 'string') { throw new TypeError('Subtitle metadata config description must be a string.'); } }; ===== src/isobmff/isobmff-misc.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { bytesToHexString, toDataView, uint8ArraysAreEqual } from '../misc'; export const buildIsobmffMimeType = (info: { isQuickTime: boolean; hasVideo: boolean; hasAudio: boolean; codecStrings: string[]; }) => { const base = info.hasVideo ? 'video/' : info.hasAudio ? 'audio/' : 'application/'; let string = base + (info.isQuickTime ? 'quicktime' : 'mp4'); if (info.codecStrings.length > 0) { const uniqueCodecMimeTypes = [...new Set(info.codecStrings)]; string += `; codecs="${uniqueCodecMimeTypes.join(', ')}"`; } return string; }; /** * Represents a Protection System Specific Header box as used by ISOBMFF Common Encryption. Contains * DRM system-specific data that can be used to obtain a decryption key. * * @group Miscellaneous * @public */ export type PsshBox = { /** The system ID as a 32-bit lowercase hex string. */ systemId: string; /** * The list of key IDs (32-bit lowercase hex strings) this box applies to, or `null` if it applies to all key IDs. */ keyIds: string[] | null; /** The content protection system-specific data. */ data: Uint8Array; }; export const parsePsshBoxContents = (contents: Uint8Array): PsshBox => { const view = toDataView(contents); let pos = 0; const version = view.getUint8(pos); pos += 1; pos += 3; // Flags const systemId = bytesToHexString(contents.subarray(pos, pos + 16)); pos += 16; let keyIds: string[] | null = null; if (version > 0) { const kidCount = view.getUint32(pos); pos += 4; if (kidCount > 0) { keyIds = []; for (let i = 0; i < kidCount; i++) { keyIds.push(bytesToHexString(contents.subarray(pos, pos + 16))); pos += 16; } } } const dataSize = view.getUint32(pos); pos += 4; return { systemId, keyIds, data: contents.slice(pos, pos + dataSize), }; }; export const psshBoxesAreEqual = (a: PsshBox, b: PsshBox) => ( a.systemId === b.systemId && uint8ArraysAreEqual(a.data, b.data) ); ===== src/isobmff/isobmff-muxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Box, free, ftyp, IsobmffBoxWriter, mdat, mfra, moof, moov, sidx, styp, vtta, vttc, vtte, } from './isobmff-boxes'; import { Muxer } from '../muxer'; import { Output, OutputAudioTrack, OutputSubtitleTrack, OutputTrack, OutputVideoTrack, TrackType } from '../output'; import { Writer } from '../writer'; import { BufferTarget } from '../target'; import { assert, computeRationalApproximation, last, promiseWithResolvers, Rational, simplifyRational } from '../misc'; import { IsobmffOutputFormatOptions, IsobmffOutputFormat, MovOutputFormat, CmafOutputFormat } from '../output-format'; import { inlineTimestampRegex, SubtitleConfig, SubtitleCue, SubtitleMetadata } from '../subtitles'; import { aacChannelMap, aacFrequencyTable, buildAacAudioSpecificConfig } from '../../shared/aac-misc'; import { parsePcmCodec, PCM_AUDIO_CODECS, PcmAudioCodec, SubtitleCodec, validateAudioChunkMetadata, validateSubtitleMetadata, validateVideoChunkMetadata, } from '../codec'; import { MAX_ADTS_FRAME_HEADER_SIZE, MIN_ADTS_FRAME_HEADER_SIZE, readAdtsFrameHeader } from '../adts/adts-reader'; import { FileSlice } from '../reader'; import { EncodedPacket, PacketType } from '../packet'; import { concatNalUnitsInLengthPrefixed, extractAvcDecoderConfigurationRecord, extractHevcDecoderConfigurationRecord, iterateNalUnitsInAnnexB, serializeAvcDecoderConfigurationRecord, serializeHevcDecoderConfigurationRecord, } from '../codec-data'; import { buildIsobmffMimeType } from './isobmff-misc'; import { MAX_BOX_HEADER_SIZE, MIN_BOX_HEADER_SIZE } from './isobmff-reader'; export const GLOBAL_TIMESCALE = 57600; // LCM of a bunch of common frame rates (24, 25, 30, 60, 144, ...) const TIMESTAMP_OFFSET = 2_082_844_800; // Seconds between Jan 1 1904 and Jan 1 1970 export type Sample = { timestamp: number; decodeTimestamp: number; duration: number; data: Uint8Array | null; size: number; type: PacketType; timescaleUnitsToNextSample: number; }; type Chunk = { /** The lowest presentation timestamp in this chunk */ startTimestamp: number; samples: Sample[]; offset: number | null; // In the case of a fragmented file, this indicates the position of the moof box pointing to the data in this chunk moofOffset: number | null; }; export type IsobmffTrackData = { muxer: IsobmffMuxer; timescale: number; samples: Sample[]; sampleQueue: Sample[]; // For fragmented files timestampProcessingQueue: Sample[]; timeToSampleTable: { sampleCount: number; sampleDelta: number }[]; compositionTimeOffsetTable: { sampleCount: number; sampleCompositionTimeOffset: number }[]; lastTimescaleUnits: number | null; lastSample: Sample | null; startTimestampOffset: number | null; finalizedChunks: Chunk[]; currentChunk: Chunk | null; compactlyCodedChunkTable: { firstChunk: number; samplesPerChunk: number; }[]; closed: boolean; } & ({ track: OutputVideoTrack; type: 'video'; info: { width: number; height: number; pixelAspectRatio: Rational; decoderConfig: VideoDecoderConfig; /** * The "Annex B transformation" involves converting the raw packet data from Annex B to * "MP4" (length-prefixed) format. * https://stackoverflow.com/questions/24884827 */ requiresAnnexBTransformation: boolean; }; } | { track: OutputAudioTrack; type: 'audio'; info: { numberOfChannels: number; sampleRate: number; decoderConfig: AudioDecoderConfig; /** * The "PCM transformation" is making every sample in the sample table be exactly one PCM audio sample long. * Some players expect this for PCM audio. */ requiresPcmTransformation: boolean; expectedNextPcmPacketTimestamp: number | null; /** * The "ADTS stripping" involves removing the ADTS header from each AAC packet. SOBMFF stores raw AAC data, not * ADTS-wrapped data. */ requiresAdtsStripping: boolean; firstPacket: EncodedPacket; }; } | { track: OutputSubtitleTrack; type: 'subtitle'; info: { config: SubtitleConfig; }; lastCueEndTimestamp: number; cueQueue: SubtitleCue[]; nextSourceId: number; cueToSourceId: WeakMap; }); export type IsobmffVideoTrackData = IsobmffTrackData & { type: 'video' }; export type IsobmffAudioTrackData = IsobmffTrackData & { type: 'audio' }; export type IsobmffSubtitleTrackData = IsobmffTrackData & { type: 'subtitle' }; export type IsobmffMetadata = { name?: string; }; export const getTrackMetadata = (trackData: IsobmffTrackData) => { const metadata: IsobmffMetadata = {}; const track = trackData.track as OutputTrack; if (track.metadata.name !== undefined) { metadata.name = track.metadata.name; } return metadata; }; export const intoTimescale = (timeInSeconds: number, timescale: number, round = true) => { const value = timeInSeconds * timescale; return round ? Math.round(value) : value; }; export class IsobmffMuxer extends Muxer { format: IsobmffOutputFormat; private writer: Writer | null = null; private boxWriter: IsobmffBoxWriter | null = null; private initWriter: Writer | null = null; private initBoxWriter: IsobmffBoxWriter | null = null; private fastStart!: NonNullable; isFragmented!: boolean; isQuickTime: boolean; isCmaf: boolean; private auxTarget = new BufferTarget(); private auxWriter = new Writer(this.auxTarget, false); private auxBoxWriter = new IsobmffBoxWriter(this.auxWriter); private mdat: Box | null = null; private ftypSize: number | null = null; trackDatas: IsobmffTrackData[] = []; private allTracksKnown = promiseWithResolvers(); creationTime = Math.floor(Date.now() / 1000) + TIMESTAMP_OFFSET; private finalizedChunks: Chunk[] = []; private nextFragmentNumber = 1; // Only relevant for fragmented files, to make sure new fragments start with the highest timestamp seen so far private maxWrittenTimestamp = -Infinity; minWrittenTimestamp = Infinity; maxWrittenEndTimestamp = -Infinity; private minimumFragmentDuration: number; private segmentHeaderSize: number | null = null; constructor(output: Output, format: IsobmffOutputFormat) { super(output); this.format = format; this.isQuickTime = format instanceof MovOutputFormat; this.isCmaf = format instanceof CmafOutputFormat; this.minimumFragmentDuration = format._options.minimumFragmentDuration ?? (format instanceof CmafOutputFormat ? Infinity : 1); } async start() { const release = await this.mutex.acquire(); if (!this.isCmaf) { this.writer = await this.output._getRootWriter(target => ( this.format._options.fastStart !== undefined ? this.format._options.fastStart === 'fragmented' : target instanceof BufferTarget // Since if this is the case we'll use 'in-memory' )); this.boxWriter = new IsobmffBoxWriter(this.writer); // If the fastStart option isn't defined, enable in-memory fast start if the target is an ArrayBuffer, as // the memory usage remains identical this.fastStart = this.format._options.fastStart ?? (this.writer.target instanceof BufferTarget ? 'in-memory' : false); this.isFragmented = this.fastStart === 'fragmented'; } else { this.fastStart = 'fragmented'; this.isFragmented = true; } if (this.isCmaf) { if (!this.output._hasInitTarget()) { throw new Error( `CMAF outputs require the initTarget field in OutputOptions to be set; the init segment` + ` will be written to it.`, ); } // Set up the init writer to which we'll write the init segment const initTarget = await this.output._getInitTarget(); const initWriter = new Writer(initTarget, true); initWriter.start(); this.initWriter = initWriter; this.initBoxWriter = new IsobmffBoxWriter(initWriter); } const holdsAvc = this.output._tracks.some(x => x.isVideoTrack() && x.source._codec === 'avc'); // Write the header { const boxWriter = this.initBoxWriter ?? this.boxWriter; assert(boxWriter); if (this.format._options.onFtyp) { boxWriter.writer.startTrackingWrites(); } boxWriter.writeBox(ftyp({ isQuickTime: this.isQuickTime, holdsAvc: holdsAvc, fragmented: this.isFragmented, cmaf: this.isCmaf, })); if (this.format._options.onFtyp) { const { data, start } = boxWriter.writer.stopTrackingWrites(); this.format._options.onFtyp(data, start); } this.ftypSize = boxWriter.writer.getPos(); if (this.isCmaf) { await this.initWriter!.flush(); } } if (this.fastStart === 'in-memory') { // We're write at finalization } else if (this.fastStart === 'reserve') { // Validate that all tracks have set maximumPacketCount for (const track of this.output._tracks) { if (track.metadata.maximumPacketCount === undefined) { throw new Error( 'All tracks must specify maximumPacketCount in their metadata when using' + ' fastStart: \'reserve\'.', ); } } // We'll start writing once we know all tracks } else if (this.isFragmented) { // We write the moov box once we write out the first fragment to make sure we get the decoder configs } else { assert(this.writer); assert(this.boxWriter); if (this.format._options.onMdat) { this.writer.startTrackingWrites(); } this.mdat = mdat(true); // Reserve large size by default, can refine this when finalizing. this.boxWriter.writeBox(this.mdat); } await this.writer?.flush(); release(); } private allTracksAreKnown() { for (const track of this.output._tracks) { if (!track.source._closed && !this.trackDatas.some(x => x.track === track)) { return false; // We haven't seen a sample from this open track yet } } return true; } async getMimeType() { await this.allTracksKnown.promise; const codecStrings = this.trackDatas.map((trackData) => { if (trackData.type === 'video') { return trackData.info.decoderConfig.codec; } else if (trackData.type === 'audio') { return trackData.info.decoderConfig.codec; } else { const map: Record = { webvtt: 'wvtt', }; return map[trackData.track.source._codec]; } }); return buildIsobmffMimeType({ isQuickTime: this.isQuickTime, hasVideo: this.trackDatas.some(x => x.type === 'video'), hasAudio: this.trackDatas.some(x => x.type === 'audio'), codecStrings, }); } private getVideoTrackData(track: OutputVideoTrack, packet: EncodedPacket, meta?: EncodedVideoChunkMetadata) { const existingTrackData = this.trackDatas.find(x => x.track === track); if (existingTrackData) { return existingTrackData as IsobmffVideoTrackData; } validateVideoChunkMetadata(meta); assert(meta); assert(meta.decoderConfig); const decoderConfig = { ...meta.decoderConfig }; assert(decoderConfig.codedWidth !== undefined); assert(decoderConfig.codedHeight !== undefined); let requiresAnnexBTransformation = false; if (track.source._codec === 'avc' && !decoderConfig.description) { // ISOBMFF can only hold AVC in the AVCC format, not in Annex B, but the missing description indicates // Annex B. This means we'll need to do some converterino. const decoderConfigurationRecord = extractAvcDecoderConfigurationRecord(packet.data); if (!decoderConfigurationRecord) { throw new Error( 'Couldn\'t extract an AVCDecoderConfigurationRecord from the AVC packet. Make sure the packets are' + ' in Annex B format (as specified in ITU-T-REC-H.264) when not providing a description, or' + ' provide a description (must be an AVCDecoderConfigurationRecord as specified in ISO 14496-15)' + ' and ensure the packets are in AVCC format.', ); } decoderConfig.description = serializeAvcDecoderConfigurationRecord(decoderConfigurationRecord); requiresAnnexBTransformation = true; } else if (track.source._codec === 'hevc' && !decoderConfig.description) { // ISOBMFF can only hold HEVC in the HEVC format, not in Annex B, but the missing description indicates // Annex B. This means we'll need to do some converterino. const decoderConfigurationRecord = extractHevcDecoderConfigurationRecord(packet.data); if (!decoderConfigurationRecord) { throw new Error( 'Couldn\'t extract an HEVCDecoderConfigurationRecord from the HEVC packet. Make sure the packets' + ' are in Annex B format (as specified in ITU-T-REC-H.265) when not providing a description, or' + ' provide a description (must be an HEVCDecoderConfigurationRecord as specified in ISO 14496-15)' + ' and ensure the packets are in HEVC format.', ); } decoderConfig.description = serializeHevcDecoderConfigurationRecord(decoderConfigurationRecord); requiresAnnexBTransformation = true; } // The frame rate set by the user may not be an integer. Since timescale is an integer, we'll approximate the // frame time (inverse of frame rate) with a rational number, then use that approximation's denominator // as the timescale. const timescale = computeRationalApproximation( 1 / (track.metadata.frameRate ?? GLOBAL_TIMESCALE), 1e6, ).den; const displayAspectWidth = decoderConfig.displayAspectWidth; const displayAspectHeight = decoderConfig.displayAspectHeight; const pixelAspectRatio = displayAspectWidth === undefined || displayAspectHeight === undefined ? { num: 1, den: 1 } : simplifyRational({ num: displayAspectWidth * decoderConfig.codedHeight, den: displayAspectHeight * decoderConfig.codedWidth, }); const newTrackData: IsobmffVideoTrackData = { muxer: this, track, type: 'video', info: { width: decoderConfig.codedWidth, height: decoderConfig.codedHeight, pixelAspectRatio, decoderConfig: decoderConfig, requiresAnnexBTransformation, }, timescale, samples: [], sampleQueue: [], timestampProcessingQueue: [], timeToSampleTable: [], compositionTimeOffsetTable: [], lastTimescaleUnits: null, lastSample: null, startTimestampOffset: null, finalizedChunks: [], currentChunk: null, compactlyCodedChunkTable: [], closed: false, }; this.trackDatas.push(newTrackData); this.trackDatas.sort((a, b) => a.track.id - b.track.id); if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } return newTrackData; } private getAudioTrackData(track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata) { const existingTrackData = this.trackDatas.find(x => x.track === track); if (existingTrackData) { return existingTrackData as IsobmffAudioTrackData; } validateAudioChunkMetadata(meta); assert(meta); assert(meta.decoderConfig); const decoderConfig = { ...meta.decoderConfig }; let requiresAdtsStripping = false; if (track.source._codec === 'aac' && !decoderConfig.description) { // ISOBMFF can only hold AAC in raw format, not ADTS, but the missing description indicates ADTS. // Parse the first packet to extract the AudioSpecificConfig. const adtsFrame = readAdtsFrameHeader(FileSlice.tempFromBytes(packet.data)); if (!adtsFrame) { throw new Error( 'Couldn\'t parse ADTS header from the AAC packet. Make sure the packets are in ADTS format' + ' (as specified in ISO 13818-7) when not providing a description, or provide a description' + ' (must be an AudioSpecificConfig as specified in ISO 14496-3) and ensure the packets' + ' are raw AAC data.', ); } const sampleRate = aacFrequencyTable[adtsFrame.samplingFrequencyIndex]; const numberOfChannels = aacChannelMap[adtsFrame.channelConfiguration]; if (sampleRate === undefined || numberOfChannels === undefined) { throw new Error('Invalid ADTS frame header.'); } decoderConfig.description = buildAacAudioSpecificConfig({ objectType: adtsFrame.objectType, sampleRate, numberOfChannels, }); requiresAdtsStripping = true; } const newTrackData: IsobmffAudioTrackData = { muxer: this, track, type: 'audio', info: { numberOfChannels: meta.decoderConfig.numberOfChannels, sampleRate: meta.decoderConfig.sampleRate, decoderConfig, requiresPcmTransformation: !this.isFragmented && (PCM_AUDIO_CODECS as readonly string[]).includes(track.source._codec), expectedNextPcmPacketTimestamp: null, requiresAdtsStripping, firstPacket: packet, }, timescale: decoderConfig.sampleRate, samples: [], sampleQueue: [], timestampProcessingQueue: [], timeToSampleTable: [], compositionTimeOffsetTable: [], lastTimescaleUnits: null, lastSample: null, startTimestampOffset: null, finalizedChunks: [], currentChunk: null, compactlyCodedChunkTable: [], closed: false, }; this.trackDatas.push(newTrackData); this.trackDatas.sort((a, b) => a.track.id - b.track.id); if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } return newTrackData; } private getSubtitleTrackData(track: OutputSubtitleTrack, meta?: SubtitleMetadata) { const existingTrackData = this.trackDatas.find(x => x.track === track); if (existingTrackData) { return existingTrackData as IsobmffSubtitleTrackData; } validateSubtitleMetadata(meta); assert(meta); assert(meta.config); const newTrackData: IsobmffSubtitleTrackData = { muxer: this, track, type: 'subtitle', info: { config: meta.config, }, timescale: 1000, // Reasonable samples: [], sampleQueue: [], timestampProcessingQueue: [], timeToSampleTable: [], compositionTimeOffsetTable: [], lastTimescaleUnits: null, lastSample: null, startTimestampOffset: null, finalizedChunks: [], currentChunk: null, compactlyCodedChunkTable: [], closed: false, lastCueEndTimestamp: 0, cueQueue: [], nextSourceId: 0, cueToSourceId: new WeakMap(), }; this.trackDatas.push(newTrackData); this.trackDatas.sort((a, b) => a.track.id - b.track.id); if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } return newTrackData; } async addEncodedVideoPacket(track: OutputVideoTrack, packet: EncodedPacket, meta?: EncodedVideoChunkMetadata) { const release = await this.mutex.acquire(); try { const trackData = this.getVideoTrackData(track, packet, meta); let packetData = packet.data; if (trackData.info.requiresAnnexBTransformation) { const nalUnits = [...iterateNalUnitsInAnnexB(packetData)] .map(loc => packetData.subarray(loc.offset, loc.offset + loc.length)); if (nalUnits.length === 0) { // It's not valid Annex B data throw new Error( 'Failed to transform packet data. Make sure all packets are provided in Annex B format, as' + ' specified in ITU-T-REC-H.264 and ITU-T-REC-H.265.', ); } // We don't strip things like SPS or PPS NALUs here, mainly because they can also appear in the middle // of a stream and potentially modify the parameters of it. So, let's just leave them in to be sure. packetData = concatNalUnitsInLengthPrefixed(nalUnits, 4); } this.validateTimestamp( trackData.track, packet.timestamp, packet.type === 'key', ); const internalSample = this.createSampleForTrack( trackData, packetData, packet.timestamp, packet.duration, packet.type, ); await this.registerSample(trackData, internalSample); } finally { release(); } } async addEncodedAudioPacket(track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata) { const release = await this.mutex.acquire(); try { const trackData = this.getAudioTrackData(track, packet, meta); let packetData = packet.data; if (trackData.info.requiresAdtsStripping) { const adtsFrame = readAdtsFrameHeader(FileSlice.tempFromBytes(packetData)); if (!adtsFrame) { throw new Error('Expected ADTS frame, didn\'t get one.'); } const headerLength = adtsFrame.crcCheck === null ? MIN_ADTS_FRAME_HEADER_SIZE : MAX_ADTS_FRAME_HEADER_SIZE; packetData = packetData.subarray(headerLength); } this.validateTimestamp( trackData.track, packet.timestamp, packet.type === 'key', ); let timestamp = packet.timestamp; let duration = packet.duration; if (trackData.info.requiresPcmTransformation) { // Packets may have only approximate timestamp/duration information, but for our PCM logic, we need it // to be precise. So here, we refine the values. const pcmInfo = parsePcmCodec( trackData.info.decoderConfig.codec as PcmAudioCodec, ); const frameSize = pcmInfo.sampleSize * trackData.info.numberOfChannels; // Compute the precise duration duration = packetData.byteLength / frameSize / trackData.info.sampleRate; if (trackData.info.expectedNextPcmPacketTimestamp !== null) { const diff = timestamp - trackData.info.expectedNextPcmPacketTimestamp; if (diff < 0.01) { timestamp = trackData.info.expectedNextPcmPacketTimestamp; } else { const paddedDuration = await this.padWithSilence( trackData, trackData.info.expectedNextPcmPacketTimestamp, diff, ); timestamp = trackData.info.expectedNextPcmPacketTimestamp + paddedDuration; } } trackData.info.expectedNextPcmPacketTimestamp = timestamp + duration; } const internalSample = this.createSampleForTrack( trackData, packetData, timestamp, duration, packet.type, ); await this.registerSample(trackData, internalSample); } finally { release(); } } private async padWithSilence(trackData: IsobmffAudioTrackData, timestamp: number, duration: number) { const deltaInTimescale = intoTimescale(duration, trackData.timescale); duration = deltaInTimescale / trackData.timescale; if (deltaInTimescale > 0) { const { sampleSize, silentValue } = parsePcmCodec( trackData.info.decoderConfig.codec as PcmAudioCodec, ); const samplesNeeded = deltaInTimescale * trackData.info.numberOfChannels; const data = new Uint8Array(sampleSize * samplesNeeded).fill(silentValue); const paddingSample = this.createSampleForTrack( trackData, new Uint8Array(data.buffer), timestamp, duration, 'key', ); await this.registerSample(trackData, paddingSample); } return duration; } async addSubtitleCue(track: OutputSubtitleTrack, cue: SubtitleCue, meta?: SubtitleMetadata) { const release = await this.mutex.acquire(); try { const trackData = this.getSubtitleTrackData(track, meta); this.validateTimestamp(trackData.track, cue.timestamp, true); if (track.source._codec === 'webvtt') { trackData.cueQueue.push(cue); await this.processWebVTTCues(trackData, cue.timestamp); } else { // TODO } } finally { release(); } } private async processWebVTTCues(trackData: IsobmffSubtitleTrackData, until: number) { // WebVTT cues need to undergo special processing as empty sections need to be padded out with samples, and // overlapping samples require special logic. The algorithm produces the format specified in ISO 14496-30. while (trackData.cueQueue.length > 0) { const timestamps = new Set([]); for (const cue of trackData.cueQueue) { assert(cue.timestamp <= until); assert(trackData.lastCueEndTimestamp <= cue.timestamp + cue.duration); timestamps.add(Math.max(cue.timestamp, trackData.lastCueEndTimestamp)); // Start timestamp timestamps.add(cue.timestamp + cue.duration); // End timestamp } const sortedTimestamps = [...timestamps].sort((a, b) => a - b); // These are the timestamps of the next sample we'll create: const sampleStart = sortedTimestamps[0]!; const sampleEnd = sortedTimestamps[1] ?? sampleStart; if (until < sampleEnd) { break; } // We may need to pad out empty space with an vtte box if (trackData.lastCueEndTimestamp < sampleStart) { this.auxWriter.seek(0); const box = vtte(); this.auxBoxWriter.writeBox(box); const body = this.auxTarget._getSlice(0, this.auxWriter.getPos()); const sample = this.createSampleForTrack( trackData, body, trackData.lastCueEndTimestamp, sampleStart - trackData.lastCueEndTimestamp, 'key', ); await this.registerSample(trackData, sample); trackData.lastCueEndTimestamp = sampleStart; } this.auxWriter.seek(0); for (let i = 0; i < trackData.cueQueue.length; i++) { const cue = trackData.cueQueue[i]!; if (cue.timestamp >= sampleEnd) { break; } inlineTimestampRegex.lastIndex = 0; const containsTimestamp = inlineTimestampRegex.test(cue.text); const endTimestamp = cue.timestamp + cue.duration; let sourceId = trackData.cueToSourceId.get(cue); if (sourceId === undefined && sampleEnd < endTimestamp) { // We know this cue will appear in more than one sample, therefore we need to mark it with a // unique ID sourceId = trackData.nextSourceId++; trackData.cueToSourceId.set(cue, sourceId); } if (cue.notes) { // Any notes/comments are included in a special vtta box const box = vtta(cue.notes); this.auxBoxWriter.writeBox(box); } const box = vttc( cue.text, containsTimestamp ? sampleStart : null, cue.identifier ?? null, cue.settings ?? null, sourceId ?? null, ); this.auxBoxWriter.writeBox(box); if (endTimestamp === sampleEnd) { // The cue won't appear in any future sample, so we're done with it trackData.cueQueue.splice(i--, 1); } } const body = this.auxTarget._getSlice(0, this.auxWriter.getPos()); const sample = this.createSampleForTrack(trackData, body, sampleStart, sampleEnd - sampleStart, 'key'); await this.registerSample(trackData, sample); trackData.lastCueEndTimestamp = sampleEnd; } } private createSampleForTrack( trackData: IsobmffTrackData, data: Uint8Array, timestamp: number, duration: number, type: PacketType, ) { const sample: Sample = { timestamp, decodeTimestamp: timestamp, // This may be refined later duration, data, size: data.byteLength, type, timescaleUnitsToNextSample: intoTimescale(duration, trackData.timescale), // Will be refined }; return sample; } private processTimestamps(trackData: IsobmffTrackData, nextSample?: Sample) { if (trackData.timestampProcessingQueue.length === 0) { return; } if (trackData.type === 'audio' && trackData.info.requiresPcmTransformation) { if (!this.isFragmented) { // The first timestamp is the lowest trackData.startTimestampOffset ??= trackData.timestampProcessingQueue[0]!.timestamp; } let totalDuration = 0; // Compute the total duration in the track timescale (which is equal to the amount of PCM audio samples) // and simply say that's how many new samples there are. for (let i = 0; i < trackData.timestampProcessingQueue.length; i++) { const sample = trackData.timestampProcessingQueue[i]!; const duration = intoTimescale(sample.duration, trackData.timescale); totalDuration += duration; } if (trackData.timeToSampleTable.length === 0) { trackData.timeToSampleTable.push({ sampleCount: totalDuration, sampleDelta: 1, }); } else { const lastEntry = last(trackData.timeToSampleTable)!; lastEntry.sampleCount += totalDuration; } trackData.timestampProcessingQueue.length = 0; return; } const sortedTimestamps = trackData.timestampProcessingQueue.map(x => x.timestamp).sort((a, b) => a - b); if (!this.isFragmented) { trackData.startTimestampOffset ??= sortedTimestamps[0]!; } for (let i = 0; i < trackData.timestampProcessingQueue.length; i++) { const sample = trackData.timestampProcessingQueue[i]!; // Since the user only supplies presentation time, but these may be out of order, we reverse-engineer from // that a sensible decode timestamp. The notion of a decode timestamp doesn't really make sense // (presentation timestamp & decode order are all you need), but it is a concept in ISOBMFF so we need to // model it. sample.decodeTimestamp = sortedTimestamps[i]!; const sampleCompositionTimeOffset = intoTimescale(sample.timestamp - sample.decodeTimestamp, trackData.timescale); const durationInTimescale = intoTimescale(sample.duration, trackData.timescale); if (trackData.lastTimescaleUnits !== null) { assert(trackData.lastSample); const timescaleUnits = intoTimescale(sample.decodeTimestamp, trackData.timescale, false); const delta = Math.round(timescaleUnits - trackData.lastTimescaleUnits); assert(delta >= 0); trackData.lastTimescaleUnits += delta; trackData.lastSample.timescaleUnitsToNextSample = delta; if (!this.isFragmented) { let lastTableEntry = last(trackData.timeToSampleTable); assert(lastTableEntry); if (lastTableEntry.sampleCount === 1) { lastTableEntry.sampleDelta = delta; const entryBefore = trackData.timeToSampleTable[trackData.timeToSampleTable.length - 2]; if (entryBefore && entryBefore.sampleDelta === delta) { // If the delta is the same as the previous one, merge the two entries entryBefore.sampleCount++; trackData.timeToSampleTable.pop(); lastTableEntry = entryBefore; } } else if (lastTableEntry.sampleDelta !== delta) { // The delta has changed, so we need a new entry to reach the current sample lastTableEntry.sampleCount--; trackData.timeToSampleTable.push(lastTableEntry = { sampleCount: 1, sampleDelta: delta, }); } if (lastTableEntry.sampleDelta === durationInTimescale) { // The sample's duration matches the delta, so we can increment the count lastTableEntry.sampleCount++; } else { // Add a new entry in order to maintain the last sample's true duration trackData.timeToSampleTable.push({ sampleCount: 1, sampleDelta: durationInTimescale, }); } const lastCompositionTimeOffsetTableEntry = last(trackData.compositionTimeOffsetTable); assert(lastCompositionTimeOffsetTableEntry); if ( lastCompositionTimeOffsetTableEntry.sampleCompositionTimeOffset === sampleCompositionTimeOffset ) { // Simply increment the count lastCompositionTimeOffsetTableEntry.sampleCount++; } else { // The composition time offset has changed, so create a new entry with the new composition time // offset trackData.compositionTimeOffsetTable.push({ sampleCount: 1, sampleCompositionTimeOffset: sampleCompositionTimeOffset, }); } } } else { // Decode timestamp of the first sample trackData.lastTimescaleUnits = intoTimescale(sample.decodeTimestamp, trackData.timescale, false); if (!this.isFragmented) { trackData.timeToSampleTable.push({ sampleCount: 1, sampleDelta: durationInTimescale, }); trackData.compositionTimeOffsetTable.push({ sampleCount: 1, sampleCompositionTimeOffset: sampleCompositionTimeOffset, }); } } trackData.lastSample = sample; } trackData.timestampProcessingQueue.length = 0; assert(trackData.lastSample); assert(trackData.lastTimescaleUnits !== null); if (nextSample !== undefined && trackData.lastSample.timescaleUnitsToNextSample === 0) { assert(nextSample.type === 'key'); // Given the next sample, we can make a guess about the duration of the last sample. This avoids having // the last sample's duration in each fragment be "0" for fragmented files. The guess we make here is // actually correct most of the time, since typically, no delta frame with a lower timestamp follows the key // frame (although it can happen). const timescaleUnits = intoTimescale(nextSample.timestamp, trackData.timescale, false); const delta = Math.round(timescaleUnits - trackData.lastTimescaleUnits); trackData.lastSample.timescaleUnitsToNextSample = delta; } } private async registerSample(trackData: IsobmffTrackData, sample: Sample) { if (sample.type === 'key') { this.processTimestamps(trackData, sample); } trackData.timestampProcessingQueue.push(sample); if (this.isFragmented) { trackData.sampleQueue.push(sample); await this.interleaveSamples(); } else if (this.fastStart === 'reserve') { await this.registerSampleFastStartReserve(trackData, sample); } else { await this.addSampleToTrack(trackData, sample); } } private async addSampleToTrack(trackData: IsobmffTrackData, sample: Sample) { if (!this.isFragmented) { trackData.samples.push(sample); if (this.fastStart === 'reserve') { const maximumPacketCount = trackData.track.metadata.maximumPacketCount; assert(maximumPacketCount !== undefined); if (trackData.samples.length > maximumPacketCount) { throw new Error( `Track #${trackData.track.id} has already reached the maximum packet count` + ` (${maximumPacketCount}). Either add less packets or increase the maximum packet count.`, ); } } } let beginNewChunk = false; if (!trackData.currentChunk) { beginNewChunk = true; } else { // Timestamp don't need to be monotonic (think B-frames), so we may need to update the start timestamp of // the chunk trackData.currentChunk.startTimestamp = Math.min( trackData.currentChunk.startTimestamp, sample.timestamp, ); const currentChunkDuration = sample.timestamp - trackData.currentChunk.startTimestamp; if (this.isFragmented) { // We can only finalize this fragment (and begin a new one) if we know that each track will be able to // start the new one with a key frame. const keyFrameQueuedEverywhere = this.trackDatas.every((otherTrackData) => { if (trackData === otherTrackData) { return sample.type === 'key'; } const firstQueuedSample = otherTrackData.sampleQueue[0]; if (firstQueuedSample) { return firstQueuedSample.type === 'key'; } return otherTrackData.closed; }); if ( currentChunkDuration >= this.minimumFragmentDuration && keyFrameQueuedEverywhere && sample.timestamp > this.maxWrittenTimestamp ) { beginNewChunk = true; await this.finalizeFragment(); } } else { beginNewChunk = currentChunkDuration >= 0.5; // Chunk is long enough, we need a new one } } if (beginNewChunk) { if (trackData.currentChunk) { await this.finalizeCurrentChunk(trackData); } trackData.currentChunk = { startTimestamp: sample.timestamp, samples: [], offset: null, moofOffset: null, }; } assert(trackData.currentChunk); trackData.currentChunk.samples.push(sample); if (this.isFragmented) { this.maxWrittenTimestamp = Math.max(this.maxWrittenTimestamp, sample.timestamp); this.maxWrittenEndTimestamp = Math.max(this.maxWrittenEndTimestamp, sample.timestamp + sample.duration); this.minWrittenTimestamp = Math.min(this.minWrittenTimestamp, sample.timestamp); } } private async finalizeCurrentChunk(trackData: IsobmffTrackData) { assert(!this.isFragmented); assert(this.writer); if (!trackData.currentChunk) return; trackData.finalizedChunks.push(trackData.currentChunk); this.finalizedChunks.push(trackData.currentChunk); let sampleCount = trackData.currentChunk.samples.length; if (trackData.type === 'audio' && trackData.info.requiresPcmTransformation) { sampleCount = trackData.currentChunk.samples .reduce((acc, sample) => acc + intoTimescale(sample.duration, trackData.timescale), 0); } if ( trackData.compactlyCodedChunkTable.length === 0 || last(trackData.compactlyCodedChunkTable)!.samplesPerChunk !== sampleCount ) { trackData.compactlyCodedChunkTable.push({ firstChunk: trackData.finalizedChunks.length, // 1-indexed samplesPerChunk: sampleCount, }); } if (this.fastStart === 'in-memory') { trackData.currentChunk.offset = 0; // We'll compute the proper offset when finalizing return; } // Write out the data trackData.currentChunk.offset = this.writer.getPos(); for (const sample of trackData.currentChunk.samples) { assert(sample.data); this.writer.write(sample.data); sample.data = null; // Can be GC'd } await this.writer.flush(); } private async interleaveSamples(isFinalCall = false) { assert(this.isFragmented); if (!isFinalCall && !this.allTracksAreKnown()) { return; // We can't interleave yet as we don't yet know how many tracks we'll truly have } outer: while (true) { let trackWithMinTimestamp: IsobmffTrackData | null = null; let minTimestamp = Infinity; for (const trackData of this.trackDatas) { if (!isFinalCall && trackData.sampleQueue.length === 0 && !trackData.closed) { break outer; } if (trackData.sampleQueue.length > 0 && trackData.sampleQueue[0]!.timestamp < minTimestamp) { trackWithMinTimestamp = trackData; minTimestamp = trackData.sampleQueue[0]!.timestamp; } } if (!trackWithMinTimestamp) { break; } const sample = trackWithMinTimestamp.sampleQueue.shift()!; await this.addSampleToTrack(trackWithMinTimestamp, sample); } } private async finalizeFragment(flushWriter = !this.isCmaf) { assert(this.isFragmented); const fragmentNumber = this.nextFragmentNumber++; if (fragmentNumber === 1) { const boxWriter = this.initBoxWriter ?? this.boxWriter; assert(boxWriter); if (this.format._options.onMoov) { boxWriter.writer.startTrackingWrites(); } this.ensureOneEnabledTrack(); // Write the moov box now that we have all decoder configs const movieBox = moov(this); boxWriter.writeBox(movieBox); if (this.format._options.onMoov) { const { data, start } = boxWriter.writer.stopTrackingWrites(); this.format._options.onMoov(data, start); } if (this.isCmaf) { assert(this.initWriter); await this.initWriter.flush(); await this.initWriter.finalize(); // Init segment is done // Only now, init the main writer; this way the init writer is fully done before the main writer is // even acquired this.writer = await this.output._getRootWriter(true); this.boxWriter = new IsobmffBoxWriter(this.writer); const stypSize = this.boxWriter.measureBox(styp()); const sidxSize = this.boxWriter.measureBox(sidx(this, 0)); this.segmentHeaderSize = stypSize + sidxSize; this.writer.seek(this.segmentHeaderSize); // Make room for the header to be written later } } assert(this.writer); assert(this.boxWriter); // Not all tracks need to be present in every fragment const tracksInFragment = this.trackDatas.filter(x => x.currentChunk); // Create an initial moof box and measure it; we need this to know where the following mdat box will begin const moofBox = moof(fragmentNumber, tracksInFragment); const moofOffset = this.writer.getPos(); const mdatStartPos = moofOffset + this.boxWriter.measureBox(moofBox); let currentPos = mdatStartPos + MIN_BOX_HEADER_SIZE; let fragmentStartTimestamp = Infinity; for (const trackData of tracksInFragment) { trackData.currentChunk!.offset = currentPos; trackData.currentChunk!.moofOffset = moofOffset; for (const sample of trackData.currentChunk!.samples) { currentPos += sample.size; } fragmentStartTimestamp = Math.min(fragmentStartTimestamp, trackData.currentChunk!.startTimestamp); } const mdatSize = currentPos - mdatStartPos; const needsLargeMdatSize = mdatSize >= 2 ** 32; if (needsLargeMdatSize) { // Shift all offsets by 8. Previously, all chunks were shifted assuming the large box size, but due to what // I suspect is a bug in WebKit, it failed in Safari (when livestreaming with MSE, not for static playback). for (const trackData of tracksInFragment) { trackData.currentChunk!.offset! += MAX_BOX_HEADER_SIZE - MIN_BOX_HEADER_SIZE; } } if (this.format._options.onMoof) { this.writer.startTrackingWrites(); } const newMoofBox = moof(fragmentNumber, tracksInFragment); this.boxWriter.writeBox(newMoofBox); if (this.format._options.onMoof) { const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onMoof(data, start, fragmentStartTimestamp); } assert(this.writer.getPos() === mdatStartPos); if (this.format._options.onMdat) { this.writer.startTrackingWrites(); } const mdatBox = mdat(needsLargeMdatSize); mdatBox.size = mdatSize; this.boxWriter.writeBox(mdatBox); this.writer.seek(mdatStartPos + (needsLargeMdatSize ? MAX_BOX_HEADER_SIZE : MIN_BOX_HEADER_SIZE)); // Write sample data for (const trackData of tracksInFragment) { for (const sample of trackData.currentChunk!.samples) { this.writer.write(sample.data!); sample.data = null; // Can be GC'd } } if (this.format._options.onMdat) { const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onMdat(data, start); } for (const trackData of tracksInFragment) { trackData.finalizedChunks.push(trackData.currentChunk!); this.finalizedChunks.push(trackData.currentChunk!); trackData.currentChunk = null; } if (flushWriter) { await this.writer.flush(); } } private async registerSampleFastStartReserve(trackData: IsobmffTrackData, sample: Sample) { assert(this.writer); assert(this.boxWriter); if (this.allTracksAreKnown()) { if (!this.mdat) { this.ensureOneEnabledTrack(); // We finally know all tracks, let's reserve space for the moov box const moovBox = moov(this); const moovSize = this.boxWriter.measureBox(moovBox); const reservedSize = moovSize + this.computeSampleTableSizeUpperBound() + 4096; // Just a little extra headroom assert(this.ftypSize !== null); this.writer.seek(this.ftypSize + reservedSize); if (this.format._options.onMdat) { this.writer.startTrackingWrites(); } this.mdat = mdat(true); this.boxWriter.writeBox(this.mdat); // Now write everything that was queued for (const trackData of this.trackDatas) { for (const sample of trackData.sampleQueue) { await this.addSampleToTrack(trackData, sample); } trackData.sampleQueue.length = 0; } } await this.addSampleToTrack(trackData, sample); } else { // Queue it for when we know all tracks trackData.sampleQueue.push(sample); } } private computeSampleTableSizeUpperBound() { assert(this.fastStart === 'reserve'); let upperBound = 0; for (const trackData of this.trackDatas) { const n = trackData.track.metadata.maximumPacketCount; assert(n !== undefined); // We validated this earlier // Given the max allowed packet count, compute the space they'll take up in the Sample Table Box, assuming // the worst case for each individual box: // stts box - since it is compactly coded, the maximum length of this table will be 2/3n upperBound += (4 + 4) * Math.ceil(2 / 3 * n); // stss box - 1 entry per sample upperBound += 4 * n; // ctts box - since it is compactly coded, the maximum length of this table will be 2/3n upperBound += (4 + 4) * Math.ceil(2 / 3 * n); // stsc box - since it is compactly coded, the maximum length of this table will be 2/3n upperBound += (4 + 4 + 4) * Math.ceil(2 / 3 * n); // stsz box - 1 entry per sample upperBound += 4 * n; // co64 box - we assume 1 sample per chunk and 64-bit chunk offsets (co64 instead of stco) upperBound += 8 * n; } return upperBound; } // eslint-disable-next-line @typescript-eslint/no-misused-promises override async onTrackClose(track: OutputTrack) { const release = await this.mutex.acquire(); const trackData = this.trackDatas.find(x => x.track === track); if (trackData) { trackData.closed = true; if (trackData.type === 'subtitle' && track.source._codec === 'webvtt') { await this.processWebVTTCues(trackData, Infinity); } this.processTimestamps(trackData); } if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } if (this.isFragmented) { // Since a track is now closed, we may be able to write out chunks that were previously waiting await this.interleaveSamples(); } release(); } ensureOneEnabledTrack() { // If no track of a given type is enabled, force the first one to be enabled. Otherwise the video won't play in // players like QuickTime. // https://github.com/Vanilagy/mediabunny/pull/391 for (const type of ['video', 'audio', 'subtitle'] as TrackType[]) { const tracks = this.trackDatas.filter(t => t.type === type); if (tracks.length === 0) { continue; } const hasEnabled = tracks.some(t => t.track.metadata.disposition?.default !== false); if (!hasEnabled) { const firstTrack = tracks[0]!; firstTrack.track.metadata.disposition = { ...firstTrack.track.metadata.disposition, default: true, }; } } } /** Finalizes the file, making it ready for use. Must be called after all video and audio chunks have been added. */ async finalize() { const release = await this.mutex.acquire(); this.allTracksKnown.resolve(); this.ensureOneEnabledTrack(); for (const trackData of this.trackDatas) { trackData.closed = true; if (trackData.type === 'subtitle' && trackData.track.source._codec === 'webvtt') { await this.processWebVTTCues(trackData, Infinity); } this.processTimestamps(trackData); } if (this.isFragmented) { await this.interleaveSamples(true); await this.finalizeFragment(false); // Don't flush the last fragment as we will flush it with the mfra box } else { for (const trackData of this.trackDatas) { await this.finalizeCurrentChunk(trackData); // Must hold because we will have processed at least one sample assert(trackData.startTimestampOffset !== null); // Shift all of the samples by the start offset. We'll then write out an edit list that will shift them // back to their proper spot in the composition. for (let i = 0; i < trackData.samples.length; i++) { const sample = trackData.samples[i]!; sample.timestamp -= trackData.startTimestampOffset; sample.decodeTimestamp -= trackData.startTimestampOffset; } } } assert(this.writer); assert(this.boxWriter); if (this.fastStart === 'in-memory') { this.mdat = mdat(false); let mdatSize: number; // We know how many chunks there are, but computing the chunk positions requires an iterative approach: // In order to know where the first chunk should go, we first need to know the size of the moov box. But we // cannot write a proper moov box without first knowing all chunk positions. So, we generate a tentative // moov box with placeholder values (0) for the chunk offsets to be able to compute its size. If it then // turns out that appending all chunks exceeds 4 GiB, we need to repeat this process, now with the co64 box // being used in the moov box instead, which will make it larger. After that, we definitely know the final // size of the moov box and can compute the proper chunk positions. for (let i = 0; i < 2; i++) { const movieBox = moov(this); const movieBoxSize = this.boxWriter.measureBox(movieBox); mdatSize = this.boxWriter.measureBox(this.mdat); let currentChunkPos = this.writer.getPos() + movieBoxSize + mdatSize; for (const chunk of this.finalizedChunks) { chunk.offset = currentChunkPos; for (const { data } of chunk.samples) { assert(data); currentChunkPos += data.byteLength; mdatSize += data.byteLength; } } if (currentChunkPos < 2 ** 32) break; if (mdatSize >= 2 ** 32) this.mdat.largeSize = true; } if (this.format._options.onMoov) { this.writer.startTrackingWrites(); } const movieBox = moov(this); this.boxWriter.writeBox(movieBox); if (this.format._options.onMoov) { const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onMoov(data, start); } if (this.format._options.onMdat) { this.writer.startTrackingWrites(); } this.mdat.size = mdatSize!; this.boxWriter.writeBox(this.mdat); for (const chunk of this.finalizedChunks) { for (const sample of chunk.samples) { assert(sample.data); this.writer.write(sample.data); sample.data = null; } } if (this.format._options.onMdat) { const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onMdat(data, start); } } else if (this.isFragmented) { if (this.isCmaf) { const contentSize = this.segmentHeaderSize !== null ? this.writer.getPos() - this.segmentHeaderSize : 0; this.writer.seek(0); // Write styp and sidx to the start; we recently made space for these this.boxWriter.writeBox(styp()); this.boxWriter.writeBox(sidx(this, contentSize)); } else { // Append the mfra box to the end of the file for better random access const startPos = this.writer.getPos(); const mfraBox = mfra(this.trackDatas); this.boxWriter.writeBox(mfraBox); // Patch the 'size' field of the mfro box at the end of the mfra box now that we know its actual size const mfraBoxSize = this.writer.getPos() - startPos; this.writer.seek(this.writer.getPos() - 4); this.boxWriter.writeU32(mfraBoxSize); } } else { assert(this.mdat); const mdatPos = this.boxWriter.offsets.get(this.mdat); assert(mdatPos !== undefined); const mdatSize = this.writer.getPos() - mdatPos; this.mdat.size = mdatSize; this.mdat.largeSize = mdatSize >= 2 ** 32; // Only use the large size if we need it this.boxWriter.patchBox(this.mdat); if (this.format._options.onMdat) { const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onMdat(data, start); } const movieBox = moov(this); if (this.fastStart === 'reserve') { assert(this.ftypSize !== null); this.writer.seek(this.ftypSize); if (this.format._options.onMoov) { this.writer.startTrackingWrites(); } this.boxWriter.writeBox(movieBox); // Fill the remaining space with a free box. If there are less than 8 bytes left, sucks I guess const remainingSpace = this.boxWriter.offsets.get(this.mdat)! - this.writer.getPos(); this.boxWriter.writeBox(free(remainingSpace)); } else { if (this.format._options.onMoov) { this.writer.startTrackingWrites(); } this.boxWriter.writeBox(movieBox); } if (this.format._options.onMoov) { const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onMoov(data, start); } } release(); } } ===== src/isobmff/isobmff-demuxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { TrackType } from '../output'; import { parseAacAudioSpecificConfig } from '../../shared/aac-misc'; import { AacCodecInfo, AudioCodec, extractAudioCodecString, extractVideoCodecString, MediaCodec, OPUS_SAMPLE_RATE, parsePcmCodec, PCM_AUDIO_CODECS, PcmAudioCodec, VideoCodec, } from '../codec'; import { Av1CodecInfo, AvcDecoderConfigurationRecord, extractAv1CodecInfoFromPacket, extractVp9CodecInfoFromPacket, FlacBlockType, HevcDecoderConfigurationRecord, Vp9CodecInfo, parseEac3Config, getEac3SampleRate, getEac3ChannelCount, AC3_ACMOD_CHANNEL_COUNTS, } from '../codec-data'; import { Demuxer } from '../demuxer'; import { Input } from '../input'; import { InputAudioTrackBacking, InputTrackBacking, InputVideoTrackBacking, } from '../input-track'; import { PacketRetrievalOptions } from '../media-sink'; import { assert, binarySearchExact, binarySearchLessOrEqual, bytesToHexString, COLOR_PRIMARIES_MAP_INVERSE, findLastIndex, isIso639Dash2LanguageCode, last, MATRIX_COEFFICIENTS_MAP_INVERSE, normalizeRotation, roundToMultiple, Rotation, textDecoder, TransformationMatrix, TRANSFER_CHARACTERISTICS_MAP_INVERSE, UNDETERMINED_LANGUAGE, toDataView, roundIfAlmostInteger, hexStringToBytes, HEX_STRING_REGEX, } from '../misc'; import { EncodedPacket, PLACEHOLDER_DATA } from '../packet'; import { buildIsobmffMimeType, parsePsshBoxContents, psshBoxesAreEqual, PsshBox } from './isobmff-misc'; import { MAX_BOX_HEADER_SIZE, MIN_BOX_HEADER_SIZE, readBoxHeader, readDataBox, readFixed_16_16, readFixed_2_30, readIsomVariableInteger, readMetadataStringShort, } from './isobmff-reader'; import { FileSlice, readBytes, readF64Be, readI16Be, readI32Be, readI64Be, Reader, readU16Be, readU24Be, readU32Be, readU64Be, readU8, readAscii, } from '../reader'; import { DEFAULT_TRACK_DISPOSITION, MetadataTags, RichImageData, TrackDisposition } from '../metadata'; import { AC3_SAMPLE_RATES } from '../../shared/ac3-misc'; import { Bitstream } from '../../shared/bitstream'; import { Aes128CbcContext } from '../aes'; import { Logging } from '../logging'; type InternalTrack = { id: number; demuxer: IsobmffDemuxer; trackBacking: InputTrackBacking | null; disposition: TrackDisposition; timescale: number; durationInMovieTimescale: number; durationInMediaTimescale: number; rotation: Rotation; internalCodecId: string | null; name: string | null; languageCode: string; sampleTableByteOffset: number | null; // null when the track's sample table is another file (ominous ik 👀) sampleTable: SampleTable | null; fragmentLookupTable: FragmentLookupTableEntry[]; currentFragmentState: FragmentTrackState | null; /** * List of all encountered fragment offsets alongside their timestamps. This list never gets truncated, but memory * consumption should be negligible. */ fragmentPositionCache: { moofOffset: number; startTimestamp: number; endTimestamp: number; }[]; /** The segment durations of all edit list entries leading up to the main one (from which the offset is taken.) */ editListPreviousSegmentDurations: number; /** The media time offset of the main edit list entry (with media time !== -1) */ editListOffset: number; /** Set when the track's samples are encrypted using a supported scheme (cenc/cens/cbcs), parsed from sinf/tenc. */ encryptionInfo: TrackEncryptionInfo | null; /** For non-fragmented encrypted tracks: parsed saiz+saio from stbl; aux info is fetched lazily on first use. */ encryptionAuxInfo: SampleEncryptionAuxInfo | null; frmaCodecString: string | null; } & ({ info: null; } | { info: { type: 'video'; width: number; height: number; squarePixelWidth: number; squarePixelHeight: number; codec: VideoCodec | null; codecDescription: Uint8Array | null; colorSpace: VideoColorSpaceInit | null; avcType: 1 | 3 | null; avcCodecInfo: AvcDecoderConfigurationRecord | null; hevcCodecInfo: HevcDecoderConfigurationRecord | null; vp9CodecInfo: Vp9CodecInfo | null; av1CodecInfo: Av1CodecInfo | null; }; } | { info: { type: 'audio'; numberOfChannels: number; sampleRate: number; codec: AudioCodec | null; codecDescription: Uint8Array | null; aacCodecInfo: AacCodecInfo | null; pcmLittleEndian: boolean; pcmSampleSize: number | null; }; }); type InternalVideoTrack = InternalTrack & { info: { type: 'video' } }; type InternalAudioTrack = InternalTrack & { info: { type: 'audio' } }; type SampleTable = { sampleTimingEntries: SampleTimingEntry[]; sampleCompositionTimeOffsets: SampleCompositionTimeOffsetEntry[]; sampleSizes: number[]; keySampleIndices: number[] | null; // Samples that are keyframes chunkOffsets: number[]; sampleToChunk: SampleToChunkEntry[]; presentationTimestamps: { presentationTimestamp: number; sampleIndex: number; }[] | null; /** * Provides a fast map from sample index to index in the sorted presentation timestamps array - so, a fast map from * decode order to presentation order. */ presentationTimestampIndexMap: number[] | null; }; type SampleTimingEntry = { startIndex: number; startDecodeTimestamp: number; count: number; delta: number; }; type SampleCompositionTimeOffsetEntry = { startIndex: number; count: number; offset: number; }; type SampleToChunkEntry = { startSampleIndex: number; startChunkIndex: number; samplesPerChunk: number; sampleDescriptionIndex: number; }; type FragmentTrackDefaults = { trackId: number; defaultSampleDescriptionIndex: number; defaultSampleDuration: number; defaultSampleSize: number; defaultSampleFlags: number; }; type FragmentLookupTableEntry = { timestamp: number; moofOffset: number; }; type FragmentTrackState = { baseDataOffset: number; sampleDescriptionIndex: number | null; defaultSampleDuration: number | null; defaultSampleSize: number | null; defaultSampleFlags: number | null; startTimestamp: number | null; encryptionAuxInfo: SampleEncryptionAuxInfo | null; }; type FragmentTrackData = { track: InternalTrack; // Kept as state for the presence of multiple trun boxes currentTimestamp: number; currentOffset: number; startTimestamp: number; endTimestamp: number; firstKeyFrameTimestamp: number | null; samples: FragmentTrackSample[]; presentationTimestamps: { presentationTimestamp: number; sampleIndex: number; }[]; startTimestampIsFinal: boolean; encryptionAuxInfo: SampleEncryptionAuxInfo | null; }; type FragmentTrackSample = { presentationTimestamp: number; duration: number; byteOffset: number; byteSize: number; isKeyFrame: boolean; encryption: SampleEncryptionInfo | null; }; type Fragment = { moofOffset: number; moofSize: number; implicitBaseDataOffset: number; trackData: Map; psshBoxes: PsshBox[]; }; type TrackEncryptionInfo = { scheme: 'cenc' | 'cens' | 'cbcs'; defaultKid: string | null; defaultIsProtected: boolean | null; defaultPerSampleIvSize: number | null; defaultConstantIv: Uint8Array | null; defaultCryptByteBlock: number | null; defaultSkipByteBlock: number | null; }; type SampleEncryptionInfo = { iv: Uint8Array; subsamples: { clearLen: number; protectedLen: number; }[] | null; }; /** * Holds parsed saiz+saio state. The encryption info itself lives at a file offset and is fetched lazily. * For fragmented files this state is per-traf; for non-fragmented files it's per-track (on stbl). */ type SampleEncryptionAuxInfo = { defaultSampleInfoSize: number; sampleSizes: Uint8Array | null; sampleCount: number; offset: number | null; // Absolute file offset of the first sample's aux info resolved: SampleEncryptionInfo[] | null; }; export class IsobmffDemuxer extends Demuxer { reader: Reader; moovSlice: FileSlice | null = null; currentTrack: InternalTrack | null = null; tracks: InternalTrack[] = []; metadataPromise: Promise | null = null; movieTimescale = -1; movieDurationInTimescale = -1; isQuickTime = false; metadataTags: MetadataTags = {}; currentMetadataKeys: Map | null = null; isFragmented = false; fragmentTrackDefaults: FragmentTrackDefaults[] = []; psshBoxes: PsshBox[] = []; currentFragment: Fragment | null = null; /** * Caches the last fragment that was read. Based on the assumption that there will be multiple reads to the * same fragment in quick succession. */ lastReadFragment: Fragment | null = null; decryptionKeyCache = new Map>(); constructor(input: Input) { super(input); this.reader = input._reader; } override async getTrackBackings() { await this.readMetadata(); return this.tracks.map(track => track.trackBacking!); } override async getMimeType() { await this.readMetadata(); const backings = await this.getTrackBackings(); const codecStrings = await Promise.all(backings.map( x => x.getDecoderConfig().then(c => c?.codec ?? null), )); return buildIsobmffMimeType({ isQuickTime: this.isQuickTime, hasVideo: this.tracks.some(x => x.info?.type === 'video'), hasAudio: this.tracks.some(x => x.info?.type === 'audio'), codecStrings: codecStrings.filter(Boolean) as string[], }); } async getMetadataTags() { await this.readMetadata(); return this.metadataTags; } readMetadata() { return this.metadataPromise ??= (async () => { let currentPos = 0; let lookForMfraBox = false; while (true) { let slice = this.reader.requestSliceRange(currentPos, MIN_BOX_HEADER_SIZE, MAX_BOX_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const startPos = currentPos; const boxInfo = readBoxHeader(slice); if (!boxInfo) { break; } if (boxInfo.name === 'ftyp' || boxInfo.name === 'styp') { const majorBrand = readAscii(slice, 4); this.isQuickTime = majorBrand === 'qt '; } else if (boxInfo.name === 'moov') { // Found moov, load it let moovSlice = this.reader.requestSlice(slice.filePos, boxInfo.contentSize); if (moovSlice instanceof Promise) moovSlice = await moovSlice; if (!moovSlice) break; this.moovSlice = moovSlice; this.readContiguousBoxes(this.moovSlice); for (const track of this.tracks) { // Modify the edit list offset based on the previous segment durations. They are in different // timescales, so we first convert to seconds and then into the track timescale. const previousSegmentDurationsInSeconds = track.editListPreviousSegmentDurations / this.movieTimescale; track.editListOffset -= Math.round(previousSegmentDurationsInSeconds * track.timescale); } lookForMfraBox = this.isFragmented && this.reader.fileSize !== null && this.reader.fileSize > startPos + boxInfo.totalSize; // There's more after the moov box break; } else if (boxInfo.name === 'moof') { if (!this.input._initInput) { throw new Error( '"moof" box encountered with no "moov" box present; this file is likely a Segment as' + ' described in ISO/IEC 14496-12 Section 8.16. A separate init file that contains a "moov"' + ' box is required to read this file, please provide it using InputOptions.initInput.', ); } const initDemuxer = (await this.input._initInput._getDemuxer()) as IsobmffDemuxer; if (initDemuxer.constructor !== IsobmffDemuxer) { throw new Error('Init input must match the input\'s format.'); } await initDemuxer.readMetadata(); this.movieTimescale = initDemuxer.movieTimescale; this.movieDurationInTimescale = initDemuxer.movieDurationInTimescale; this.metadataTags = initDemuxer.metadataTags; this.isFragmented = true; this.fragmentTrackDefaults = initDemuxer.fragmentTrackDefaults; this.psshBoxes = initDemuxer.psshBoxes; // Create tracks from the init input's tracks for (const foreignTrack of initDemuxer.tracks) { const track: InternalTrack = { id: foreignTrack.id, demuxer: this, trackBacking: null, disposition: foreignTrack.disposition, timescale: foreignTrack.timescale, durationInMediaTimescale: foreignTrack.durationInMediaTimescale, durationInMovieTimescale: foreignTrack.durationInMovieTimescale, rotation: foreignTrack.rotation, internalCodecId: foreignTrack.internalCodecId, name: foreignTrack.name, languageCode: foreignTrack.languageCode, sampleTableByteOffset: null, sampleTable: null, fragmentLookupTable: [], currentFragmentState: null, fragmentPositionCache: [], editListPreviousSegmentDurations: foreignTrack.editListPreviousSegmentDurations, editListOffset: foreignTrack.editListOffset, encryptionInfo: foreignTrack.encryptionInfo, encryptionAuxInfo: null, frmaCodecString: null, info: foreignTrack.info, }; if (foreignTrack.trackBacking) { assert(track.info); if (track.info.type === 'video' && track.info.width !== -1) { const videoTrack = track as InternalVideoTrack; track.trackBacking = new IsobmffVideoTrackBacking(videoTrack); this.tracks.push(track); } else if (track.info.type === 'audio' && track.info.numberOfChannels !== -1) { const audioTrack = track as InternalAudioTrack; track.trackBacking = new IsobmffAudioTrackBacking(audioTrack); this.tracks.push(track); } } else { // The track didn't have enough info to warrant a backing } } lookForMfraBox = false; // No point in doing it for segment files break; } currentPos = startPos + boxInfo.totalSize; } if (lookForMfraBox) { assert(this.reader.fileSize !== null); // The last 4 bytes may contain the size of the mfra box at the end of the file let lastWordSlice = this.reader.requestSlice(this.reader.fileSize - 4, 4); if (lastWordSlice instanceof Promise) lastWordSlice = await lastWordSlice; assert(lastWordSlice); const lastWord = readU32Be(lastWordSlice); const potentialMfraPos = this.reader.fileSize - lastWord; if (potentialMfraPos >= 0 && potentialMfraPos <= this.reader.fileSize - MAX_BOX_HEADER_SIZE) { let mfraHeaderSlice = this.reader.requestSliceRange( potentialMfraPos, MIN_BOX_HEADER_SIZE, MAX_BOX_HEADER_SIZE, ); if (mfraHeaderSlice instanceof Promise) mfraHeaderSlice = await mfraHeaderSlice; if (mfraHeaderSlice) { const boxInfo = readBoxHeader(mfraHeaderSlice); if (boxInfo && boxInfo.name === 'mfra') { // We found the mfra box, allowing for much better random access. Let's parse it. let mfraSlice = this.reader.requestSlice(mfraHeaderSlice.filePos, boxInfo.contentSize); if (mfraSlice instanceof Promise) mfraSlice = await mfraSlice; if (mfraSlice) { this.readContiguousBoxes(mfraSlice); } } } } } })(); } getSampleTableForTrack(internalTrack: InternalTrack) { if (internalTrack.sampleTable) { return internalTrack.sampleTable; } const sampleTable: SampleTable = { sampleTimingEntries: [], sampleCompositionTimeOffsets: [], sampleSizes: [], keySampleIndices: null, chunkOffsets: [], sampleToChunk: [], presentationTimestamps: null, presentationTimestampIndexMap: null, }; internalTrack.sampleTable = sampleTable; if (internalTrack.sampleTableByteOffset === null) { // There's no sample table to read, it's in another file (happens with segments) return sampleTable; } assert(this.moovSlice); const stblContainerSlice = this.moovSlice.slice(internalTrack.sampleTableByteOffset); this.currentTrack = internalTrack; this.traverseBox(stblContainerSlice); this.currentTrack = null; const isPcmCodec = internalTrack.info?.type === 'audio' && internalTrack.info.codec && (PCM_AUDIO_CODECS as readonly string[]).includes(internalTrack.info.codec); if (isPcmCodec && sampleTable.sampleCompositionTimeOffsets.length === 0) { // If the audio has PCM samples, the way the samples are defined in the sample table is somewhat // suboptimal: Each individual audio sample is its own sample, meaning we can have 48000 samples per second. // Because we treat each sample as its own atomic unit that can be decoded, this would lead to a huge // amount of very short samples for PCM audio. So instead, we make a transformation: If the audio is in PCM, // we say that each chunk (that normally holds many samples) now is one big sample. We can this because // the samples in the chunk are contiguous and the format is PCM, so the entire chunk as one thing still // encodes valid audio information. assert(internalTrack.info?.type === 'audio'); const pcmInfo = parsePcmCodec(internalTrack.info.codec as PcmAudioCodec); const newSampleTimingEntries: SampleTimingEntry[] = []; const newSampleSizes: number[] = []; for (let i = 0; i < sampleTable.sampleToChunk.length; i++) { const chunkEntry = sampleTable.sampleToChunk[i]!; const nextEntry = sampleTable.sampleToChunk[i + 1]; const chunkCount = (nextEntry ? nextEntry.startChunkIndex : sampleTable.chunkOffsets.length) - chunkEntry.startChunkIndex; for (let j = 0; j < chunkCount; j++) { const startSampleIndex = chunkEntry.startSampleIndex + j * chunkEntry.samplesPerChunk; const endSampleIndex = startSampleIndex + chunkEntry.samplesPerChunk; // Exclusive, outside of chunk const startTimingEntryIndex = binarySearchLessOrEqual( sampleTable.sampleTimingEntries, startSampleIndex, x => x.startIndex, ); const startTimingEntry = sampleTable.sampleTimingEntries[startTimingEntryIndex]!; const endTimingEntryIndex = binarySearchLessOrEqual( sampleTable.sampleTimingEntries, endSampleIndex, x => x.startIndex, ); const endTimingEntry = sampleTable.sampleTimingEntries[endTimingEntryIndex]!; const firstSampleTimestamp = startTimingEntry.startDecodeTimestamp + (startSampleIndex - startTimingEntry.startIndex) * startTimingEntry.delta; const lastSampleTimestamp = endTimingEntry.startDecodeTimestamp + (endSampleIndex - endTimingEntry.startIndex) * endTimingEntry.delta; const delta = lastSampleTimestamp - firstSampleTimestamp; const lastSampleTimingEntry = last(newSampleTimingEntries); if (lastSampleTimingEntry && lastSampleTimingEntry.delta === delta) { lastSampleTimingEntry.count++; } else { // One sample for the entire chunk newSampleTimingEntries.push({ startIndex: chunkEntry.startChunkIndex + j, startDecodeTimestamp: firstSampleTimestamp, count: 1, delta, }); } // Instead of determining the chunk's size by looping over the samples sizes in the sample table, we // can directly compute it as we know how many PCM frames are in this chunk, and the size of each // PCM frame. This also improves compatibility with some files which fail to write proper sample // size values into their sample tables in the PCM case. const chunkSize = chunkEntry.samplesPerChunk * pcmInfo.sampleSize * internalTrack.info.numberOfChannels; newSampleSizes.push(chunkSize); } chunkEntry.startSampleIndex = chunkEntry.startChunkIndex; chunkEntry.samplesPerChunk = 1; } sampleTable.sampleTimingEntries = newSampleTimingEntries; sampleTable.sampleSizes = newSampleSizes; } if (sampleTable.sampleCompositionTimeOffsets.length > 0) { // If composition time offsets are defined, we must build a list of all presentation timestamps and then // sort them sampleTable.presentationTimestamps = []; for (const entry of sampleTable.sampleTimingEntries) { for (let i = 0; i < entry.count; i++) { sampleTable.presentationTimestamps.push({ presentationTimestamp: entry.startDecodeTimestamp + i * entry.delta, sampleIndex: entry.startIndex + i, }); } } for (const entry of sampleTable.sampleCompositionTimeOffsets) { for (let i = 0; i < entry.count; i++) { const sampleIndex = entry.startIndex + i; const sample = sampleTable.presentationTimestamps[sampleIndex]; if (!sample) { continue; } sample.presentationTimestamp += entry.offset; } } sampleTable.presentationTimestamps.sort((a, b) => a.presentationTimestamp - b.presentationTimestamp); sampleTable.presentationTimestampIndexMap = Array(sampleTable.presentationTimestamps.length).fill(-1); for (let i = 0; i < sampleTable.presentationTimestamps.length; i++) { sampleTable.presentationTimestampIndexMap[sampleTable.presentationTimestamps[i]!.sampleIndex] = i; } } else { // If they're not defined, we can simply use the decode timestamps as presentation timestamps } return sampleTable; } async readFragment(startPos: number): Promise { if (this.lastReadFragment?.moofOffset === startPos) { return this.lastReadFragment; } let headerSlice = this.reader.requestSliceRange(startPos, MIN_BOX_HEADER_SIZE, MAX_BOX_HEADER_SIZE); if (headerSlice instanceof Promise) headerSlice = await headerSlice; assert(headerSlice); const moofBoxInfo = readBoxHeader(headerSlice); assert(moofBoxInfo?.name === 'moof'); let entireSlice = this.reader.requestSlice(startPos, moofBoxInfo.totalSize); if (entireSlice instanceof Promise) entireSlice = await entireSlice; assert(entireSlice); this.traverseBox(entireSlice); const fragment = this.lastReadFragment; assert(fragment && fragment.moofOffset === startPos); for (const [, trackData] of fragment.trackData) { const track = trackData.track; const { fragmentPositionCache } = track; if (!trackData.startTimestampIsFinal) { // It may be that some tracks don't define the base decode time, i.e. when the fragment begins. This // we'll need to figure out the start timestamp another way. We'll compute the timestamp by accessing // the lookup entries and fragment cache, which works out nicely with the lookup algorithm: If these // exist, then the lookup will automatically start at the furthest possible point. If they don't, the // lookup starts sequentially from the start, incrementally summing up all fragment durations. It's sort // of implicit, but it ends up working nicely. const lookupEntry = track.fragmentLookupTable.find(x => x.moofOffset === fragment.moofOffset); if (lookupEntry) { // There's a lookup entry, let's use its timestamp offsetFragmentTrackDataByTimestamp(trackData, lookupEntry.timestamp); } else { const lastCacheIndex = binarySearchLessOrEqual( fragmentPositionCache, fragment.moofOffset - 1, x => x.moofOffset, ); if (lastCacheIndex !== -1) { // Let's use the timestamp of the previous fragment in the cache const lastCache = fragmentPositionCache[lastCacheIndex]!; offsetFragmentTrackDataByTimestamp(trackData, lastCache.endTimestamp); } else { // We're the first fragment I guess, "offset by 0" } } trackData.startTimestampIsFinal = true; } // Let's remember that a fragment with a given timestamp is here, speeding up future lookups if no // lookup table exists const insertionIndex = binarySearchLessOrEqual( fragmentPositionCache, trackData.startTimestamp, x => x.startTimestamp, ); if ( insertionIndex === -1 || fragmentPositionCache[insertionIndex]!.moofOffset !== fragment.moofOffset ) { fragmentPositionCache.splice(insertionIndex + 1, 0, { moofOffset: fragment.moofOffset, startTimestamp: trackData.startTimestamp, endTimestamp: trackData.endTimestamp, }); } // If senc wasn't parsed but saiz+saio were, fetch the aux info now and stamp each sample with it if (trackData.encryptionAuxInfo && track.encryptionInfo) { const entries = await resolveEncryptionAuxInfo( this.reader, track.encryptionInfo, trackData.encryptionAuxInfo, ); for (let i = 0; i < Math.min(trackData.samples.length, entries.length); i++) { const entry = entries[i]!; trackData.samples[i]!.encryption = entry; } } } return fragment; } readContiguousBoxes(slice: FileSlice) { const startIndex = slice.filePos; while (slice.filePos - startIndex <= slice.length - MIN_BOX_HEADER_SIZE) { const foundBox = this.traverseBox(slice); if (!foundBox) { break; } } } // eslint-disable-next-line @stylistic/generator-star-spacing *iterateContiguousBoxes(slice: FileSlice) { const startIndex = slice.filePos; while (slice.filePos - startIndex <= slice.length - MIN_BOX_HEADER_SIZE) { const startPos = slice.filePos; const boxInfo = readBoxHeader(slice); if (!boxInfo) { break; } yield { boxInfo, slice }; slice.filePos = startPos + boxInfo.totalSize; } } traverseBox(slice: FileSlice): boolean { const startPos = slice.filePos; const boxInfo = readBoxHeader(slice); if (!boxInfo) { return false; } const contentStartPos = slice.filePos; const boxEndPos = startPos + boxInfo.totalSize; switch (boxInfo.name) { case 'mdia': case 'minf': case 'dinf': case 'mfra': case 'edts': case 'sinf': case 'schi': { this.readContiguousBoxes(slice.slice(contentStartPos, boxInfo.contentSize)); }; break; case 'mvhd': { const version = readU8(slice); slice.skip(3); // Flags if (version === 1) { slice.skip(8 + 8); this.movieTimescale = readU32Be(slice); this.movieDurationInTimescale = readU64Be(slice); } else { slice.skip(4 + 4); this.movieTimescale = readU32Be(slice); this.movieDurationInTimescale = readU32Be(slice); } }; break; case 'trak': { const track = { id: -1, demuxer: this, trackBacking: null, disposition: { ...DEFAULT_TRACK_DISPOSITION, primary: false, }, info: null, timescale: -1, durationInMovieTimescale: -1, durationInMediaTimescale: -1, rotation: 0, internalCodecId: null, name: null, languageCode: UNDETERMINED_LANGUAGE, sampleTableByteOffset: -1, sampleTable: null, fragmentLookupTable: [], currentFragmentState: null, fragmentPositionCache: [], editListPreviousSegmentDurations: 0, editListOffset: 0, encryptionInfo: null, encryptionAuxInfo: null, frmaCodecString: null, } satisfies InternalTrack as InternalTrack; this.currentTrack = track; this.readContiguousBoxes(slice.slice(contentStartPos, boxInfo.contentSize)); if (track.id !== -1 && track.timescale !== -1 && track.info !== null) { if (track.info.type === 'video' && track.info.width !== -1) { const videoTrack = track as InternalVideoTrack; track.trackBacking = new IsobmffVideoTrackBacking(videoTrack); this.tracks.push(track); } else if (track.info.type === 'audio' && track.info.numberOfChannels !== -1) { const audioTrack = track as InternalAudioTrack; track.trackBacking = new IsobmffAudioTrackBacking(audioTrack); this.tracks.push(track); } } this.currentTrack = null; }; break; case 'tkhd': { const track = this.currentTrack; if (!track) { break; } const version = readU8(slice); const flags = readU24Be(slice); // Spec says disabled tracks are to be treated like they don't exist, but in practice, they are treated // more like non-default tracks. const trackEnabled = !!(flags & 0x1); track.disposition.default = trackEnabled; // Skip over creation & modification time to reach the track ID if (version === 0) { slice.skip(8); track.id = readU32Be(slice); slice.skip(4); track.durationInMovieTimescale = readU32Be(slice); } else if (version === 1) { slice.skip(16); track.id = readU32Be(slice); slice.skip(4); track.durationInMovieTimescale = readU64Be(slice); } else { throw new Error(`Incorrect track header version ${version}.`); } slice.skip(2 * 4 + 2 + 2 + 2 + 2); const matrix: TransformationMatrix = [ readFixed_16_16(slice), readFixed_16_16(slice), readFixed_2_30(slice), readFixed_16_16(slice), readFixed_16_16(slice), readFixed_2_30(slice), readFixed_16_16(slice), readFixed_16_16(slice), readFixed_2_30(slice), ]; const rotation = normalizeRotation(roundToMultiple(extractRotationFromMatrix(matrix), 90)); assert(rotation === 0 || rotation === 90 || rotation === 180 || rotation === 270); track.rotation = rotation; }; break; case 'elst': { const track = this.currentTrack; if (!track) { break; } const version = readU8(slice); slice.skip(3); // Flags let relevantEntryFound = false; let previousSegmentDurations = 0; const entryCount = readU32Be(slice); for (let i = 0; i < entryCount; i++) { const segmentDuration = version === 1 ? readU64Be(slice) : readU32Be(slice); const mediaTime = version === 1 ? readI64Be(slice) : readI32Be(slice); const mediaRate = readFixed_16_16(slice); if (segmentDuration === 0) { // Don't care continue; } if (relevantEntryFound) { Logging._warn( 'Unsupported edit list: multiple edits are not currently supported. Only using first edit.', ); break; } if (mediaTime === -1) { previousSegmentDurations += segmentDuration; continue; } if (mediaRate !== 1) { Logging._warn('Unsupported edit list entry: media rate must be 1.'); break; } track.editListPreviousSegmentDurations = previousSegmentDurations; track.editListOffset = mediaTime; relevantEntryFound = true; } }; break; case 'mdhd': { const track = this.currentTrack; if (!track) { break; } const version = readU8(slice); slice.skip(3); // Flags if (version === 0) { slice.skip(8); track.timescale = readU32Be(slice); track.durationInMediaTimescale = readU32Be(slice); } else if (version === 1) { slice.skip(16); track.timescale = readU32Be(slice); track.durationInMediaTimescale = readU64Be(slice); } let language = readU16Be(slice); if (language > 0) { track.languageCode = ''; for (let i = 0; i < 3; i++) { track.languageCode = String.fromCharCode(0x60 + (language & 0b11111)) + track.languageCode; language >>= 5; } if (!isIso639Dash2LanguageCode(track.languageCode)) { // Sometimes the bytes are garbage track.languageCode = UNDETERMINED_LANGUAGE; } } }; break; case 'hdlr': { const track = this.currentTrack; if (!track) { break; } slice.skip(8); // Version + flags + pre-defined const handlerType = readAscii(slice, 4); if (handlerType === 'vide') { track.info = { type: 'video', width: -1, height: -1, squarePixelWidth: -1, squarePixelHeight: -1, codec: null, codecDescription: null, colorSpace: null, avcType: null, avcCodecInfo: null, hevcCodecInfo: null, vp9CodecInfo: null, av1CodecInfo: null, }; } else if (handlerType === 'soun') { track.info = { type: 'audio', numberOfChannels: -1, sampleRate: -1, codec: null, codecDescription: null, aacCodecInfo: null, pcmLittleEndian: false, pcmSampleSize: null, }; } }; break; case 'stbl': { const track = this.currentTrack; if (!track) { break; } track.sampleTableByteOffset = startPos; this.readContiguousBoxes(slice.slice(contentStartPos, boxInfo.contentSize)); }; break; case 'stsd': { const track = this.currentTrack; if (!track) { break; } if (track.info === null || track.sampleTable) { break; } const stsdVersion = readU8(slice); slice.skip(3); // Flags const entries = readU32Be(slice); for (let i = 0; i < entries; i++) { const sampleBoxStartPos = slice.filePos; const sampleBoxInfo = readBoxHeader(slice); if (!sampleBoxInfo) { break; } track.internalCodecId = sampleBoxInfo.name; const lowercaseBoxName = sampleBoxInfo.name.toLowerCase(); if (track.info.type === 'video') { slice.skip(6 * 1 + 2 + 2 + 2 + 3 * 4); track.info.width = readU16Be(slice); track.info.height = readU16Be(slice); track.info.squarePixelWidth = track.info.width; track.info.squarePixelHeight = track.info.height; slice.skip(4 + 4 + 4 + 2 + 32 + 2 + 2); track.frmaCodecString = null; this.readContiguousBoxes( slice.slice( slice.filePos, (sampleBoxStartPos + sampleBoxInfo.totalSize) - slice.filePos, ), ); const codecName = lowercaseBoxName === 'encv' ? track.frmaCodecString : lowercaseBoxName; track.frmaCodecString = null; if (codecName === 'avc1' || codecName === 'avc3') { track.info.codec = 'avc'; track.info.avcType = codecName === 'avc1' ? 1 : 3; } else if (codecName === 'hvc1' || codecName === 'hev1') { track.info.codec = 'hevc'; } else if (codecName === 'vp08') { track.info.codec = 'vp8'; } else if (codecName === 'vp09') { track.info.codec = 'vp9'; } else if (codecName === 'av01') { track.info.codec = 'av1'; } else if (codecName === null) { Logging._warn(`Unknown encrypted video codec due to missing frma box.`); } else { Logging._warn(`Unsupported video codec (sample entry type '${sampleBoxInfo.name}').`); } } else { slice.skip(6 * 1 + 2); const version = readU16Be(slice); slice.skip(3 * 2); let channelCount = readU16Be(slice); let sampleSize = readU16Be(slice); slice.skip(2 * 2); // Can't use fixed16_16 as that's signed let sampleRate = readU32Be(slice) / 0x10000; let lpcmFlags: number | null = null; if (stsdVersion === 0 && version > 0) { // Additional QuickTime fields if (version === 1) { slice.skip(4); sampleSize = 8 * readU32Be(slice); slice.skip(2 * 4); } else if (version === 2) { slice.skip(4); sampleRate = readF64Be(slice); channelCount = readU32Be(slice); slice.skip(4); // Always 0x7f000000 sampleSize = readU32Be(slice); lpcmFlags = readU32Be(slice); slice.skip(2 * 4); } } track.info.numberOfChannels = channelCount; track.info.sampleRate = sampleRate; track.frmaCodecString = null; this.readContiguousBoxes( slice.slice( slice.filePos, (sampleBoxStartPos + sampleBoxInfo.totalSize) - slice.filePos, ), ); const codecName = lowercaseBoxName === 'enca' ? track.frmaCodecString : lowercaseBoxName; track.frmaCodecString = null; // developer.apple.com/documentation/quicktime-file-format/sound_sample_descriptions/ if (codecName === 'mp4a') { // The codec is set by the esds box } else if (codecName === 'opus') { track.info.codec = 'opus'; track.info.sampleRate = OPUS_SAMPLE_RATE; // Always the same } else if (codecName === 'flac') { track.info.codec = 'flac'; } else if (codecName === 'ulaw') { track.info.codec = 'ulaw'; } else if (codecName === 'alaw') { track.info.codec = 'alaw'; } else if (codecName === 'ac-3') { track.info.codec = 'ac3'; } else if (codecName === 'ec-3') { track.info.codec = 'eac3'; } else if (codecName === 'twos') { if (sampleSize === 8) { track.info.codec = 'pcm-s8'; } else if (sampleSize === 16) { track.info.codec = track.info.pcmLittleEndian ? 'pcm-s16' : 'pcm-s16be'; } else { Logging._warn(`Unsupported sample size ${sampleSize} for codec 'twos'.`); track.info.codec = null; } } else if (codecName === 'sowt') { if (sampleSize === 8) { track.info.codec = 'pcm-s8'; } else if (sampleSize === 16) { track.info.codec = 'pcm-s16'; } else { Logging._warn(`Unsupported sample size ${sampleSize} for codec 'sowt'.`); track.info.codec = null; } } else if (codecName === 'raw ') { track.info.codec = 'pcm-u8'; } else if (codecName === 'in24') { track.info.codec = track.info.pcmLittleEndian ? 'pcm-s24' : 'pcm-s24be'; } else if (codecName === 'in32') { track.info.codec = track.info.pcmLittleEndian ? 'pcm-s32' : 'pcm-s32be'; } else if (codecName === 'fl32') { track.info.codec = track.info.pcmLittleEndian ? 'pcm-f32' : 'pcm-f32be'; } else if (codecName === 'fl64') { track.info.codec = track.info.pcmLittleEndian ? 'pcm-f64' : 'pcm-f64be'; } else if (codecName === 'ipcm') { const pcmSampleSize = track.info.pcmSampleSize; if (track.info.pcmLittleEndian) { if (pcmSampleSize === 16) { track.info.codec = 'pcm-s16'; } else if (pcmSampleSize === 24) { track.info.codec = 'pcm-s24'; } else if (pcmSampleSize === 32) { track.info.codec = 'pcm-s32'; } else { Logging._warn(`Invalid ipcm sample size ${pcmSampleSize}.`); track.info.codec = null; } } else { if (pcmSampleSize === 16) { track.info.codec = 'pcm-s16be'; } else if (pcmSampleSize === 24) { track.info.codec = 'pcm-s24be'; } else if (pcmSampleSize === 32) { track.info.codec = 'pcm-s32be'; } else { Logging._warn(`Invalid ipcm sample size ${pcmSampleSize}.`); track.info.codec = null; } } } else if (codecName === 'fpcm') { const pcmSampleSize = track.info.pcmSampleSize; if (track.info.pcmLittleEndian) { if (pcmSampleSize === 32) { track.info.codec = 'pcm-f32'; } else if (pcmSampleSize === 64) { track.info.codec = 'pcm-f64'; } else { Logging._warn(`Invalid fpcm sample size ${pcmSampleSize}.`); track.info.codec = null; } } else { if (pcmSampleSize === 32) { track.info.codec = 'pcm-f32be'; } else if (pcmSampleSize === 64) { track.info.codec = 'pcm-f64be'; } else { Logging._warn(`Invalid fpcm sample size ${pcmSampleSize}.`); track.info.codec = null; } } } else if (codecName === 'lpcm' && lpcmFlags !== null) { const bytesPerSample = (sampleSize + 7) >> 3; const isFloat = Boolean(lpcmFlags & 1); const isBigEndian = Boolean(lpcmFlags & 2); const sFlags = lpcmFlags & 4 ? -1 : 0; // I guess it means "signed flags" or something? if (sampleSize > 0 && sampleSize <= 64) { if (isFloat) { if (sampleSize === 32) { track.info.codec = isBigEndian ? 'pcm-f32be' : 'pcm-f32'; } } else { if (sFlags & (1 << (bytesPerSample - 1))) { if (bytesPerSample === 1) { track.info.codec = 'pcm-s8'; } else if (bytesPerSample === 2) { track.info.codec = isBigEndian ? 'pcm-s16be' : 'pcm-s16'; } else if (bytesPerSample === 3) { track.info.codec = isBigEndian ? 'pcm-s24be' : 'pcm-s24'; } else if (bytesPerSample === 4) { track.info.codec = isBigEndian ? 'pcm-s32be' : 'pcm-s32'; } } else { if (bytesPerSample === 1) { track.info.codec = 'pcm-u8'; } } } } if (track.info.codec === null) { Logging._warn('Unsupported PCM format.'); } } else if (codecName === null) { Logging._warn(`Unknown encrypted audio codec due to missing frma box.`); } else { Logging._warn(`Unsupported audio codec (sample entry type '${sampleBoxInfo.name}').`); } } slice.filePos = sampleBoxStartPos + sampleBoxInfo.totalSize; } }; break; case 'frma': { const track = this.currentTrack; if (!track) { break; } const format = readAscii(slice, 4); const lowercase = format.toLowerCase(); // Tells us what codec the encrypted track actually uses track.frmaCodecString = lowercase; }; break; case 'schm': { const track = this.currentTrack; if (!track) { break; } slice.skip(4); // Version + flags const schemeType = readAscii(slice, 4); if (schemeType === 'cenc' || schemeType === 'cens' || schemeType === 'cbcs') { track.encryptionInfo = { scheme: schemeType, defaultKid: null, defaultIsProtected: null, defaultPerSampleIvSize: null, defaultConstantIv: null, defaultCryptByteBlock: null, defaultSkipByteBlock: null, }; } else { Logging._warn(`Unsupported encryption scheme '${schemeType}'.`); } }; break; case 'tenc': { const track = this.currentTrack; if (!track || !track.encryptionInfo) { break; } const version = readU8(slice); slice.skip(3); // Flags slice.skip(1); // Reserved const patternByte = readU8(slice); if (version > 0) { track.encryptionInfo.defaultCryptByteBlock = patternByte >> 4; track.encryptionInfo.defaultSkipByteBlock = patternByte & 0xf; } else { track.encryptionInfo.defaultCryptByteBlock = 0; track.encryptionInfo.defaultSkipByteBlock = 0; } track.encryptionInfo.defaultIsProtected = readU8(slice) !== 0; track.encryptionInfo.defaultPerSampleIvSize = readU8(slice); track.encryptionInfo.defaultKid = bytesToHexString(readBytes(slice, 16)); if (track.encryptionInfo.defaultIsProtected && track.encryptionInfo.defaultPerSampleIvSize === 0) { const constantIvSize = readU8(slice); const constantIv = new Uint8Array(16); constantIv.set(readBytes(slice, constantIvSize), 0); track.encryptionInfo.defaultConstantIv = constantIv; } }; break; case 'avcC': { const track = this.currentTrack; if (!track) { break; } assert(track.info); track.info.codecDescription = readBytes(slice, boxInfo.contentSize); }; break; case 'hvcC': { const track = this.currentTrack; if (!track) { break; } assert(track.info); track.info.codecDescription = readBytes(slice, boxInfo.contentSize); }; break; case 'vpcC': { const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'video'); slice.skip(4); // Version + flags const profile = readU8(slice); const level = readU8(slice); const thirdByte = readU8(slice); const bitDepth = thirdByte >> 4; const chromaSubsampling = (thirdByte >> 1) & 0b111; const videoFullRangeFlag = thirdByte & 1; const colourPrimaries = readU8(slice); const transferCharacteristics = readU8(slice); const matrixCoefficients = readU8(slice); track.info.vp9CodecInfo = { profile, level, bitDepth, chromaSubsampling, videoFullRangeFlag, colourPrimaries, transferCharacteristics, matrixCoefficients, }; }; break; case 'av1C': { const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'video'); slice.skip(1); // Marker + version const secondByte = readU8(slice); const profile = secondByte >> 5; const level = secondByte & 0b11111; const thirdByte = readU8(slice); const tier = thirdByte >> 7; const highBitDepth = (thirdByte >> 6) & 1; const twelveBit = (thirdByte >> 5) & 1; const monochrome = (thirdByte >> 4) & 1; const chromaSubsamplingX = (thirdByte >> 3) & 1; const chromaSubsamplingY = (thirdByte >> 2) & 1; const chromaSamplePosition = thirdByte & 0b11; // Logic from https://aomediacodec.github.io/av1-spec/av1-spec.pdf const bitDepth = profile === 2 && highBitDepth ? (twelveBit ? 12 : 10) : (highBitDepth ? 10 : 8); track.info.av1CodecInfo = { profile, level, tier, bitDepth, monochrome, chromaSubsamplingX, chromaSubsamplingY, chromaSamplePosition, }; }; break; case 'colr': { const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'video'); const colourType = readAscii(slice, 4); if (colourType !== 'nclx' && colourType !== 'nclc') { break; } const colourPrimaries = readU16Be(slice); const transferCharacteristics = readU16Be(slice); const matrixCoefficients = readU16Be(slice); let fullRange: boolean | undefined = undefined; if (colourType === 'nclx') { fullRange = Boolean(readU8(slice) & 0x80); } track.info.colorSpace = { primaries: COLOR_PRIMARIES_MAP_INVERSE[colourPrimaries], transfer: TRANSFER_CHARACTERISTICS_MAP_INVERSE[transferCharacteristics], matrix: MATRIX_COEFFICIENTS_MAP_INVERSE[matrixCoefficients], fullRange, } as VideoColorSpaceInit; }; break; case 'pasp': { const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'video'); const num = readU32Be(slice); const den = readU32Be(slice); // https://github.com/Vanilagy/mediabunny/issues/362 if (num > 0 && den > 0) { if (num > den) { track.info.squarePixelWidth = Math.round(track.info.width * num / den); } else { track.info.squarePixelHeight = Math.round(track.info.height * den / num); } } }; break; case 'wave': { this.readContiguousBoxes(slice.slice(contentStartPos, boxInfo.contentSize)); }; break; case 'esds': { const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'audio'); slice.skip(4); // Version + flags const tag = readU8(slice); assert(tag === 0x03); // ES Descriptor readIsomVariableInteger(slice); // Length slice.skip(2); // ES ID const mixed = readU8(slice); const streamDependenceFlag = (mixed & 0x80) !== 0; const urlFlag = (mixed & 0x40) !== 0; const ocrStreamFlag = (mixed & 0x20) !== 0; if (streamDependenceFlag) { slice.skip(2); } if (urlFlag) { const urlLength = readU8(slice); slice.skip(urlLength); } if (ocrStreamFlag) { slice.skip(2); } const decoderConfigTag = readU8(slice); assert(decoderConfigTag === 0x04); // DecoderConfigDescriptor const decoderConfigDescriptorLength = readIsomVariableInteger(slice); // Length const payloadStart = slice.filePos; const objectTypeIndication = readU8(slice); if (objectTypeIndication === 0x40 || objectTypeIndication === 0x67) { track.info.codec = 'aac'; track.info.aacCodecInfo = { isMpeg2: objectTypeIndication === 0x67, objectType: null, }; } else if (objectTypeIndication === 0x69 || objectTypeIndication === 0x6b) { track.info.codec = 'mp3'; } else if (objectTypeIndication === 0xdd) { track.info.codec = 'vorbis'; // "nonstandard, gpac uses it" - FFmpeg } else { Logging._warn( `Unsupported audio codec (objectTypeIndication ${objectTypeIndication}) - discarding track.`, ); } slice.skip(1 + 3 + 4 + 4); if (decoderConfigDescriptorLength > slice.filePos - payloadStart) { // There's a DecoderSpecificInfo at the end, let's read it const decoderSpecificInfoTag = readU8(slice); assert(decoderSpecificInfoTag === 0x05); // DecoderSpecificInfo const decoderSpecificInfoLength = readIsomVariableInteger(slice); track.info.codecDescription = readBytes(slice, decoderSpecificInfoLength); if (track.info.codec === 'aac') { // Let's try to deduce more accurate values directly from the AudioSpecificConfig: const audioSpecificConfig = parseAacAudioSpecificConfig(track.info.codecDescription); if (audioSpecificConfig.numberOfChannels !== null) { track.info.numberOfChannels = audioSpecificConfig.numberOfChannels; } if (audioSpecificConfig.sampleRate !== null) { track.info.sampleRate = audioSpecificConfig.sampleRate; } } } }; break; case 'enda': { const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'audio'); track.info.pcmLittleEndian = !!(readU16Be(slice) & 0xff); // 0xff is from FFmpeg }; break; case 'pcmC': { const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'audio'); slice.skip(1 + 3); // Version + flags // ISO/IEC 23003-5 const formatFlags = readU8(slice); track.info.pcmLittleEndian = Boolean(formatFlags & 0x01); track.info.pcmSampleSize = readU8(slice); }; break; case 'dOps': { // Used for Opus audio const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'audio'); slice.skip(1); // Version // https://www.opus-codec.org/docs/opus_in_isobmff.html const outputChannelCount = readU8(slice); const preSkip = readU16Be(slice); const inputSampleRate = readU32Be(slice); const outputGain = readI16Be(slice); const channelMappingFamily = readU8(slice); let channelMappingTable: Uint8Array; if (channelMappingFamily !== 0) { channelMappingTable = readBytes(slice, 2 + outputChannelCount); } else { channelMappingTable = new Uint8Array(0); } // https://datatracker.ietf.org/doc/html/draft-ietf-codec-oggopus-06 const description = new Uint8Array(8 + 1 + 1 + 2 + 4 + 2 + 1 + channelMappingTable.byteLength); const view = new DataView(description.buffer); view.setUint32(0, 0x4f707573, false); // 'Opus' view.setUint32(4, 0x48656164, false); // 'Head' view.setUint8(8, 1); // Version view.setUint8(9, outputChannelCount); view.setUint16(10, preSkip, true); view.setUint32(12, inputSampleRate, true); view.setInt16(16, outputGain, true); view.setUint8(18, channelMappingFamily); description.set(channelMappingTable, 19); track.info.codecDescription = description; track.info.numberOfChannels = outputChannelCount; // Don't copy the input sample rate, irrelevant, and output sample rate is fixed }; break; case 'dfLa': { // Used for FLAC audio const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'audio'); slice.skip(4); // Version + flags // https://datatracker.ietf.org/doc/rfc9639/ const BLOCK_TYPE_MASK = 0x7f; const LAST_METADATA_BLOCK_FLAG_MASK = 0x80; const startPos = slice.filePos; while (slice.filePos < boxEndPos) { const flagAndType = readU8(slice); const metadataBlockLength = readU24Be(slice); const type = flagAndType & BLOCK_TYPE_MASK; // It's a STREAMINFO block; let's extract the actual sample rate and channel count if (type === FlacBlockType.STREAMINFO) { slice.skip(10); // Extract sample rate and channel count const word = readU32Be(slice); const sampleRate = word >>> 12; const numberOfChannels = ((word >> 9) & 0b111) + 1; track.info.sampleRate = sampleRate; track.info.numberOfChannels = numberOfChannels; slice.skip(20); } else { // Simply skip ahead to the next block slice.skip(metadataBlockLength); } if (flagAndType & LAST_METADATA_BLOCK_FLAG_MASK) { break; } } const endPos = slice.filePos; slice.filePos = startPos; const bytes = readBytes(slice, endPos - startPos); const description = new Uint8Array(4 + bytes.byteLength); const view = new DataView(description.buffer); view.setUint32(0, 0x664c6143, false); // 'fLaC' description.set(bytes, 4); // Set the codec description to be 'fLaC' + all metadata blocks track.info.codecDescription = description; }; break; case 'dac3': { // AC3SpecificBox const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'audio'); const bytes = readBytes(slice, 3); const bitstream = new Bitstream(bytes); const fscod = bitstream.readBits(2); bitstream.skipBits(5 + 3); // Skip bsid and bsmod const acmod = bitstream.readBits(3); const lfeon = bitstream.readBits(1); if (fscod < 3) { track.info.sampleRate = AC3_SAMPLE_RATES[fscod]!; } track.info.numberOfChannels = AC3_ACMOD_CHANNEL_COUNTS[acmod]! + lfeon; }; break; case 'dec3': { // EC3SpecificBox const track = this.currentTrack; if (!track) { break; } assert(track.info?.type === 'audio'); const bytes = readBytes(slice, boxInfo.contentSize); const config = parseEac3Config(bytes); if (!config) { Logging._warn('Invalid dec3 box contents, ignoring.'); break; } const sampleRate = getEac3SampleRate(config); if (sampleRate !== null) { track.info.sampleRate = sampleRate; } track.info.numberOfChannels = getEac3ChannelCount(config); }; break; case 'stts': { const track = this.currentTrack; if (!track) { break; } if (!track.sampleTable) { break; } slice.skip(4); // Version + flags const entryCount = readU32Be(slice); let currentIndex = 0; let currentTimestamp = 0; for (let i = 0; i < entryCount; i++) { const sampleCount = readU32Be(slice); const sampleDelta = readU32Be(slice); track.sampleTable.sampleTimingEntries.push({ startIndex: currentIndex, startDecodeTimestamp: currentTimestamp, count: sampleCount, delta: sampleDelta, }); currentIndex += sampleCount; currentTimestamp += sampleCount * sampleDelta; } }; break; case 'ctts': { const track = this.currentTrack; if (!track) { break; } if (!track.sampleTable) { break; } slice.skip(1 + 3); // Version + flags const entryCount = readU32Be(slice); let sampleIndex = 0; for (let i = 0; i < entryCount; i++) { const sampleCount = readU32Be(slice); const sampleOffset = readI32Be(slice); track.sampleTable.sampleCompositionTimeOffsets.push({ startIndex: sampleIndex, count: sampleCount, offset: sampleOffset, }); sampleIndex += sampleCount; } }; break; case 'stsz': { const track = this.currentTrack; if (!track) { break; } if (!track.sampleTable) { break; } slice.skip(4); // Version + flags const sampleSize = readU32Be(slice); const sampleCount = readU32Be(slice); if (sampleSize === 0) { for (let i = 0; i < sampleCount; i++) { const sampleSize = readU32Be(slice); track.sampleTable.sampleSizes.push(sampleSize); } } else { track.sampleTable.sampleSizes.push(sampleSize); } }; break; case 'stz2': { const track = this.currentTrack; if (!track) { break; } if (!track.sampleTable) { break; } slice.skip(4); // Version + flags slice.skip(3); // Reserved const fieldSize = readU8(slice); // in bits const sampleCount = readU32Be(slice); const bytes = readBytes(slice, Math.ceil(sampleCount * fieldSize / 8)); const bitstream = new Bitstream(bytes); for (let i = 0; i < sampleCount; i++) { const sampleSize = bitstream.readBits(fieldSize); track.sampleTable.sampleSizes.push(sampleSize); } }; break; case 'stss': { const track = this.currentTrack; if (!track) { break; } if (!track.sampleTable) { break; } slice.skip(4); // Version + flags track.sampleTable.keySampleIndices = []; const entryCount = readU32Be(slice); for (let i = 0; i < entryCount; i++) { const sampleIndex = readU32Be(slice) - 1; // Convert to 0-indexed track.sampleTable.keySampleIndices.push(sampleIndex); } if (track.sampleTable.keySampleIndices[0] !== 0) { // Some files don't mark the first sample a key sample, which is basically almost always incorrect. // Here, we correct for that mistake: track.sampleTable.keySampleIndices.unshift(0); } }; break; case 'stsc': { const track = this.currentTrack; if (!track) { break; } if (!track.sampleTable) { break; } slice.skip(4); const entryCount = readU32Be(slice); for (let i = 0; i < entryCount; i++) { const startChunkIndex = readU32Be(slice) - 1; // Convert to 0-indexed const samplesPerChunk = readU32Be(slice); const sampleDescriptionIndex = readU32Be(slice); track.sampleTable.sampleToChunk.push({ startSampleIndex: -1, startChunkIndex, samplesPerChunk, sampleDescriptionIndex, }); } let startSampleIndex = 0; for (let i = 0; i < track.sampleTable.sampleToChunk.length; i++) { track.sampleTable.sampleToChunk[i]!.startSampleIndex = startSampleIndex; if (i < track.sampleTable.sampleToChunk.length - 1) { const nextChunk = track.sampleTable.sampleToChunk[i + 1]!; const chunkCount = nextChunk.startChunkIndex - track.sampleTable.sampleToChunk[i]!.startChunkIndex; startSampleIndex += chunkCount * track.sampleTable.sampleToChunk[i]!.samplesPerChunk; } } }; break; case 'stco': { const track = this.currentTrack; if (!track) { break; } if (!track.sampleTable) { break; } slice.skip(4); // Version + flags const entryCount = readU32Be(slice); for (let i = 0; i < entryCount; i++) { const chunkOffset = readU32Be(slice); track.sampleTable.chunkOffsets.push(chunkOffset); } }; break; case 'co64': { const track = this.currentTrack; if (!track) { break; } if (!track.sampleTable) { break; } slice.skip(4); // Version + flags const entryCount = readU32Be(slice); for (let i = 0; i < entryCount; i++) { const chunkOffset = readU64Be(slice); track.sampleTable.chunkOffsets.push(chunkOffset); } }; break; case 'mvex': { this.isFragmented = true; this.readContiguousBoxes(slice.slice(contentStartPos, boxInfo.contentSize)); }; break; case 'mehd': { const version = readU8(slice); slice.skip(3); // Flags const fragmentDuration = version === 1 ? readU64Be(slice) : readU32Be(slice); this.movieDurationInTimescale = fragmentDuration; }; break; case 'trex': { slice.skip(4); // Version + flags const trackId = readU32Be(slice); const defaultSampleDescriptionIndex = readU32Be(slice); const defaultSampleDuration = readU32Be(slice); const defaultSampleSize = readU32Be(slice); const defaultSampleFlags = readU32Be(slice); // We store these separately rather than in the tracks since the tracks may not exist yet this.fragmentTrackDefaults.push({ trackId, defaultSampleDescriptionIndex, defaultSampleDuration, defaultSampleSize, defaultSampleFlags, }); }; break; case 'tfra': { const version = readU8(slice); slice.skip(3); // Flags const trackId = readU32Be(slice); const track = this.tracks.find(x => x.id === trackId); if (!track) { break; } const word = readU32Be(slice); const lengthSizeOfTrafNum = (word & 0b110000) >> 4; const lengthSizeOfTrunNum = (word & 0b001100) >> 2; const lengthSizeOfSampleNum = word & 0b000011; const functions = [readU8, readU16Be, readU24Be, readU32Be]; const readTrafNum = functions[lengthSizeOfTrafNum]!; const readTrunNum = functions[lengthSizeOfTrunNum]!; const readSampleNum = functions[lengthSizeOfSampleNum]!; const numberOfEntries = readU32Be(slice); for (let i = 0; i < numberOfEntries; i++) { const time = version === 1 ? readU64Be(slice) : readU32Be(slice); const moofOffset = version === 1 ? readU64Be(slice) : readU32Be(slice); readTrafNum(slice); readTrunNum(slice); readSampleNum(slice); track.fragmentLookupTable.push({ timestamp: time, moofOffset, }); } // Sort by timestamp in case it's not naturally sorted track.fragmentLookupTable.sort((a, b) => a.timestamp - b.timestamp); // Remove multiple entries for the same time for (let i = 0; i < track.fragmentLookupTable.length - 1; i++) { const entry1 = track.fragmentLookupTable[i]!; const entry2 = track.fragmentLookupTable[i + 1]!; if (entry1.timestamp === entry2.timestamp) { track.fragmentLookupTable.splice(i + 1, 1); i--; } } }; break; case 'moof': { this.currentFragment = { moofOffset: startPos, moofSize: boxInfo.totalSize, implicitBaseDataOffset: startPos, trackData: new Map(), psshBoxes: [], }; this.readContiguousBoxes(slice.slice(contentStartPos, boxInfo.contentSize)); this.lastReadFragment = this.currentFragment; this.currentFragment = null; }; break; case 'traf': { assert(this.currentFragment); this.readContiguousBoxes(slice.slice(contentStartPos, boxInfo.contentSize)); // It is possible that there is no current track, for example when we don't care about the track // referenced in the track fragment header. if (this.currentTrack) { const trackData = this.currentFragment.trackData.get(this.currentTrack.id); cond: if (trackData) { if (trackData.samples.length === 0) { // Don't associate the fragment with the track if it has no samples, this simplifies // other code this.currentFragment.trackData.delete(this.currentTrack.id); break cond; } trackData.presentationTimestamps = trackData.samples .map((x, i) => ({ presentationTimestamp: x.presentationTimestamp, sampleIndex: i })) .sort((a, b) => a.presentationTimestamp - b.presentationTimestamp); for (let i = 0; i < trackData.presentationTimestamps.length; i++) { const currentEntry = trackData.presentationTimestamps[i]!; const currentSample = trackData.samples[currentEntry.sampleIndex]!; if (trackData.firstKeyFrameTimestamp === null && currentSample.isKeyFrame) { trackData.firstKeyFrameTimestamp = currentSample.presentationTimestamp; } if (i < trackData.presentationTimestamps.length - 1) { // Update sample durations based on presentation order const nextEntry = trackData.presentationTimestamps[i + 1]!; const duration = nextEntry.presentationTimestamp - currentEntry.presentationTimestamp; currentSample.duration = duration; } } const firstSample = trackData.samples[trackData.presentationTimestamps[0]!.sampleIndex]!; const lastSample = trackData.samples[last(trackData.presentationTimestamps)!.sampleIndex]!; trackData.startTimestamp = firstSample.presentationTimestamp; trackData.endTimestamp = lastSample.presentationTimestamp + lastSample.duration; const { currentFragmentState } = this.currentTrack; assert(currentFragmentState); if (currentFragmentState.startTimestamp !== null) { offsetFragmentTrackDataByTimestamp(trackData, currentFragmentState.startTimestamp); trackData.startTimestampIsFinal = true; } // Transfer the buffered saiz+saio state onto the track data, so readFragment can resolve it // once all boxes are parsed. Only relevant if senc wasn't already used to populate samples. if (currentFragmentState.encryptionAuxInfo && !trackData.samples[0]!.encryption) { trackData.encryptionAuxInfo = currentFragmentState.encryptionAuxInfo; } } this.currentTrack.currentFragmentState = null; this.currentTrack = null; } }; break; case 'pssh': { if (this.input._formatOptions.isobmff?._suppressPsshParsing) { break; } const psshBox = parsePsshBoxContents(readBytes(slice, boxInfo.contentSize)); if (this.currentFragment) { this.currentFragment.psshBoxes.push(psshBox); } else if (!this.currentTrack) { this.psshBoxes.push(psshBox); } }; break; case 'tfhd': { assert(this.currentFragment); slice.skip(1); // Version const flags = readU24Be(slice); const baseDataOffsetPresent = Boolean(flags & 0x000001); const sampleDescriptionIndexPresent = Boolean(flags & 0x000002); const defaultSampleDurationPresent = Boolean(flags & 0x000008); const defaultSampleSizePresent = Boolean(flags & 0x000010); const defaultSampleFlagsPresent = Boolean(flags & 0x000020); const durationIsEmpty = Boolean(flags & 0x010000); const defaultBaseIsMoof = Boolean(flags & 0x020000); const trackId = readU32Be(slice); const track = this.tracks.find(x => x.id === trackId); if (!track) { // We don't care about this track break; } const defaults = this.fragmentTrackDefaults.find(x => x.trackId === trackId); this.currentTrack = track; track.currentFragmentState = { baseDataOffset: this.currentFragment.implicitBaseDataOffset, sampleDescriptionIndex: defaults?.defaultSampleDescriptionIndex ?? null, defaultSampleDuration: defaults?.defaultSampleDuration ?? null, defaultSampleSize: defaults?.defaultSampleSize ?? null, defaultSampleFlags: defaults?.defaultSampleFlags ?? null, startTimestamp: null, encryptionAuxInfo: null, }; if (baseDataOffsetPresent) { track.currentFragmentState.baseDataOffset = readU64Be(slice); } else if (defaultBaseIsMoof) { track.currentFragmentState.baseDataOffset = this.currentFragment.moofOffset; } if (sampleDescriptionIndexPresent) { track.currentFragmentState.sampleDescriptionIndex = readU32Be(slice); } if (defaultSampleDurationPresent) { track.currentFragmentState.defaultSampleDuration = readU32Be(slice); } if (defaultSampleSizePresent) { track.currentFragmentState.defaultSampleSize = readU32Be(slice); } if (defaultSampleFlagsPresent) { track.currentFragmentState.defaultSampleFlags = readU32Be(slice); } if (durationIsEmpty) { track.currentFragmentState.defaultSampleDuration = 0; } }; break; case 'tfdt': { const track = this.currentTrack; if (!track) { break; } assert(track.currentFragmentState); const version = readU8(slice); slice.skip(3); // Flags const baseMediaDecodeTime = version === 0 ? readU32Be(slice) : readU64Be(slice); track.currentFragmentState.startTimestamp = baseMediaDecodeTime; }; break; case 'trun': { const track = this.currentTrack; if (!track) { break; } assert(this.currentFragment); assert(track.currentFragmentState); const version = readU8(slice); const flags = readU24Be(slice); const dataOffsetPresent = Boolean(flags & 0x000001); const firstSampleFlagsPresent = Boolean(flags & 0x000004); const sampleDurationPresent = Boolean(flags & 0x000100); const sampleSizePresent = Boolean(flags & 0x000200); const sampleFlagsPresent = Boolean(flags & 0x000400); const sampleCompositionTimeOffsetsPresent = Boolean(flags & 0x000800); const sampleCount = readU32Be(slice); let dataOffset: number | null = null; if (dataOffsetPresent) { dataOffset = readI32Be(slice); } let firstSampleFlags: number | null = null; if (firstSampleFlagsPresent) { firstSampleFlags = readU32Be(slice); } let trackData: FragmentTrackData; if (this.currentFragment.trackData.has(track.id)) { trackData = this.currentFragment.trackData.get(track.id)!; if (dataOffset !== null) { trackData.currentOffset = track.currentFragmentState.baseDataOffset + dataOffset; } else { // "If the data-offset is not present, then the data for this run starts immediately after the // data of the previous run" } } else { trackData = { track, currentTimestamp: 0, currentOffset: track.currentFragmentState.baseDataOffset + (dataOffset ?? 0), startTimestamp: 0, endTimestamp: 0, firstKeyFrameTimestamp: null, samples: [], presentationTimestamps: [], startTimestampIsFinal: false, encryptionAuxInfo: null, }; this.currentFragment.trackData.set(track.id, trackData); } for (let i = 0; i < sampleCount; i++) { let sampleDuration: number; if (sampleDurationPresent) { sampleDuration = readU32Be(slice); } else { assert(track.currentFragmentState.defaultSampleDuration !== null); sampleDuration = track.currentFragmentState.defaultSampleDuration; } let sampleSize: number; if (sampleSizePresent) { sampleSize = readU32Be(slice); } else { assert(track.currentFragmentState.defaultSampleSize !== null); sampleSize = track.currentFragmentState.defaultSampleSize; } let sampleFlags: number; if (sampleFlagsPresent) { sampleFlags = readU32Be(slice); } else { assert(track.currentFragmentState.defaultSampleFlags !== null); sampleFlags = track.currentFragmentState.defaultSampleFlags; } if (i === 0 && firstSampleFlags !== null) { sampleFlags = firstSampleFlags; } let sampleCompositionTimeOffset = 0; if (sampleCompositionTimeOffsetsPresent) { if (version === 0) { sampleCompositionTimeOffset = readU32Be(slice); } else { sampleCompositionTimeOffset = readI32Be(slice); } } const isKeyFrame = !(sampleFlags & 0x00010000); trackData.samples.push({ presentationTimestamp: trackData.currentTimestamp + sampleCompositionTimeOffset, duration: sampleDuration, byteOffset: trackData.currentOffset, byteSize: sampleSize, isKeyFrame, encryption: null, }); trackData.currentOffset += sampleSize; trackData.currentTimestamp += sampleDuration; } this.currentFragment.implicitBaseDataOffset = trackData.currentOffset; }; break; case 'saiz': { // Sample Auxiliary Information Sizes - per-sample sizes of (typically) the encryption aux info. const track = this.currentTrack; if (!track || !track.encryptionInfo) { break; } slice.skip(1); // Version const flags = readU24Be(slice); if (flags & 0x01) { const auxInfoType = readAscii(slice, 4); const auxInfoTypeParam = readU32Be(slice); if (auxInfoType !== track.encryptionInfo.scheme || auxInfoTypeParam !== 0) { // Not the encryption aux info break; } } const defaultSampleInfoSize = readU8(slice); const sampleCount = readU32Be(slice); let sampleSizes: Uint8Array | null = null; if (defaultSampleInfoSize === 0 && sampleCount > 0) { sampleSizes = readBytes(slice, sampleCount); } const aux = getOrCreateEncryptionAuxInfo(track); aux.defaultSampleInfoSize = defaultSampleInfoSize; aux.sampleSizes = sampleSizes; aux.sampleCount = sampleCount; }; break; case 'saio': { // Sample Auxiliary Information Offsets - file offset(s) where the aux info lives. const track = this.currentTrack; if (!track || !track.encryptionInfo) { break; } const version = readU8(slice); const flags = readU24Be(slice); if (flags & 0x01) { const auxInfoType = readAscii(slice, 4); const auxInfoTypeParam = readU32Be(slice); if (auxInfoType !== track.encryptionInfo.scheme || auxInfoTypeParam !== 0) { break; } } const entryCount = readU32Be(slice); if (entryCount === 0) { break; } if (entryCount > 1) { Logging._warn('Multiple saio entries are not supported; using the first offset only.'); } let offset = version === 0 ? readU32Be(slice) : Number(readU64Be(slice)); // Per ISO/IEC 23001-7: when saio is inside a moof, offsets are relative to the start of the moof box. if (this.currentFragment) { offset += this.currentFragment.moofOffset; } const aux = getOrCreateEncryptionAuxInfo(track); aux.offset = offset; }; break; case 'senc': { // Per-sample encryption info inside a 'traf'. Holds per-sample IV and optional subsample breakdown // for CENC-protected samples const track = this.currentTrack; if (!track || !track.encryptionInfo) { break; } assert(this.currentFragment); const trackData = this.currentFragment.trackData.get(track.id); if (!trackData) { break; } slice.skip(1); // Version const flags = readU24Be(slice); const useSubsamples = Boolean(flags & 0x000002); const sampleCount = readU32Be(slice); const ivSize = track.encryptionInfo.defaultPerSampleIvSize; assert(ivSize !== null); for (let i = 0; i < Math.min(sampleCount, trackData.samples.length); i++) { // Normalize the IV to 16 bytes so downstream code can assume a full-length buffer. For CTR with // an 8-byte per-sample IV the lower 8 bytes are zero (that's the CENC spec's block counter start); // for CBC/cbcs the IV is always 16 bytes by spec. const iv = new Uint8Array(16); if (ivSize > 0) { iv.set(readBytes(slice, ivSize), 0); } else { iv.set(track.encryptionInfo.defaultConstantIv!, 0); } let subsamples: SampleEncryptionInfo['subsamples'] = null; if (useSubsamples) { const subsampleCount = readU16Be(slice); subsamples = []; for (let j = 0; j < subsampleCount; j++) { const clearLen = readU16Be(slice); const protectedLen = readU32Be(slice); subsamples.push({ clearLen, protectedLen }); } } const sample = trackData.samples[i]!; sample.encryption = { iv, subsamples }; } }; break; // Metadata section // https://exiftool.org/TagNames/QuickTime.html // https://mp4workshop.com/about case 'udta': { // Contains either movie metadata or track metadata const iterator = this.iterateContiguousBoxes(slice.slice(contentStartPos, boxInfo.contentSize)); for (const { boxInfo, slice } of iterator) { if (boxInfo.name !== 'meta' && !this.currentTrack) { const startPos = slice.filePos; this.metadataTags.raw ??= {}; if (boxInfo.name[0] === '©') { // https://mp4workshop.com/about // Box name starting with © indicates "international text" this.metadataTags.raw[boxInfo.name] ??= readMetadataStringShort(slice); } else { this.metadataTags.raw[boxInfo.name] ??= readBytes(slice, boxInfo.contentSize); } slice.filePos = startPos; } switch (boxInfo.name) { case 'meta': { slice.skip(-boxInfo.headerSize); this.traverseBox(slice); }; break; case '©nam': case 'name': { if (this.currentTrack) { this.currentTrack.name = textDecoder.decode(readBytes(slice, boxInfo.contentSize)); } else { this.metadataTags.title ??= readMetadataStringShort(slice); } }; break; case '©des': { if (!this.currentTrack) { this.metadataTags.description ??= readMetadataStringShort(slice); } }; break; case '©ART': { if (!this.currentTrack) { this.metadataTags.artist ??= readMetadataStringShort(slice); } }; break; case '©alb': { if (!this.currentTrack) { this.metadataTags.album ??= readMetadataStringShort(slice); } }; break; case 'albr': { if (!this.currentTrack) { this.metadataTags.albumArtist ??= readMetadataStringShort(slice); } }; break; case '©gen': { if (!this.currentTrack) { this.metadataTags.genre ??= readMetadataStringShort(slice); } }; break; case '©day': { if (!this.currentTrack) { const date = new Date(readMetadataStringShort(slice)); if (!Number.isNaN(date.getTime())) { this.metadataTags.date ??= date; } } }; break; case '©cmt': { if (!this.currentTrack) { this.metadataTags.comment ??= readMetadataStringShort(slice); } }; break; case '©lyr': { if (!this.currentTrack) { this.metadataTags.lyrics ??= readMetadataStringShort(slice); } }; break; } } }; break; case 'meta': { if (this.currentTrack) { break; // Only care about movie-level metadata for now } // The 'meta' box comes in two flavors, one with flags/version and one without. To know which is which, // let's read the next 4 bytes, which are either the version or the size of the first subbox. const word = readU32Be(slice); const isQuickTime = word !== 0; this.currentMetadataKeys = new Map(); if (isQuickTime) { this.readContiguousBoxes(slice.slice(contentStartPos, boxInfo.contentSize)); } else { this.readContiguousBoxes(slice.slice(contentStartPos + 4, boxInfo.contentSize - 4)); } this.currentMetadataKeys = null; }; break; case 'keys': { if (!this.currentMetadataKeys) { break; } slice.skip(4); // Version + flags const entryCount = readU32Be(slice); for (let i = 0; i < entryCount; i++) { const keySize = readU32Be(slice); slice.skip(4); // Key namespace const keyName = textDecoder.decode(readBytes(slice, keySize - 8)); this.currentMetadataKeys.set(i + 1, keyName); } }; break; case 'ilst': { if (!this.currentMetadataKeys) { break; } const iterator = this.iterateContiguousBoxes(slice.slice(contentStartPos, boxInfo.contentSize)); for (const { boxInfo, slice } of iterator) { let metadataKey = boxInfo.name; // Interpret the box name as a u32be const nameAsNumber = (metadataKey.charCodeAt(0) << 24) + (metadataKey.charCodeAt(1) << 16) + (metadataKey.charCodeAt(2) << 8) + metadataKey.charCodeAt(3); if (this.currentMetadataKeys.has(nameAsNumber)) { // An entry exists for this number metadataKey = this.currentMetadataKeys.get(nameAsNumber)!; } const data = readDataBox(slice); this.metadataTags.raw ??= {}; this.metadataTags.raw[metadataKey] ??= data; switch (metadataKey) { case '©nam': case 'titl': case 'com.apple.quicktime.title': case 'title': { if (typeof data === 'string') { this.metadataTags.title ??= data; } }; break; case '©des': case 'desc': case 'dscp': case 'com.apple.quicktime.description': case 'description': { if (typeof data === 'string') { this.metadataTags.description ??= data; } }; break; case '©ART': case 'com.apple.quicktime.artist': case 'artist': { if (typeof data === 'string') { this.metadataTags.artist ??= data; } }; break; case '©alb': case 'albm': case 'com.apple.quicktime.album': case 'album': { if (typeof data === 'string') { this.metadataTags.album ??= data; } }; break; case 'aART': case 'album_artist': { if (typeof data === 'string') { this.metadataTags.albumArtist ??= data; } }; break; case '©cmt': case 'com.apple.quicktime.comment': case 'comment': { if (typeof data === 'string') { this.metadataTags.comment ??= data; } }; break; case '©gen': case 'gnre': case 'com.apple.quicktime.genre': case 'genre': { if (typeof data === 'string') { this.metadataTags.genre ??= data; } }; break; case '©lyr': case 'lyrics': { if (typeof data === 'string') { this.metadataTags.lyrics ??= data; } }; break; case '©day': case 'rldt': case 'com.apple.quicktime.creationdate': case 'date': { if (typeof data === 'string') { const date = new Date(data); if (!Number.isNaN(date.getTime())) { this.metadataTags.date ??= date; } } }; break; case 'covr': case 'com.apple.quicktime.artwork': { if (data instanceof RichImageData) { this.metadataTags.images ??= []; this.metadataTags.images.push({ data: data.data, kind: 'coverFront', mimeType: data.mimeType, }); } else if (data instanceof Uint8Array) { this.metadataTags.images ??= []; this.metadataTags.images.push({ data, kind: 'coverFront', mimeType: 'image/*', }); } }; break; case 'track': { if (typeof data === 'string') { const parts = data.split('/'); const trackNum = Number.parseInt(parts[0]!, 10); const tracksTotal = parts[1] && Number.parseInt(parts[1], 10); if (Number.isInteger(trackNum) && trackNum > 0) { this.metadataTags.trackNumber ??= trackNum; } if (tracksTotal && Number.isInteger(tracksTotal) && tracksTotal > 0) { this.metadataTags.tracksTotal ??= tracksTotal; } } }; break; case 'trkn': { if (data instanceof Uint8Array && data.length >= 6) { const view = toDataView(data); const trackNumber = view.getUint16(2, false); const tracksTotal = view.getUint16(4, false); if (trackNumber > 0) { this.metadataTags.trackNumber ??= trackNumber; } if (tracksTotal > 0) { this.metadataTags.tracksTotal ??= tracksTotal; } } }; break; case 'disc': case 'disk': { if (data instanceof Uint8Array && data.length >= 6) { const view = toDataView(data); const discNumber = view.getUint16(2, false); const discNumberMax = view.getUint16(4, false); if (discNumber > 0) { this.metadataTags.discNumber ??= discNumber; } if (discNumberMax > 0) { this.metadataTags.discsTotal ??= discNumberMax; } } }; break; } } }; break; } slice.filePos = boxEndPos; return true; } } abstract class IsobmffTrackBacking implements InputTrackBacking { packetToSampleIndex = new WeakMap(); packetToFragmentLocation = new WeakMap(); constructor(public internalTrack: InternalTrack) {} abstract getType(): TrackType; abstract getDecoderConfig(): Promise; getId() { return this.internalTrack.id; } getNumber() { const demuxer = this.internalTrack.demuxer; const trackType = this.internalTrack.trackBacking!.getType(); let number = 0; for (const track of demuxer.tracks) { if (track.trackBacking!.getType() === trackType) { number++; } if (track === this.internalTrack) { break; } } return number; } getCodec(): MediaCodec | null { throw new Error('Not implemented on base class.'); } getInternalCodecId() { return this.internalTrack.internalCodecId; } getName() { return this.internalTrack.name; } getLanguageCode() { return this.internalTrack.languageCode; } getTimeResolution() { return this.internalTrack.timescale; } isRelativeToUnixEpoch() { return false; } getUnixTimeForTimestamp() { return null; } getDisposition() { return this.internalTrack.disposition; } getPairingMask() { return 1n; } getBitrate() { return null; } getAverageBitrate() { return null; } async getDurationFromMetadata() { const track = this.internalTrack; if (track.durationInMediaTimescale <= 0) { // The duration is often zero for fragmented files for example; return `null` to signal that the duration // must be computed instead. return null; } assert(track.trackBacking); const firstPacket = await track.trackBacking.getFirstPacket({ metadataOnly: true }); return (firstPacket?.timestamp ?? 0) + track.durationInMediaTimescale / track.timescale; } async getLiveRefreshInterval() { return null; } async getFirstPacket(options: PacketRetrievalOptions) { const regularPacket = await this.fetchPacketForSampleIndex(0, options); if (regularPacket || !this.internalTrack.demuxer.isFragmented) { // If there's a non-fragmented packet, always prefer that return regularPacket; } return this.performFragmentedLookup( null, (fragment) => { const trackData = fragment.trackData.get(this.internalTrack.id); if (trackData) { return { sampleIndex: 0, correctSampleFound: true, }; } return { sampleIndex: -1, correctSampleFound: false, }; }, -Infinity, // Use -Infinity as a search timestamp to avoid using the lookup entries Infinity, options, ); } private mapTimestampIntoTimescale(timestamp: number) { // Do a little rounding to catch cases where the result is very close to an integer. If it is, it's likely // that the number was originally an integer divided by the timescale. For stability, it's best // to return the integer in this case. return roundIfAlmostInteger(timestamp * this.internalTrack.timescale) + this.internalTrack.editListOffset; } async getPacket(timestamp: number, options: PacketRetrievalOptions) { const timestampInTimescale = this.mapTimestampIntoTimescale(timestamp); const sampleTable = this.internalTrack.demuxer.getSampleTableForTrack(this.internalTrack); const sampleIndex = getSampleIndexForTimestamp(sampleTable, timestampInTimescale); const regularPacket = await this.fetchPacketForSampleIndex(sampleIndex, options); if (!sampleTableIsEmpty(sampleTable) || !this.internalTrack.demuxer.isFragmented) { // Prefer the non-fragmented packet return regularPacket; } return this.performFragmentedLookup( null, (fragment) => { const trackData = fragment.trackData.get(this.internalTrack.id); if (!trackData) { return { sampleIndex: -1, correctSampleFound: false }; } const index = binarySearchLessOrEqual( trackData.presentationTimestamps, timestampInTimescale, x => x.presentationTimestamp, ); const sampleIndex = index !== -1 ? trackData.presentationTimestamps[index]!.sampleIndex : -1; const correctSampleFound = index !== -1 && timestampInTimescale < trackData.endTimestamp; return { sampleIndex, correctSampleFound }; }, timestampInTimescale, timestampInTimescale, options, ); } async getNextPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { const regularSampleIndex = this.packetToSampleIndex.get(packet); if (regularSampleIndex !== undefined) { // Prefer the non-fragmented packet return this.fetchPacketForSampleIndex(regularSampleIndex + 1, options); } const locationInFragment = this.packetToFragmentLocation.get(packet); if (locationInFragment === undefined) { throw new Error('Packet was not created from this track.'); } return this.performFragmentedLookup( locationInFragment.fragment, (fragment) => { if (fragment === locationInFragment.fragment) { const trackData = fragment.trackData.get(this.internalTrack.id)!; if (locationInFragment.sampleIndex + 1 < trackData.samples.length) { // We can simply take the next sample in the fragment return { sampleIndex: locationInFragment.sampleIndex + 1, correctSampleFound: true, }; } } else { const trackData = fragment.trackData.get(this.internalTrack.id); if (trackData) { return { sampleIndex: 0, correctSampleFound: true, }; } } return { sampleIndex: -1, correctSampleFound: false, }; }, -Infinity, // Use -Infinity as a search timestamp to avoid using the lookup entries Infinity, options, ); } async getKeyPacket(timestamp: number, options: PacketRetrievalOptions) { const timestampInTimescale = this.mapTimestampIntoTimescale(timestamp); const sampleTable = this.internalTrack.demuxer.getSampleTableForTrack(this.internalTrack); const sampleIndex = getKeyframeSampleIndexForTimestamp(sampleTable, timestampInTimescale); const regularPacket = await this.fetchPacketForSampleIndex(sampleIndex, options); if (!sampleTableIsEmpty(sampleTable) || !this.internalTrack.demuxer.isFragmented) { // Prefer the non-fragmented packet return regularPacket; } return this.performFragmentedLookup( null, (fragment) => { const trackData = fragment.trackData.get(this.internalTrack.id); if (!trackData) { return { sampleIndex: -1, correctSampleFound: false }; } const index = findLastIndex(trackData.presentationTimestamps, (x) => { const sample = trackData.samples[x.sampleIndex]!; return sample.isKeyFrame && x.presentationTimestamp <= timestampInTimescale; }); const sampleIndex = index !== -1 ? trackData.presentationTimestamps[index]!.sampleIndex : -1; const correctSampleFound = index !== -1 && timestampInTimescale < trackData.endTimestamp; return { sampleIndex, correctSampleFound }; }, timestampInTimescale, timestampInTimescale, options, ); } async getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { const regularSampleIndex = this.packetToSampleIndex.get(packet); if (regularSampleIndex !== undefined) { // Prefer the non-fragmented packet const sampleTable = this.internalTrack.demuxer.getSampleTableForTrack(this.internalTrack); const nextKeyFrameSampleIndex = getNextKeyframeIndexForSample(sampleTable, regularSampleIndex); return this.fetchPacketForSampleIndex(nextKeyFrameSampleIndex, options); } const locationInFragment = this.packetToFragmentLocation.get(packet); if (locationInFragment === undefined) { throw new Error('Packet was not created from this track.'); } return this.performFragmentedLookup( locationInFragment.fragment, (fragment) => { if (fragment === locationInFragment.fragment) { const trackData = fragment.trackData.get(this.internalTrack.id)!; const nextKeyFrameIndex = trackData.samples.findIndex( (x, i) => x.isKeyFrame && i > locationInFragment.sampleIndex, ); if (nextKeyFrameIndex !== -1) { // We can simply take the next key frame in the fragment return { sampleIndex: nextKeyFrameIndex, correctSampleFound: true, }; } } else { const trackData = fragment.trackData.get(this.internalTrack.id); if (trackData && trackData.firstKeyFrameTimestamp !== null) { const keyFrameIndex = trackData.samples.findIndex(x => x.isKeyFrame); assert(keyFrameIndex !== -1); // There must be one return { sampleIndex: keyFrameIndex, correctSampleFound: true, }; } } return { sampleIndex: -1, correctSampleFound: false, }; }, -Infinity, // Use -Infinity as a search timestamp to avoid using the lookup entries Infinity, options, ); } private async fetchPacketForSampleIndex(sampleIndex: number, options: PacketRetrievalOptions) { if (sampleIndex === -1) { return null; } const sampleTable = this.internalTrack.demuxer.getSampleTableForTrack(this.internalTrack); const sampleInfo = getSampleInfo(sampleTable, sampleIndex); if (!sampleInfo) { return null; } let data: Uint8Array; if (options.metadataOnly) { data = PLACEHOLDER_DATA; } else { let slice = this.internalTrack.demuxer.reader.requestSlice( sampleInfo.sampleOffset, sampleInfo.sampleSize, ); if (slice instanceof Promise) slice = await slice; if (!slice) { return null; // Data is outside } data = readBytes(slice, sampleInfo.sampleSize); if (this.internalTrack.encryptionAuxInfo) { assert(this.internalTrack.encryptionInfo); const entries = await resolveEncryptionAuxInfo( this.internalTrack.demuxer.reader, this.internalTrack.encryptionInfo, this.internalTrack.encryptionAuxInfo, ); if (sampleIndex < entries.length) { data = await decryptSample(this.internalTrack, entries[sampleIndex]!, data, null); } } } const timestamp = (sampleInfo.presentationTimestamp - this.internalTrack.editListOffset) / this.internalTrack.timescale; const duration = sampleInfo.duration / this.internalTrack.timescale; const packet = new EncodedPacket( data, sampleInfo.isKeyFrame ? 'key' : 'delta', timestamp, duration, sampleIndex, sampleInfo.sampleSize, ); this.packetToSampleIndex.set(packet, sampleIndex); return packet; } private async fetchPacketInFragment(fragment: Fragment, sampleIndex: number, options: PacketRetrievalOptions) { if (sampleIndex === -1) { return null; } const trackData = fragment.trackData.get(this.internalTrack.id)!; const fragmentSample = trackData.samples[sampleIndex]; assert(fragmentSample); let data: Uint8Array; if (options.metadataOnly) { data = PLACEHOLDER_DATA; } else { let slice = this.internalTrack.demuxer.reader.requestSlice( fragmentSample.byteOffset, fragmentSample.byteSize, ); if (slice instanceof Promise) slice = await slice; if (!slice) { return null; // Data is outside } data = readBytes(slice, fragmentSample.byteSize); if (fragmentSample.encryption) { data = await decryptSample(this.internalTrack, fragmentSample.encryption, data, fragment); } } const timestamp = (fragmentSample.presentationTimestamp - this.internalTrack.editListOffset) / this.internalTrack.timescale; const duration = fragmentSample.duration / this.internalTrack.timescale; const packet = new EncodedPacket( data, fragmentSample.isKeyFrame ? 'key' : 'delta', timestamp, duration, fragment.moofOffset + sampleIndex, fragmentSample.byteSize, ); this.packetToFragmentLocation.set(packet, { fragment, sampleIndex }); return packet; } /** Looks for a packet in the fragments while trying to load as few fragments as possible to retrieve it. */ private async performFragmentedLookup( // The fragment where we start looking startFragment: Fragment | null, // This function returns the best-matching sample in a given fragment getMatchInFragment: (fragment: Fragment) => { sampleIndex: number; correctSampleFound: boolean }, // The timestamp with which we can search the lookup table searchTimestamp: number, // The timestamp for which we know the correct sample will not come after it latestTimestamp: number, options: PacketRetrievalOptions, ): Promise { const demuxer = this.internalTrack.demuxer; let currentFragment: Fragment | null = null; let bestFragment: Fragment | null = null; let bestSampleIndex = -1; if (startFragment) { const { sampleIndex, correctSampleFound } = getMatchInFragment(startFragment); if (correctSampleFound) { return this.fetchPacketInFragment(startFragment, sampleIndex, options); } if (sampleIndex !== -1) { bestFragment = startFragment; bestSampleIndex = sampleIndex; } } // Search for a lookup entry; this way, we won't need to start searching from the start of the file // but can jump right into the correct fragment (or at least nearby). const lookupEntryIndex = binarySearchLessOrEqual( this.internalTrack.fragmentLookupTable, searchTimestamp, x => x.timestamp, ); const lookupEntry = lookupEntryIndex !== -1 ? this.internalTrack.fragmentLookupTable[lookupEntryIndex]! : null; const positionCacheIndex = binarySearchLessOrEqual( this.internalTrack.fragmentPositionCache, searchTimestamp, x => x.startTimestamp, ); const positionCacheEntry = positionCacheIndex !== -1 ? this.internalTrack.fragmentPositionCache[positionCacheIndex]! : null; const lookupEntryPosition = Math.max( lookupEntry?.moofOffset ?? 0, positionCacheEntry?.moofOffset ?? 0, ) || null; let currentPos: number; if (!startFragment) { currentPos = lookupEntryPosition ?? 0; } else { if (lookupEntryPosition === null || startFragment.moofOffset >= lookupEntryPosition) { currentPos = startFragment.moofOffset + startFragment.moofSize; currentFragment = startFragment; } else { // Use the lookup entry currentPos = lookupEntryPosition; } } while (true) { if (currentFragment) { const trackData = currentFragment.trackData.get(this.internalTrack.id); if (trackData && trackData.startTimestamp > latestTimestamp) { // We're already past the upper bound, no need to keep searching break; } } // Load the header let slice = demuxer.reader.requestSliceRange(currentPos, MIN_BOX_HEADER_SIZE, MAX_BOX_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const boxStartPos = currentPos; const boxInfo = readBoxHeader(slice); if (!boxInfo) { break; } if (boxInfo.name === 'moof') { currentFragment = await demuxer.readFragment(boxStartPos); const { sampleIndex, correctSampleFound } = getMatchInFragment(currentFragment); if (correctSampleFound) { return this.fetchPacketInFragment(currentFragment, sampleIndex, options); } if (sampleIndex !== -1) { bestFragment = currentFragment; bestSampleIndex = sampleIndex; } } currentPos = boxStartPos + boxInfo.totalSize; } // Catch faulty lookup table entries if (lookupEntry && (!bestFragment || bestFragment.moofOffset < lookupEntry.moofOffset)) { // The lookup table entry lied to us! We found a lookup entry but no fragment there that satisfied // the match. In this case, let's search again but using the lookup entry before that. const previousLookupEntry = this.internalTrack.fragmentLookupTable[lookupEntryIndex - 1]; assert(!previousLookupEntry || previousLookupEntry.timestamp < lookupEntry.timestamp); const newSearchTimestamp = previousLookupEntry?.timestamp ?? -Infinity; return this.performFragmentedLookup( null, getMatchInFragment, newSearchTimestamp, latestTimestamp, options, ); } if (bestFragment) { // If we finished looping but didn't find a perfect match, still return the best match we found return this.fetchPacketInFragment(bestFragment, bestSampleIndex, options); } return null; } } class IsobmffVideoTrackBacking extends IsobmffTrackBacking implements InputVideoTrackBacking { override internalTrack: InternalVideoTrack; decoderConfigPromise: Promise | null = null; constructor(internalTrack: InternalVideoTrack) { super(internalTrack); this.internalTrack = internalTrack; } getType() { return 'video' as const; } override getCodec(): VideoCodec | null { return this.internalTrack.info.codec; } getCodedWidth() { return this.internalTrack.info.width; } getCodedHeight() { return this.internalTrack.info.height; } getSquarePixelWidth() { return this.internalTrack.info.squarePixelWidth; } getSquarePixelHeight() { return this.internalTrack.info.squarePixelHeight; } getRotation() { return this.internalTrack.rotation; } async getColorSpace(): Promise { return { primaries: this.internalTrack.info.colorSpace?.primaries, transfer: this.internalTrack.info.colorSpace?.transfer, matrix: this.internalTrack.info.colorSpace?.matrix, fullRange: this.internalTrack.info.colorSpace?.fullRange, }; } async canBeTransparent() { return false; } async getDecoderConfig(): Promise { if (!this.internalTrack.info.codec) { return null; } return this.decoderConfigPromise ??= (async (): Promise => { if (this.internalTrack.info.codec === 'vp9' && !this.internalTrack.info.vp9CodecInfo) { const firstPacket = await this.getFirstPacket({}); this.internalTrack.info.vp9CodecInfo = firstPacket && extractVp9CodecInfoFromPacket(firstPacket.data); } else if (this.internalTrack.info.codec === 'av1' && !this.internalTrack.info.av1CodecInfo) { const firstPacket = await this.getFirstPacket({}); this.internalTrack.info.av1CodecInfo = firstPacket && extractAv1CodecInfoFromPacket(firstPacket.data); } const config: VideoDecoderConfig = { codec: extractVideoCodecString(this.internalTrack.info), codedWidth: this.internalTrack.info.width, codedHeight: this.internalTrack.info.height, description: this.internalTrack.info.codecDescription ?? undefined, colorSpace: this.internalTrack.info.colorSpace ?? undefined, }; if ( this.internalTrack.info.width !== this.internalTrack.info.squarePixelWidth || this.internalTrack.info.height !== this.internalTrack.info.squarePixelHeight ) { config.displayAspectWidth = this.internalTrack.info.squarePixelWidth; config.displayAspectHeight = this.internalTrack.info.squarePixelHeight; } return config; })(); } } class IsobmffAudioTrackBacking extends IsobmffTrackBacking implements InputAudioTrackBacking { override internalTrack: InternalAudioTrack; decoderConfig: AudioDecoderConfig | null = null; constructor(internalTrack: InternalAudioTrack) { super(internalTrack); this.internalTrack = internalTrack; } getType() { return 'audio' as const; } override getCodec(): AudioCodec | null { return this.internalTrack.info.codec; } getNumberOfChannels() { return this.internalTrack.info.numberOfChannels; } getSampleRate() { return this.internalTrack.info.sampleRate; } async getDecoderConfig(): Promise { if (!this.internalTrack.info.codec) { return null; } return this.decoderConfig ??= { codec: extractAudioCodecString(this.internalTrack.info), numberOfChannels: this.internalTrack.info.numberOfChannels, sampleRate: this.internalTrack.info.sampleRate, description: this.internalTrack.info.codecDescription ?? undefined, }; } } const getSampleIndexForTimestamp = (sampleTable: SampleTable, timescaleUnits: number) => { if (sampleTable.presentationTimestamps) { const index = binarySearchLessOrEqual( sampleTable.presentationTimestamps, timescaleUnits, x => x.presentationTimestamp, ); if (index === -1) { return -1; } return sampleTable.presentationTimestamps[index]!.sampleIndex; } else { const index = binarySearchLessOrEqual( sampleTable.sampleTimingEntries, timescaleUnits, x => x.startDecodeTimestamp, ); if (index === -1) { return -1; } const entry = sampleTable.sampleTimingEntries[index]!; return entry.startIndex + Math.min( Math.floor((timescaleUnits - entry.startDecodeTimestamp) / entry.delta), entry.count - 1, ); } }; const getKeyframeSampleIndexForTimestamp = (sampleTable: SampleTable, timescaleUnits: number) => { if (!sampleTable.keySampleIndices) { // Every sample is a keyframe return getSampleIndexForTimestamp(sampleTable, timescaleUnits); } if (sampleTable.presentationTimestamps) { const index = binarySearchLessOrEqual( sampleTable.presentationTimestamps, timescaleUnits, x => x.presentationTimestamp, ); if (index === -1) { return -1; } // Walk the samples in presentation order until we find one that's a keyframe for (let i = index; i >= 0; i--) { const sampleIndex = sampleTable.presentationTimestamps[i]!.sampleIndex; const isKeyFrame = binarySearchExact(sampleTable.keySampleIndices, sampleIndex, x => x) !== -1; if (isKeyFrame) { return sampleIndex; } } return -1; } else { const sampleIndex = getSampleIndexForTimestamp(sampleTable, timescaleUnits); const index = binarySearchLessOrEqual(sampleTable.keySampleIndices, sampleIndex, x => x); return sampleTable.keySampleIndices[index] ?? -1; } }; type SampleInfo = { presentationTimestamp: number; duration: number; sampleOffset: number; sampleSize: number; chunkOffset: number; chunkSize: number; isKeyFrame: boolean; }; const getSampleInfo = (sampleTable: SampleTable, sampleIndex: number): SampleInfo | null => { const timingEntryIndex = binarySearchLessOrEqual(sampleTable.sampleTimingEntries, sampleIndex, x => x.startIndex); const timingEntry = sampleTable.sampleTimingEntries[timingEntryIndex]; if (!timingEntry || timingEntry.startIndex + timingEntry.count <= sampleIndex) { return null; } const decodeTimestamp = timingEntry.startDecodeTimestamp + (sampleIndex - timingEntry.startIndex) * timingEntry.delta; let presentationTimestamp = decodeTimestamp; const offsetEntryIndex = binarySearchLessOrEqual( sampleTable.sampleCompositionTimeOffsets, sampleIndex, x => x.startIndex, ); const offsetEntry = sampleTable.sampleCompositionTimeOffsets[offsetEntryIndex]; if (offsetEntry && sampleIndex - offsetEntry.startIndex < offsetEntry.count) { presentationTimestamp += offsetEntry.offset; } const sampleSize = sampleTable.sampleSizes[Math.min(sampleIndex, sampleTable.sampleSizes.length - 1)]!; const chunkEntryIndex = binarySearchLessOrEqual(sampleTable.sampleToChunk, sampleIndex, x => x.startSampleIndex); const chunkEntry = sampleTable.sampleToChunk[chunkEntryIndex]; assert(chunkEntry); const chunkIndex = chunkEntry.startChunkIndex + Math.floor((sampleIndex - chunkEntry.startSampleIndex) / chunkEntry.samplesPerChunk); const chunkOffset = sampleTable.chunkOffsets[chunkIndex]!; const startSampleIndexOfChunk = chunkEntry.startSampleIndex + (chunkIndex - chunkEntry.startChunkIndex) * chunkEntry.samplesPerChunk; let chunkSize = 0; let sampleOffset = chunkOffset; if (sampleTable.sampleSizes.length === 1) { sampleOffset += sampleSize * (sampleIndex - startSampleIndexOfChunk); chunkSize += sampleSize * chunkEntry.samplesPerChunk; } else { for (let i = startSampleIndexOfChunk; i < startSampleIndexOfChunk + chunkEntry.samplesPerChunk; i++) { const sampleSize = sampleTable.sampleSizes[i]!; if (i < sampleIndex) { sampleOffset += sampleSize; } chunkSize += sampleSize; } } let duration = timingEntry.delta; if (sampleTable.presentationTimestamps) { // In order to accurately compute the duration, we need to take the duration to the next sample in presentation // order, not in decode order const presentationIndex = sampleTable.presentationTimestampIndexMap![sampleIndex]; assert(presentationIndex !== undefined); if (presentationIndex < sampleTable.presentationTimestamps.length - 1) { const nextEntry = sampleTable.presentationTimestamps[presentationIndex + 1]!; const nextPresentationTimestamp = nextEntry.presentationTimestamp; duration = nextPresentationTimestamp - presentationTimestamp; } } return { presentationTimestamp, duration, sampleOffset, sampleSize, chunkOffset, chunkSize, isKeyFrame: sampleTable.keySampleIndices ? binarySearchExact(sampleTable.keySampleIndices, sampleIndex, x => x) !== -1 : true, }; }; const getNextKeyframeIndexForSample = (sampleTable: SampleTable, sampleIndex: number) => { if (!sampleTable.keySampleIndices) { return sampleIndex + 1; } const index = binarySearchLessOrEqual(sampleTable.keySampleIndices, sampleIndex, x => x); return sampleTable.keySampleIndices[index + 1] ?? -1; }; const offsetFragmentTrackDataByTimestamp = (trackData: FragmentTrackData, timestamp: number) => { trackData.startTimestamp += timestamp; trackData.endTimestamp += timestamp; for (const sample of trackData.samples) { sample.presentationTimestamp += timestamp; } for (const entry of trackData.presentationTimestamps) { entry.presentationTimestamp += timestamp; } }; /** Extracts the rotation component from a transformation matrix, in degrees. */ const extractRotationFromMatrix = (matrix: TransformationMatrix) => { const [a, b] = matrix; // (1, 0) projects onto (a, b), so that's all we need const radians = Math.atan2(b, a); if (!Number.isFinite(radians)) { // Can happen if the entire matrix is 0, for example return 0; } return radians * (180 / Math.PI); }; const sampleTableIsEmpty = (sampleTable: SampleTable) => { return sampleTable.sampleSizes.length === 0; }; const getOrCreateEncryptionAuxInfo = (track: InternalTrack) => { if (track.currentFragmentState) { return track.currentFragmentState.encryptionAuxInfo ??= { defaultSampleInfoSize: 0, sampleSizes: null, sampleCount: 0, offset: null, resolved: null, }; } else { return track.encryptionAuxInfo ??= { defaultSampleInfoSize: 0, sampleSizes: null, sampleCount: 0, offset: null, resolved: null, }; } }; const resolveEncryptionAuxInfo = async ( reader: Reader, encryptionInfo: TrackEncryptionInfo, aux: SampleEncryptionAuxInfo, ) => { if (aux.resolved) { return aux.resolved; } if (aux.offset === null || aux.sampleCount === 0) { throw new Error('Incomplete saiz/saio info; cannot resolve encryption data.'); } let totalSize = 0; if (aux.defaultSampleInfoSize > 0) { totalSize = aux.defaultSampleInfoSize * aux.sampleCount; } else { assert(aux.sampleSizes); for (let i = 0; i < aux.sampleCount; i++) { totalSize += aux.sampleSizes[i]!; } } let slice = reader.requestSlice(aux.offset, totalSize); if (slice instanceof Promise) slice = await slice; if (!slice) { throw new Error('Failed to read auxiliary encryption info.'); } const ivSize = encryptionInfo.defaultPerSampleIvSize; assert(ivSize !== null); // Each aux entry has the same byte layout as a senc entry: IV (of size ivSize, or the constant IV from tenc // when ivSize is 0), then optionally subsample count + [clearLen, protectedLen] pairs. Subsamples are present // iff the entry is larger than the IV. const entries: SampleEncryptionInfo[] = []; for (let i = 0; i < aux.sampleCount; i++) { const entrySize = aux.defaultSampleInfoSize > 0 ? aux.defaultSampleInfoSize : aux.sampleSizes![i]!; const iv = new Uint8Array(16); if (ivSize > 0) { iv.set(readBytes(slice, ivSize), 0); } else { iv.set(encryptionInfo.defaultConstantIv!, 0); } let subsamples: { clearLen: number; protectedLen: number }[] | null = null; if (entrySize > ivSize) { const subsampleCount = readU16Be(slice); subsamples = []; for (let j = 0; j < subsampleCount; j++) { const clearLen = readU16Be(slice); const protectedLen = readU32Be(slice); subsamples.push({ clearLen, protectedLen }); } } entries.push({ iv, subsamples }); } aux.resolved = entries; return entries; }; const decryptSample = async ( track: InternalTrack, sampleEncryption: SampleEncryptionInfo, data: Uint8Array, fragment: Fragment | null, ): Promise => { assert(track.encryptionInfo); const encryptionInfo = track.encryptionInfo; assert(encryptionInfo.defaultKid !== null); const keyId = encryptionInfo.defaultKid; let keyBytes: Uint8Array; const cacheEntry = track.demuxer.decryptionKeyCache.get(keyId); if (cacheEntry) { keyBytes = await cacheEntry; } else { if (!track.demuxer.input._formatOptions.isobmff?.resolveKeyId) { throw new Error( 'Encrypted media samples encountered. To decrypt them, please provide a callback for' + ' InputOptions.formatOptions.isobmff.resolveKeyId.', ); } const promise = (async () => { let psshBoxes = track.demuxer.psshBoxes; if (fragment) { psshBoxes = [ ...psshBoxes, ...fragment.psshBoxes, ].filter(x => x.keyIds === null || x.keyIds.includes(keyId)); // Filter out duplicates for (let i = 0; i < psshBoxes.length - 1; i++) { for (let j = i + 1; j < psshBoxes.length; j++) { if (psshBoxesAreEqual(psshBoxes[i]!, psshBoxes[j]!)) { psshBoxes.splice(j, 1); j--; } } } } const keyResult = await track.demuxer.input._formatOptions.isobmff!.resolveKeyId!({ keyId, psshBoxes }); if (!( (typeof keyResult === 'string' && keyResult.length === 32 && HEX_STRING_REGEX.test(keyResult)) || (keyResult instanceof Uint8Array && keyResult.byteLength === 16) )) { throw new TypeError( 'resolveKeyId must return a 32-character hex string or a 16-byte Uint8Array containing the' + ' decryption key.', ); } return keyResult instanceof Uint8Array ? keyResult : hexStringToBytes(keyResult); })(); track.demuxer.decryptionKeyCache.set(keyId, promise); keyBytes = await promise; } if (encryptionInfo.scheme === 'cenc' || encryptionInfo.scheme === 'cens') { return decryptCtr(keyBytes, encryptionInfo, sampleEncryption, data); } else { return decryptCbcs(keyBytes, encryptionInfo, sampleEncryption, data); } }; const decryptCtr = async ( key: Uint8Array, encryptionInfo: TrackEncryptionInfo, sampleEncryption: SampleEncryptionInfo, data: Uint8Array, ) => { const counter = new Uint8Array(16); counter.set(sampleEncryption.iv, 0); const cryptoKey = await crypto.subtle.importKey( 'raw', key as BufferSource, { name: 'AES-CTR' }, false, ['decrypt'], ); const cryptApply = async (input: Uint8Array) => { const plaintext = await crypto.subtle.decrypt( { name: 'AES-CTR', counter, length: 64 }, cryptoKey, input as BufferSource, ); return new Uint8Array(plaintext); }; if (!sampleEncryption.subsamples) { // Whole sample is protected, no pattern return cryptApply(data); } assert(encryptionInfo.defaultCryptByteBlock !== null && encryptionInfo.defaultSkipByteBlock !== null); const cryptRanges = collectCryptRanges( sampleEncryption.subsamples, encryptionInfo.defaultCryptByteBlock, encryptionInfo.defaultSkipByteBlock, ); // Concatenate all crypt ranges into a single buffer so the continuous CTR counter behavior is preserved let totalCryptLen = 0; for (const range of cryptRanges) { for (const seg of range.perSubsample) { totalCryptLen += seg.length; } } const cryptBuffer = new Uint8Array(totalCryptLen); let writePos = 0; for (const range of cryptRanges) { for (const seg of range.perSubsample) { cryptBuffer.set(data.subarray(seg.offset, seg.offset + seg.length), writePos); writePos += seg.length; } } const plain = await cryptApply(cryptBuffer); // Now let's build the output const output = new Uint8Array(data); let readPos = 0; for (const range of cryptRanges) { for (const seg of range.perSubsample) { output.set(plain.subarray(readPos, readPos + seg.length), seg.offset); readPos += seg.length; } } return output; }; const decryptCbcs = ( key: Uint8Array, encryptionInfo: TrackEncryptionInfo, sampleEncryption: SampleEncryptionInfo, data: Uint8Array, ) => { const ctx = new Aes128CbcContext(); ctx.init({ key, iv: sampleEncryption.iv }); const cryptByteBlock = encryptionInfo.defaultCryptByteBlock; const skipByteBlock = encryptionInfo.defaultSkipByteBlock; assert(cryptByteBlock !== null && skipByteBlock !== null); if (!sampleEncryption.subsamples) { // Whole-sample encryption: straightforward CBC over floor(size / 16) blocks, any trailing bytes stay clear const output = new Uint8Array(data); const numBlocks = Math.floor(data.length / 16); for (let b = 0; b < numBlocks; b++) { const off = b * 16; ctx.in.set(data.subarray(off, off + 16)); ctx.decrypt(); output.set(ctx.out, off); } return output; } if (cryptByteBlock === 0 && skipByteBlock === 0) { throw new Error('cbcs with subsamples requires pattern encryption.'); } const output = new Uint8Array(data); // Pattern decryption: IV is reset at the start of each subsample. Within a subsample, the CBC chain continues // across skipped blocks (the IV after a crypt group carries over to the next crypt group's first block). const cryptRanges = collectCryptRanges(sampleEncryption.subsamples, cryptByteBlock, skipByteBlock); const ivView = new DataView(sampleEncryption.iv.buffer, sampleEncryption.iv.byteOffset, 16); for (const range of cryptRanges) { // Reset IV per subsample ctx.iv[0] = ivView.getUint32(0, false); ctx.iv[1] = ivView.getUint32(4, false); ctx.iv[2] = ivView.getUint32(8, false); ctx.iv[3] = ivView.getUint32(12, false); for (const seg of range.perSubsample) { // Decrypt length / 16 blocks at this offset const numBlocks = seg.length / 16; for (let b = 0; b < numBlocks; b++) { const offset = seg.offset + b * 16; ctx.in.set(data.subarray(offset, offset + 16)); ctx.decrypt(); output.set(ctx.out, offset); } } } return output; }; const collectCryptRanges = ( subsamples: { clearLen: number; protectedLen: number }[], cryptByteBlock: number, skipByteBlock: number, ) => { const ranges: { perSubsample: { offset: number; length: number }[] }[] = []; const hasPattern = cryptByteBlock !== 0 || skipByteBlock !== 0; let cursor = 0; for (const subsample of subsamples) { cursor += subsample.clearLen; const perSubsample: { offset: number; length: number }[] = []; if (!hasPattern) { if (subsample.protectedLen > 0) { perSubsample.push({ offset: cursor, length: subsample.protectedLen }); } cursor += subsample.protectedLen; } else { let remaining = subsample.protectedLen; let pos = cursor; while (remaining > 0) { if (remaining < 16 * cryptByteBlock) { break; // Partial final crypt group stays clear } const cryptBytes = 16 * cryptByteBlock; perSubsample.push({ offset: pos, length: cryptBytes }); pos += cryptBytes; remaining -= cryptBytes; const skipBytes = Math.min(16 * skipByteBlock, remaining); pos += skipBytes; remaining -= skipBytes; } cursor += subsample.protectedLen; } ranges.push({ perSubsample }); } return ranges; }; ===== src/isobmff/isobmff-boxes.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { toUint8Array, assert, isU32, last, TransformationMatrix, textEncoder, COLOR_PRIMARIES_MAP, TRANSFER_CHARACTERISTICS_MAP, MATRIX_COEFFICIENTS_MAP, colorSpaceIsComplete, UNDETERMINED_LANGUAGE, assertNever, keyValueIterator, } from '../misc'; import { AudioCodec, generateAv1CodecConfigurationFromCodecString, parsePcmCodec, PCM_AUDIO_CODECS, PcmAudioCodec, SubtitleCodec, VideoCodec, } from '../codec'; import { formatSubtitleTimestamp } from '../subtitles'; import { Writer } from '../writer'; import { getTrackMetadata, GLOBAL_TIMESCALE, intoTimescale, IsobmffAudioTrackData, IsobmffMuxer, IsobmffSubtitleTrackData, IsobmffTrackData, IsobmffVideoTrackData, Sample, } from './isobmff-muxer'; import { parseAc3SyncFrame, parseEac3SyncFrame, parseOpusIdentificationHeader } from '../codec-data'; import { MetadataTags, RichImageData } from '../metadata'; import { Bitstream } from '../../shared/bitstream'; export class IsobmffBoxWriter { private helper = new Uint8Array(8); private helperView = new DataView(this.helper.buffer); /** * Stores the position from the start of the file to where boxes elements have been written. This is used to * rewrite/edit elements that were already added before, and to measure sizes of things. */ offsets = new WeakMap(); constructor(public writer: Writer) {} writeU32(value: number) { this.helperView.setUint32(0, value, false); this.writer.write(this.helper.subarray(0, 4)); } writeU64(value: number) { this.helperView.setUint32(0, Math.floor(value / 2 ** 32), false); this.helperView.setUint32(4, value, false); this.writer.write(this.helper.subarray(0, 8)); } writeAscii(text: string) { for (let i = 0; i < text.length; i++) { this.helperView.setUint8(i % 8, text.charCodeAt(i)); if (i % 8 === 7) this.writer.write(this.helper); } if (text.length % 8 !== 0) { this.writer.write(this.helper.subarray(0, text.length % 8)); } } writeBox(box: Box) { this.offsets.set(box, this.writer.getPos()); if (box.contents && !box.children) { this.writeBoxHeader(box, box.size ?? box.contents.byteLength + 8); this.writer.write(box.contents); } else { const startPos = this.writer.getPos(); this.writeBoxHeader(box, 0); if (box.contents) this.writer.write(box.contents); if (box.children) for (const child of box.children) if (child) this.writeBox(child); const endPos = this.writer.getPos(); const size = box.size ?? endPos - startPos; this.writer.seek(startPos); this.writeBoxHeader(box, size); this.writer.seek(endPos); } } writeBoxHeader(box: Box, size: number) { this.writeU32(box.largeSize ? 1 : size); this.writeAscii(box.type); if (box.largeSize) this.writeU64(size); } measureBoxHeader(box: Box) { return 8 + (box.largeSize ? 8 : 0); } patchBox(box: Box) { const boxOffset = this.offsets.get(box); assert(boxOffset !== undefined); const endPos = this.writer.getPos(); this.writer.seek(boxOffset); this.writeBox(box); this.writer.seek(endPos); } measureBox(box: Box) { if (box.contents && !box.children) { const headerSize = this.measureBoxHeader(box); return headerSize + box.contents.byteLength; } else { let result = this.measureBoxHeader(box); if (box.contents) result += box.contents.byteLength; if (box.children) for (const child of box.children) if (child) result += this.measureBox(child); return result; } } } const bytes = /* #__PURE__ */ new Uint8Array(8); const view = /* #__PURE__ */ new DataView(bytes.buffer); const u8 = (value: number) => { return [(value % 0x100 + 0x100) % 0x100]; }; const u16 = (value: number) => { view.setUint16(0, value, false); return [bytes[0], bytes[1]] as number[]; }; const i16 = (value: number) => { view.setInt16(0, value, false); return [bytes[0], bytes[1]] as number[]; }; const u24 = (value: number) => { view.setUint32(0, value, false); return [bytes[1], bytes[2], bytes[3]] as number[]; }; const u32 = (value: number) => { view.setUint32(0, value, false); return [bytes[0], bytes[1], bytes[2], bytes[3]] as number[]; }; const i32 = (value: number) => { view.setInt32(0, value, false); return [bytes[0], bytes[1], bytes[2], bytes[3]] as number[]; }; const u64 = (value: number) => { view.setUint32(0, Math.floor(value / 2 ** 32), false); view.setUint32(4, value, false); return [bytes[0], bytes[1], bytes[2], bytes[3], bytes[4], bytes[5], bytes[6], bytes[7]] as number[]; }; const i64 = (value: number) => { view.setInt32(0, Math.floor(value / 2 ** 32), false); view.setUint32(4, value, false); return [bytes[0], bytes[1], bytes[2], bytes[3], bytes[4], bytes[5], bytes[6], bytes[7]] as number[]; }; const fixed_8_8 = (value: number) => { view.setInt16(0, 2 ** 8 * value, false); return [bytes[0], bytes[1]] as number[]; }; const fixed_16_16 = (value: number) => { view.setInt32(0, 2 ** 16 * value, false); return [bytes[0], bytes[1], bytes[2], bytes[3]] as number[]; }; const fixed_2_30 = (value: number) => { view.setInt32(0, 2 ** 30 * value, false); return [bytes[0], bytes[1], bytes[2], bytes[3]] as number[]; }; const variableUnsignedInt = (value: number, byteLength?: number) => { const bytes: number[] = []; let remaining = value; do { let byte = remaining & 0x7f; remaining >>= 7; // If this isn't the first byte we're adding (meaning there will be more bytes after it // when we reverse the array), set the continuation bit if (bytes.length > 0) { byte |= 0x80; } bytes.push(byte); if (byteLength !== undefined) { byteLength--; } } while (remaining > 0 || byteLength); // Reverse the array since we built it backwards return bytes.reverse(); }; const ascii = (text: string, nullTerminated = false) => { const bytes = Array(text.length).fill(null).map((_, i) => text.charCodeAt(i)); if (nullTerminated) bytes.push(0x00); return bytes; }; const rotationMatrix = (rotationInDegrees: number): TransformationMatrix => { const theta = rotationInDegrees * (Math.PI / 180); const cosTheta = Math.round(Math.cos(theta)); const sinTheta = Math.round(Math.sin(theta)); // Matrices are post-multiplied in ISOBMFF, meaning this is the transpose of your typical rotation matrix return [ cosTheta, sinTheta, 0, -sinTheta, cosTheta, 0, 0, 0, 1, ]; }; const IDENTITY_MATRIX = /* #__PURE__ */ rotationMatrix(0); const matrixToBytes = (matrix: TransformationMatrix) => { return [ fixed_16_16(matrix[0]), fixed_16_16(matrix[1]), fixed_2_30(matrix[2]), fixed_16_16(matrix[3]), fixed_16_16(matrix[4]), fixed_2_30(matrix[5]), fixed_16_16(matrix[6]), fixed_16_16(matrix[7]), fixed_2_30(matrix[8]), ]; }; export interface Box { type: string; contents?: Uint8Array; children?: (Box | null)[]; size?: number; largeSize?: boolean; } type NestedNumberArray = (number | NestedNumberArray)[]; export const box = (type: string, contents?: NestedNumberArray, children?: (Box | null)[]): Box => ({ type, contents: contents && new Uint8Array(contents.flat(10) as number[]), children, }); /** A FullBox always starts with a version byte, followed by three flag bytes. */ export const fullBox = ( type: string, version: number, flags: number, contents?: NestedNumberArray, children?: Box[], ) => box( type, [u8(version), u24(flags), contents ?? []], children, ); /** * File Type Compatibility Box: Allows the reader to determine whether this is a type of file that the * reader understands. */ export const ftyp = (details: { isQuickTime: boolean; holdsAvc: boolean; fragmented: boolean; cmaf: boolean; }) => { // You can find the full logic for this at // https://github.com/FFmpeg/FFmpeg/blob/de2fb43e785773738c660cdafb9309b1ef1bc80d/libavformat/movenc.c#L5518 // Obviously, this lib only needs a small subset of that logic. const minorVersion = 0x200; if (details.isQuickTime) { return box('ftyp', [ ascii('qt '), // Major brand u32(minorVersion), // Minor version // Compatible brands ascii('qt '), ]); } if (details.fragmented) { if (details.cmaf) { return box('ftyp', [ ascii('iso5'), // Major brand u32(minorVersion), // Minor version // Compatible brands ascii('iso5'), ascii('iso6'), ascii('mp41'), ascii('cmfc'), ascii('dash'), ]); } else { return box('ftyp', [ ascii('iso5'), // Major brand u32(minorVersion), // Minor version // Compatible brands ascii('iso5'), ascii('iso6'), ascii('mp41'), ]); } } return box('ftyp', [ ascii('isom'), // Major brand u32(minorVersion), // Minor version // Compatible brands ascii('isom'), details.holdsAvc ? ascii('avc1') : [], ascii('mp41'), ]); }; /** Segment Type Box */ export const styp = () => box('styp', [ ascii('iso5'), // Major brand u32(0), // Minor version // Compatible brands ascii('iso5'), ascii('iso6'), ascii('mp41'), ascii('cmfc'), ascii('dash'), ]); /** Segment Index Box */ export const sidx = (muxer: IsobmffMuxer, referencedSize: number) => { let duration = muxer.maxWrittenEndTimestamp - muxer.minWrittenTimestamp; if (!Number.isFinite(duration)) { duration = 0; } return fullBox('sidx', 1, 0, [ u32(1), // Reference ID u32(GLOBAL_TIMESCALE), // Timescale u64(intoTimescale(muxer.minWrittenTimestamp, GLOBAL_TIMESCALE)), // Earliest presentation time u64(0), // First offset u16(0), // Reserved u16(1), // Reference count u32(referencedSize & 0x7fffffff), // Reference type (0) + referenced size u32(intoTimescale(duration, GLOBAL_TIMESCALE)), // Subsegment duration u32(0), // Starts with SAP + SAP type + SAP delta time (no information provided) ]); }; /** Movie Sample Data Box. Contains the actual frames/samples of the media. */ export const mdat = (reserveLargeSize: boolean): Box => ({ type: 'mdat', largeSize: reserveLargeSize }); /** Free Space Box: A box that designates unused space in the movie data file. */ export const free = (size: number): Box => ({ type: 'free', size }); /** * Movie Box: Used to specify the information that defines a movie - that is, the information that allows * an application to interpret the sample data that is stored elsewhere. */ export const moov = (muxer: IsobmffMuxer) => { return box('moov', undefined, [ mvhd(muxer.creationTime, muxer.trackDatas), ...muxer.trackDatas.map(x => trak(x, muxer.creationTime)), muxer.isFragmented ? mvex(muxer.trackDatas) : null, udta(muxer), ]); }; /** Movie Header Box: Used to specify the characteristics of the entire movie, such as timescale and duration. */ export const mvhd = ( creationTime: number, trackDatas: IsobmffTrackData[], ) => { const duration = Math.max( 0, ...trackDatas .map(trackData => ( intoTimescale(presentationSpan(trackData), GLOBAL_TIMESCALE) + intoTimescale(trackData.startTimestampOffset ?? 0, GLOBAL_TIMESCALE) )), ); const nextTrackId = Math.max(0, ...trackDatas.map(x => x.track.id)) + 1; // Conditionally use u64 if u32 isn't enough const needsU64 = !isU32(creationTime) || !isU32(duration); const u32OrU64 = needsU64 ? u64 : u32; return fullBox('mvhd', +needsU64, 0, [ u32OrU64(creationTime), // Creation time u32OrU64(creationTime), // Modification time u32(GLOBAL_TIMESCALE), // Timescale u32OrU64(duration), // Duration fixed_16_16(1), // Preferred rate fixed_8_8(1), // Preferred volume Array(10).fill(0), // Reserved matrixToBytes(IDENTITY_MATRIX), // Matrix Array(24).fill(0), // Pre-defined u32(nextTrackId), // Next track ID ]); }; const presentationSpan = (trackData: IsobmffTrackData) => { if (trackData.samples.length === 0) { return 0; } let minTimestamp = Infinity; let maxEndTimestamp = -Infinity; for (let i = 0; i < trackData.samples.length; i++) { const sample = trackData.samples[i]!; if (sample.timestamp < minTimestamp) { minTimestamp = sample.timestamp; } if (sample.timestamp + sample.duration > maxEndTimestamp) { maxEndTimestamp = sample.timestamp + sample.duration; } } if (minTimestamp === Infinity) { return 0; } return maxEndTimestamp - minTimestamp; }; /** * Track Box: Defines a single track of a movie. A movie may consist of one or more tracks. Each track is * independent of the other tracks in the movie and carries its own temporal and spatial information. Each Track Box * contains its associated Media Box. */ export const trak = (trackData: IsobmffTrackData, creationTime: number) => { const trackMetadata = getTrackMetadata(trackData); const needsEditList = trackData.startTimestampOffset !== null && trackData.startTimestampOffset > 0; return box('trak', undefined, [ tkhd(trackData, creationTime), needsEditList ? edts(trackData, trackData.startTimestampOffset!) : null, mdia(trackData, creationTime), trackMetadata.name !== undefined ? box('udta', undefined, [ box('name', [ // VLC (and Mediabunny) also recognize ©nam ...textEncoder.encode(trackMetadata.name), ]), ]) : null, ]); }; /** Track Header Box: Specifies the characteristics of a single track within a movie. */ export const tkhd = ( trackData: IsobmffTrackData, creationTime: number, ) => { const durationInGlobalTimescale = intoTimescale(presentationSpan(trackData), GLOBAL_TIMESCALE) + intoTimescale(trackData.startTimestampOffset ?? 0, GLOBAL_TIMESCALE); const needsU64 = !isU32(creationTime) || !isU32(durationInGlobalTimescale); const u32OrU64 = needsU64 ? u64 : u32; let matrix: TransformationMatrix; if (trackData.type === 'video') { const rotation = trackData.track.metadata.rotation; matrix = rotationMatrix(rotation ?? 0); } else { matrix = IDENTITY_MATRIX; } let flags = 0x2; // Track in movie if (trackData.track.metadata.disposition?.default !== false) { flags |= 0x1; // Track enabled } return fullBox('tkhd', +needsU64, flags, [ u32OrU64(creationTime), // Creation time u32OrU64(creationTime), // Modification time u32(trackData.track.id), // Track ID u32(0), // Reserved u32OrU64(durationInGlobalTimescale), // Duration Array(8).fill(0), // Reserved u16(0), // Layer u16(trackData.track.id), // Alternate group fixed_8_8(trackData.type === 'audio' ? 1 : 0), // Volume u16(0), // Reserved matrixToBytes(matrix), // Matrix fixed_16_16(trackData.type === 'video' ? trackData.info.width : 0), // Track width fixed_16_16(trackData.type === 'video' ? trackData.info.height : 0), // Track height ]); }; /** Edit Box: Specifies edits to the track's media. */ export const edts = (trackData: IsobmffTrackData, offset: number) => { const startOffset = intoTimescale(offset, GLOBAL_TIMESCALE); const mediaDuration = intoTimescale(presentationSpan(trackData), GLOBAL_TIMESCALE); const needs64Bits = !isU32(startOffset) || !isU32(mediaDuration); const u32OrU64 = needs64Bits ? u64 : u32; const i32OrI64 = needs64Bits ? i64 : i32; return box('edts', undefined, [ fullBox('elst', needs64Bits ? 1 : 0, 0, [ u32(2), // Entry count // #1 u32OrU64(startOffset), // Segment duration i32OrI64(-1), // Media time fixed_16_16(1), // Media rate // #2 u32OrU64(mediaDuration), // Segment duration i32OrI64(0), // Media time fixed_16_16(1), // Media rate ]), ]); }; /** Media Box: Describes and define a track's media type and sample data. */ export const mdia = (trackData: IsobmffTrackData, creationTime: number) => box('mdia', undefined, [ mdhd(trackData, creationTime), hdlr(true, TRACK_TYPE_TO_COMPONENT_SUBTYPE[trackData.type], TRACK_TYPE_TO_HANDLER_NAME[trackData.type]), minf(trackData), ]); /** Media Header Box: Specifies the characteristics of a media, including timescale and duration. */ export const mdhd = ( trackData: IsobmffTrackData, creationTime: number, ) => { // Since the duration represents the raw media duration, edit list offsets are not taken into account here const localDuration = intoTimescale( presentationSpan(trackData), trackData.timescale, ); const needsU64 = !isU32(creationTime) || !isU32(localDuration); const u32OrU64 = needsU64 ? u64 : u32; return fullBox('mdhd', +needsU64, 0, [ u32OrU64(creationTime), // Creation time u32OrU64(creationTime), // Modification time u32(trackData.timescale), // Timescale u32OrU64(localDuration), // Duration u16(getLanguageCodeInt(trackData.track.metadata.languageCode ?? UNDETERMINED_LANGUAGE)), // Language u16(0), // Quality ]); }; const TRACK_TYPE_TO_COMPONENT_SUBTYPE: Record = { video: 'vide', audio: 'soun', subtitle: 'text', }; const TRACK_TYPE_TO_HANDLER_NAME: Record = { video: 'MediabunnyVideoHandler', audio: 'MediabunnySoundHandler', subtitle: 'MediabunnyTextHandler', }; /** Handler Reference Box. */ export const hdlr = ( hasComponentType: boolean, handlerType: string, name: string, manufacturer = '\0\0\0\0', ) => fullBox('hdlr', 0, 0, [ hasComponentType ? ascii('mhlr') : u32(0), // Component type ascii(handlerType), // Component subtype ascii(manufacturer), // Component manufacturer u32(0), // Component flags u32(0), // Component flags mask ascii(name, true), // Component name ]); /** * Media Information Box: Stores handler-specific information for a track's media data. The media handler uses this * information to map from media time to media data and to process the media data. */ export const minf = (trackData: IsobmffTrackData) => box('minf', undefined, [ TRACK_TYPE_TO_HEADER_BOX[trackData.type](), dinf(), stbl(trackData), ]); /** Video Media Information Header Box: Defines specific color and graphics mode information. */ export const vmhd = () => fullBox('vmhd', 0, 1, [ u16(0), // Graphics mode u16(0), // Opcolor R u16(0), // Opcolor G u16(0), // Opcolor B ]); /** Sound Media Information Header Box: Stores the sound media's control information, such as balance. */ export const smhd = () => fullBox('smhd', 0, 0, [ u16(0), // Balance u16(0), // Reserved ]); /** Null Media Header Box. */ export const nmhd = () => fullBox('nmhd', 0, 0); const TRACK_TYPE_TO_HEADER_BOX: Record Box> = { video: vmhd, audio: smhd, subtitle: nmhd, }; /** * Data Information Box: Contains information specifying the data handler component that provides access to the * media data. The data handler component uses the Data Information Box to interpret the media's data. */ export const dinf = () => box('dinf', undefined, [ dref(), ]); /** * Data Reference Box: Contains tabular data that instructs the data handler component how to access the media's data. */ export const dref = () => fullBox('dref', 0, 0, [ u32(1), // Entry count ], [ url(), ]); export const url = () => fullBox('url ', 0, 1); // Self-reference flag enabled /** * Sample Table Box: Contains information for converting from media time to sample number to sample location. This box * also indicates how to interpret the sample (for example, whether to decompress the video data and, if so, how). */ export const stbl = (trackData: IsobmffTrackData) => { const needsCtts = trackData.compositionTimeOffsetTable.length > 1 || trackData.compositionTimeOffsetTable.some(x => x.sampleCompositionTimeOffset !== 0); return box('stbl', undefined, [ stsd(trackData), stts(trackData), needsCtts ? ctts(trackData) : null, needsCtts ? cslg(trackData) : null, stsc(trackData), stsz(trackData), stco(trackData), stss(trackData), ]); }; /** * Sample Description Box: Stores information that allows you to decode samples in the media. The data stored in the * sample description varies, depending on the media type. */ export const stsd = (trackData: IsobmffTrackData) => { let sampleDescription: Box; if (trackData.type === 'video') { sampleDescription = videoSampleDescription( videoCodecToBoxName(trackData.track.source._codec, trackData.info.decoderConfig.codec), trackData, ); } else if (trackData.type === 'audio') { const boxName = audioCodecToBoxName(trackData.track.source._codec, trackData.muxer.isQuickTime); assert(boxName); sampleDescription = soundSampleDescription( boxName, trackData, ); } else if (trackData.type === 'subtitle') { sampleDescription = subtitleSampleDescription( SUBTITLE_CODEC_TO_BOX_NAME[trackData.track.source._codec], trackData, ); } assert(sampleDescription!); return fullBox('stsd', 0, 0, [ u32(1), // Entry count ], [ sampleDescription, ]); }; /** Video Sample Description Box: Contains information that defines how to interpret video media data. */ export const videoSampleDescription = ( compressionType: string, trackData: IsobmffVideoTrackData, ) => box(compressionType, [ Array(6).fill(0), // Reserved u16(1), // Data reference index u16(0), // Pre-defined u16(0), // Reserved Array(12).fill(0), // Pre-defined u16(trackData.info.width), // Width u16(trackData.info.height), // Height u32(0x00480000), // Horizontal resolution u32(0x00480000), // Vertical resolution u32(0), // Reserved u16(1), // Frame count Array(32).fill(0), // Compressor name u16(0x0018), // Depth i16(0xffff), // Pre-defined ], [ VIDEO_CODEC_TO_CONFIGURATION_BOX[trackData.track.source._codec](trackData), pasp(trackData), colorSpaceIsComplete(trackData.info.decoderConfig.colorSpace) ? colr(trackData) : null, ]); /** Pixel Aspect Ratio Box: Specifies pixel width:height spacing for non-square pixels. */ export const pasp = (trackData: IsobmffVideoTrackData) => { if (trackData.info.pixelAspectRatio.num === trackData.info.pixelAspectRatio.den) { return null; } return box('pasp', [ u32(trackData.info.pixelAspectRatio.num), u32(trackData.info.pixelAspectRatio.den), ]); }; /** Colour Information Box: Specifies the color space of the video. */ export const colr = (trackData: IsobmffVideoTrackData) => box('colr', [ ascii(trackData.muxer.isQuickTime ? 'nclc' : 'nclx'), // Colour type u16(COLOR_PRIMARIES_MAP[trackData.info.decoderConfig.colorSpace!.primaries!]), // Colour primaries u16(TRANSFER_CHARACTERISTICS_MAP[trackData.info.decoderConfig.colorSpace!.transfer!]), // Transfer characteristics u16(MATRIX_COEFFICIENTS_MAP[trackData.info.decoderConfig.colorSpace!.matrix!]), // Matrix coefficients trackData.muxer.isQuickTime ? [] // Doesn't have it : u8((trackData.info.decoderConfig.colorSpace!.fullRange ? 1 : 0) << 7), // Full range flag ]); /** AVC Configuration Box: Provides additional information to the decoder. */ export const avcC = (trackData: IsobmffVideoTrackData) => trackData.info.decoderConfig && box('avcC', [ // For AVC, description is an AVCDecoderConfigurationRecord, so nothing else to do here ...toUint8Array(trackData.info.decoderConfig.description!), ]); /** HEVC Configuration Box: Provides additional information to the decoder. */ export const hvcC = (trackData: IsobmffVideoTrackData) => trackData.info.decoderConfig && box('hvcC', [ // For HEVC, description is an HEVCDecoderConfigurationRecord, so nothing else to do here ...toUint8Array(trackData.info.decoderConfig.description!), ]); /** VP Configuration Box: Provides additional information to the decoder. */ export const vpcC = (trackData: IsobmffVideoTrackData) => { // Reference: https://www.webmproject.org/vp9/mp4/ if (!trackData.info.decoderConfig) { return null; } const decoderConfig = trackData.info.decoderConfig; const parts = decoderConfig.codec.split('.'); // We can derive the required values from the codec string const profile = Number(parts[1]); const level = Number(parts[2]); const bitDepth = Number(parts[3]); const chromaSubsampling = parts[4] ? Number(parts[4]) : 1; // 4:2:0 colocated with luma (0,0) const videoFullRangeFlag = parts[8] ? Number(parts[8]) : Number(decoderConfig.colorSpace?.fullRange ?? 0); const thirdByte = (bitDepth << 4) + (chromaSubsampling << 1) + videoFullRangeFlag; const colourPrimaries = parts[5] ? Number(parts[5]) : decoderConfig.colorSpace?.primaries ? COLOR_PRIMARIES_MAP[decoderConfig.colorSpace.primaries] : 2; // Default to undetermined const transferCharacteristics = parts[6] ? Number(parts[6]) : decoderConfig.colorSpace?.transfer ? TRANSFER_CHARACTERISTICS_MAP[decoderConfig.colorSpace.transfer] : 2; const matrixCoefficients = parts[7] ? Number(parts[7]) : decoderConfig.colorSpace?.matrix ? MATRIX_COEFFICIENTS_MAP[decoderConfig.colorSpace.matrix] : 2; return fullBox('vpcC', 1, 0, [ u8(profile), // Profile u8(level), // Level u8(thirdByte), // Bit depth, chroma subsampling, full range u8(colourPrimaries), // Colour primaries u8(transferCharacteristics), // Transfer characteristics u8(matrixCoefficients), // Matrix coefficients u16(0), // Codec initialization data size ]); }; /** AV1 Configuration Box: Provides additional information to the decoder. */ export const av1C = (trackData: IsobmffVideoTrackData) => { return box('av1C', generateAv1CodecConfigurationFromCodecString(trackData.info.decoderConfig.codec)); }; /** Sound Sample Description Box: Contains information that defines how to interpret sound media data. */ export const soundSampleDescription = ( compressionType: string, trackData: IsobmffAudioTrackData, ) => { let version = 0; let contents: NestedNumberArray; let sampleSizeInBits = 16; const isPcmCodec = (PCM_AUDIO_CODECS as readonly AudioCodec[]).includes(trackData.track.source._codec); if (isPcmCodec) { const codec = trackData.track.source._codec as PcmAudioCodec; const { sampleSize } = parsePcmCodec(codec); sampleSizeInBits = 8 * sampleSize; if (sampleSizeInBits > 16) { version = 1; } } if (trackData.muxer.isQuickTime) { version = 1; } if (version === 0) { contents = [ Array(6).fill(0), // Reserved u16(1), // Data reference index u16(version), // Version u16(0), // Revision level u32(0), // Vendor u16(trackData.info.numberOfChannels), // Number of channels u16(sampleSizeInBits), // Sample size (bits) u16(0), // Compression ID u16(0), // Packet size u16(trackData.info.sampleRate < 2 ** 16 ? trackData.info.sampleRate : 0), // Sample rate (upper) u16(0), // Sample rate (lower) ]; } else { const compressionId = isPcmCodec ? 0 : -2; contents = [ Array(6).fill(0), // Reserved u16(1), // Data reference index u16(version), // Version u16(0), // Revision level u32(0), // Vendor u16(trackData.info.numberOfChannels), // Number of channels u16(Math.min(sampleSizeInBits, 16)), // Sample size (bits) i16(compressionId), // Compression ID u16(0), // Packet size u16(trackData.info.sampleRate < 2 ** 16 ? trackData.info.sampleRate : 0), // Sample rate (upper) u16(0), // Sample rate (lower) isPcmCodec ? [ u32(1), // Samples per packet (must be 1 for uncompressed formats) u32(sampleSizeInBits / 8), // Bytes per packet u32(trackData.info.numberOfChannels * sampleSizeInBits / 8), // Bytes per frame ] : [ u32(0), // Samples per packet (don't bother, still works with 0) u32(0), // Bytes per packet (variable) u32(0), // Bytes per frame (variable) ], u32(2), // Bytes per sample (constant in FFmpeg) ]; } return box(compressionType, contents, [ audioCodecToConfigurationBox(trackData.track.source._codec, trackData.muxer.isQuickTime)?.(trackData) ?? null, ]); }; /** MPEG-4 Elementary Stream Descriptor Box. */ export const esds = (trackData: IsobmffAudioTrackData) => { // We build up the bytes in a layered way which reflects the nested structure let objectTypeIndication: number; switch (trackData.track.source._codec) { case 'aac': { objectTypeIndication = 0x40; }; break; case 'mp3': { objectTypeIndication = 0x6b; }; break; case 'vorbis': { objectTypeIndication = 0xdd; }; break; default: throw new Error(`Unhandled audio codec: ${trackData.track.source._codec}`); } let bytes = [ ...u8(objectTypeIndication), // Object type indication ...u8(0x15), // stream type(6bits)=5 audio, flags(2bits)=1 ...u24(0), // 24bit buffer size ...u32(0), // max bitrate ...u32(0), // avg bitrate ]; if (trackData.info.decoderConfig.description) { const description = toUint8Array(trackData.info.decoderConfig.description); // Add the decoder description to the end bytes = [ ...bytes, ...u8(0x05), // TAG(5) = DecoderSpecificInfo ...variableUnsignedInt(description.byteLength), ...description, ]; } bytes = [ ...u16(1), // ES_ID = 1 ...u8(0x00), // flags etc = 0 ...u8(0x04), // TAG(4) = ES Descriptor ...variableUnsignedInt(bytes.length), ...bytes, ...u8(0x06), // TAG(6) ...u8(0x01), // length ...u8(0x02), // data ]; bytes = [ ...u8(0x03), // TAG(3) = Object Descriptor ...variableUnsignedInt(bytes.length), ...bytes, ]; return fullBox('esds', 0, 0, bytes); }; export const wave = (trackData: IsobmffAudioTrackData) => { return box('wave', undefined, [ frma(trackData), enda(trackData), box('\x00\x00\x00\x00'), // NULL tag at the end ]); }; export const frma = (trackData: IsobmffAudioTrackData) => { return box('frma', [ ascii(audioCodecToBoxName(trackData.track.source._codec, trackData.muxer.isQuickTime)), ]); }; // This box specifies PCM endianness export const enda = (trackData: IsobmffAudioTrackData) => { const { littleEndian } = parsePcmCodec(trackData.track.source._codec as PcmAudioCodec); return box('enda', [ u16(+littleEndian), ]); }; /** Opus Specific Box. */ export const dOps = (trackData: IsobmffAudioTrackData) => { let outputChannelCount = trackData.info.numberOfChannels; // Default PreSkip, should be at least 80 milliseconds worth of playback, measured in 48000 Hz samples let preSkip = 3840; let inputSampleRate = trackData.info.sampleRate; let outputGain = 0; let channelMappingFamily = 0; let channelMappingTable: Uint8Array = new Uint8Array(0); // Read preskip and from codec private data from the encoder // https://www.rfc-editor.org/rfc/rfc7845#section-5 const description = trackData.info.decoderConfig?.description; if (description) { assert(description.byteLength >= 18); const bytes = toUint8Array(description); const header = parseOpusIdentificationHeader(bytes); outputChannelCount = header.outputChannelCount; preSkip = header.preSkip; inputSampleRate = header.inputSampleRate; outputGain = header.outputGain; channelMappingFamily = header.channelMappingFamily; if (header.channelMappingTable) { channelMappingTable = header.channelMappingTable; } } // https://www.opus-codec.org/docs/opus_in_isobmff.html return box('dOps', [ u8(0), // Version u8(outputChannelCount), // OutputChannelCount u16(preSkip), // PreSkip u32(inputSampleRate), // InputSampleRate i16(outputGain), // OutputGain u8(channelMappingFamily), // ChannelMappingFamily ...channelMappingTable, ]); }; /** FLAC specific box. */ export const dfLa = (trackData: IsobmffAudioTrackData) => { const description = trackData.info.decoderConfig?.description; assert(description); const bytes = toUint8Array(description); return fullBox('dfLa', 0, 0, [ ...bytes.subarray(4), ]); }; /** PCM Configuration Box, ISO/IEC 23003-5. */ const pcmC = (trackData: IsobmffAudioTrackData) => { const { littleEndian, sampleSize } = parsePcmCodec(trackData.track.source._codec as PcmAudioCodec); const formatFlags = +littleEndian; return fullBox('pcmC', 0, 0, [ u8(formatFlags), u8(8 * sampleSize), ]); }; /** AC3SpecificBox */ const dac3 = (trackData: IsobmffAudioTrackData) => { const frameInfo = parseAc3SyncFrame(trackData.info.firstPacket.data); if (!frameInfo) { throw new Error( 'Couldn\'t extract AC-3 frame info from the audio packet. ' + 'Ensure the packets contain valid AC-3 sync frames (as specified in ETSI TS 102 366).', ); } const bytes = new Uint8Array(3); const bitstream = new Bitstream(bytes); bitstream.writeBits(2, frameInfo.fscod); bitstream.writeBits(5, frameInfo.bsid); bitstream.writeBits(3, frameInfo.bsmod); bitstream.writeBits(3, frameInfo.acmod); bitstream.writeBits(1, frameInfo.lfeon); bitstream.writeBits(5, frameInfo.bitRateCode); bitstream.writeBits(5, 0); // reserved return box('dac3', [...bytes]); }; /** EC3SpecificBox */ const dec3 = (trackData: IsobmffAudioTrackData) => { const frameInfo = parseEac3SyncFrame(trackData.info.firstPacket.data); if (!frameInfo) { throw new Error( 'Couldn\'t extract E-AC-3 frame info from the audio packet. ' + 'Ensure the packets contain valid E-AC-3 sync frames (as specified in ETSI TS 102 366).', ); } // Calculate size let totalBits = 16; // header: data_rate (13) + num_ind_sub (3) for (const sub of frameInfo.substreams) { totalBits += 23; // fixed fields per substream if (sub.numDepSub > 0) { totalBits += 9; // chan_loc } else { totalBits += 1; // reserved } } const size = Math.ceil(totalBits / 8); const bytes = new Uint8Array(size); const bitstream = new Bitstream(bytes); bitstream.writeBits(13, frameInfo.dataRate); bitstream.writeBits(3, frameInfo.substreams.length - 1); // num_ind_sub for (const sub of frameInfo.substreams) { bitstream.writeBits(2, sub.fscod); bitstream.writeBits(5, sub.bsid); bitstream.writeBits(1, 0); // reserved bitstream.writeBits(1, 0); // asvc = 0 bitstream.writeBits(3, sub.bsmod); bitstream.writeBits(3, sub.acmod); bitstream.writeBits(1, sub.lfeon); bitstream.writeBits(3, 0); // reserved bitstream.writeBits(4, sub.numDepSub); if (sub.numDepSub > 0) { bitstream.writeBits(9, sub.chanLoc); } else { bitstream.writeBits(1, 0); // reserved } } return box('dec3', [...bytes]); }; export const subtitleSampleDescription = ( compressionType: string, trackData: IsobmffSubtitleTrackData, ) => box(compressionType, [ Array(6).fill(0), // Reserved u16(1), // Data reference index ], [ SUBTITLE_CODEC_TO_CONFIGURATION_BOX[trackData.track.source._codec](trackData), ]); export const vttC = (trackData: IsobmffSubtitleTrackData) => box('vttC', [ ...textEncoder.encode(trackData.info.config.description), ]); export const txtC = (textConfig: Uint8Array) => fullBox('txtC', 0, 0, [ ...textConfig, 0, // Text config (null-terminated) ]); /** * Time-To-Sample Box: Stores duration information for a media's samples, providing a mapping from a time in a media * to the corresponding data sample. The table is compact, meaning that consecutive samples with the same time delta * will be grouped. */ export const stts = (trackData: IsobmffTrackData) => { return fullBox('stts', 0, 0, [ u32(trackData.timeToSampleTable.length), // Number of entries trackData.timeToSampleTable.map(x => [ // Time-to-sample table u32(x.sampleCount), // Sample count u32(x.sampleDelta), // Sample duration ]), ]); }; /** Sync Sample Box: Identifies the key frames in the media, marking the random access points within a stream. */ export const stss = (trackData: IsobmffTrackData) => { if (trackData.samples.every(x => x.type === 'key')) return null; // No stss box -> every frame is a key frame const keySamples = [...trackData.samples.entries()].filter(([, sample]) => sample.type === 'key'); return fullBox('stss', 0, 0, [ u32(keySamples.length), // Number of entries keySamples.map(([index]) => u32(index + 1)), // Sync sample table ]); }; /** * Sample-To-Chunk Box: As samples are added to a media, they are collected into chunks that allow optimized data * access. A chunk contains one or more samples. Chunks in a media may have different sizes, and the samples within a * chunk may have different sizes. The Sample-To-Chunk Box stores chunk information for the samples in a media, stored * in a compactly-coded fashion. */ export const stsc = (trackData: IsobmffTrackData) => { return fullBox('stsc', 0, 0, [ u32(trackData.compactlyCodedChunkTable.length), // Number of entries trackData.compactlyCodedChunkTable.map(x => [ // Sample-to-chunk table u32(x.firstChunk), // First chunk u32(x.samplesPerChunk), // Samples per chunk u32(1), // Sample description index ]), ]); }; /** Sample Size Box: Specifies the byte size of each sample in the media. */ export const stsz = (trackData: IsobmffTrackData) => { if (trackData.type === 'audio' && trackData.info.requiresPcmTransformation) { const { sampleSize } = parsePcmCodec(trackData.track.source._codec as PcmAudioCodec); // With PCM, every sample has the same size return fullBox('stsz', 0, 0, [ u32(sampleSize * trackData.info.numberOfChannels), // Sample size u32(trackData.samples.reduce((acc, x) => acc + intoTimescale(x.duration, trackData.timescale), 0)), ]); } return fullBox('stsz', 0, 0, [ u32(0), // Sample size (0 means non-constant size) u32(trackData.samples.length), // Number of entries trackData.samples.map(x => u32(x.size)), // Sample size table ]); }; /** Chunk Offset Box: Identifies the location of each chunk of data in the media's data stream, relative to the file. */ export const stco = (trackData: IsobmffTrackData) => { if (trackData.finalizedChunks.length > 0 && last(trackData.finalizedChunks)!.offset! >= 2 ** 32) { // If the file is large, use the co64 box return fullBox('co64', 0, 0, [ u32(trackData.finalizedChunks.length), // Number of entries trackData.finalizedChunks.map(x => u64(x.offset!)), // Chunk offset table ]); } return fullBox('stco', 0, 0, [ u32(trackData.finalizedChunks.length), // Number of entries trackData.finalizedChunks.map(x => u32(x.offset!)), // Chunk offset table ]); }; /** * Composition Time to Sample Box: Stores composition time offset information (PTS-DTS) for a * media's samples. The table is compact, meaning that consecutive samples with the same time * composition time offset will be grouped. */ export const ctts = (trackData: IsobmffTrackData) => { return fullBox('ctts', 1, 0, [ u32(trackData.compositionTimeOffsetTable.length), // Number of entries trackData.compositionTimeOffsetTable.map(x => [ // Time-to-sample table u32(x.sampleCount), // Sample count i32(x.sampleCompositionTimeOffset), // Sample offset ]), ]); }; /** * Composition to Decode Box: Stores information about the composition and display times of the media samples. */ export const cslg = (trackData: IsobmffTrackData) => { let leastDecodeToDisplayDelta = Infinity; let greatestDecodeToDisplayDelta = -Infinity; let compositionStartTime = Infinity; let compositionEndTime = -Infinity; assert(trackData.compositionTimeOffsetTable.length > 0); assert(trackData.samples.length > 0); for (let i = 0; i < trackData.compositionTimeOffsetTable.length; i++) { const entry = trackData.compositionTimeOffsetTable[i]!; leastDecodeToDisplayDelta = Math.min(leastDecodeToDisplayDelta, entry.sampleCompositionTimeOffset); greatestDecodeToDisplayDelta = Math.max(greatestDecodeToDisplayDelta, entry.sampleCompositionTimeOffset); } for (let i = 0; i < trackData.samples.length; i++) { const sample = trackData.samples[i]!; compositionStartTime = Math.min( compositionStartTime, intoTimescale(sample.timestamp, trackData.timescale), ); compositionEndTime = Math.max( compositionEndTime, intoTimescale(sample.timestamp + sample.duration, trackData.timescale), ); } const compositionToDtsShift = Math.max(-leastDecodeToDisplayDelta, 0); if (compositionEndTime >= 2 ** 31) { // For very large files, the composition end time can't be represented in i32, so let's just scrap the box in // that case. QuickTime fails to read the file if there's a cslg box with version 1, so that's sadly not an // option. return null; } return fullBox('cslg', 0, 0, [ i32(compositionToDtsShift), // Composition to DTS shift i32(leastDecodeToDisplayDelta), // Least decode to display delta i32(greatestDecodeToDisplayDelta), // Greatest decode to display delta i32(compositionStartTime), // Composition start time i32(compositionEndTime), // Composition end time ]); }; /** * Movie Extends Box: This box signals to readers that the file is fragmented. Contains a single Track Extends Box * for each track in the movie. */ export const mvex = (trackDatas: IsobmffTrackData[]) => { return box('mvex', undefined, trackDatas.map(trex)); }; /** Track Extends Box: Contains the default values used by the movie fragments. */ export const trex = (trackData: IsobmffTrackData) => { return fullBox('trex', 0, 0, [ u32(trackData.track.id), // Track ID u32(1), // Default sample description index u32(0), // Default sample duration u32(0), // Default sample size u32(0), // Default sample flags ]); }; /** * Movie Fragment Box: The movie fragments extend the presentation in time. They provide the information that would * previously have been in the Movie Box. */ export const moof = (sequenceNumber: number, trackDatas: IsobmffTrackData[]) => { return box('moof', undefined, [ mfhd(sequenceNumber), ...trackDatas.map(traf), ]); }; /** Movie Fragment Header Box: Contains a sequence number as a safety check. */ export const mfhd = (sequenceNumber: number) => { return fullBox('mfhd', 0, 0, [ u32(sequenceNumber), // Sequence number ]); }; const fragmentSampleFlags = (sample: Sample) => { let byte1 = 0; let byte2 = 0; const byte3 = 0; const byte4 = 0; const sampleIsDifferenceSample = sample.type === 'delta'; byte2 |= +sampleIsDifferenceSample; if (sampleIsDifferenceSample) { byte1 |= 1; // There is redundant coding in this sample } else { byte1 |= 2; // There is no redundant coding in this sample } // Note that there are a lot of other flags to potentially set here, but most are irrelevant / non-necessary return byte1 << 24 | byte2 << 16 | byte3 << 8 | byte4; }; /** Track Fragment Box */ export const traf = (trackData: IsobmffTrackData) => { return box('traf', undefined, [ tfhd(trackData), tfdt(trackData), trun(trackData), ]); }; /** Track Fragment Header Box: Provides a reference to the extended track, and flags. */ export const tfhd = (trackData: IsobmffTrackData) => { assert(trackData.currentChunk); let tfFlags = 0; tfFlags |= 0x00008; // Default sample duration present tfFlags |= 0x00010; // Default sample size present tfFlags |= 0x00020; // Default sample flags present tfFlags |= 0x20000; // Default base is moof // Prefer the second sample over the first one, as the first one is a sync sample and therefore the "odd one out" const referenceSample = trackData.currentChunk.samples[1] ?? trackData.currentChunk.samples[0]!; const referenceSampleInfo = { duration: referenceSample.timescaleUnitsToNextSample, size: referenceSample.size, flags: fragmentSampleFlags(referenceSample), }; return fullBox('tfhd', 0, tfFlags, [ u32(trackData.track.id), // Track ID u32(referenceSampleInfo.duration), // Default sample duration u32(referenceSampleInfo.size), // Default sample size u32(referenceSampleInfo.flags), // Default sample flags ]); }; /** * Track Fragment Decode Time Box: Provides the absolute decode time of the first sample of the fragment. This is * useful for performing random access on the media file. */ export const tfdt = (trackData: IsobmffTrackData) => { assert(trackData.currentChunk); return fullBox('tfdt', 1, 0, [ u64(intoTimescale(trackData.currentChunk.startTimestamp, trackData.timescale)), // Base Media Decode Time ]); }; /** Track Run Box: Specifies a run of contiguous samples for a given track. */ export const trun = (trackData: IsobmffTrackData) => { assert(trackData.currentChunk); const allSampleDurations = trackData.currentChunk.samples.map(x => x.timescaleUnitsToNextSample); const allSampleSizes = trackData.currentChunk.samples.map(x => x.size); const allSampleFlags = trackData.currentChunk.samples.map(fragmentSampleFlags); const allSampleCompositionTimeOffsets = trackData.currentChunk.samples .map(x => intoTimescale(x.timestamp - x.decodeTimestamp, trackData.timescale)); const uniqueSampleDurations = new Set(allSampleDurations); const uniqueSampleSizes = new Set(allSampleSizes); const uniqueSampleFlags = new Set(allSampleFlags); const uniqueSampleCompositionTimeOffsets = new Set(allSampleCompositionTimeOffsets); const firstSampleFlagsPresent = uniqueSampleFlags.size === 2 && allSampleFlags[0] !== allSampleFlags[1]; const sampleDurationPresent = uniqueSampleDurations.size > 1; const sampleSizePresent = uniqueSampleSizes.size > 1; const sampleFlagsPresent = !firstSampleFlagsPresent && uniqueSampleFlags.size > 1; const sampleCompositionTimeOffsetsPresent = uniqueSampleCompositionTimeOffsets.size > 1 || [...uniqueSampleCompositionTimeOffsets].some(x => x !== 0); let flags = 0; flags |= 0x0001; // Data offset present flags |= 0x0004 * +firstSampleFlagsPresent; // First sample flags present flags |= 0x0100 * +sampleDurationPresent; // Sample duration present flags |= 0x0200 * +sampleSizePresent; // Sample size present flags |= 0x0400 * +sampleFlagsPresent; // Sample flags present flags |= 0x0800 * +sampleCompositionTimeOffsetsPresent; // Sample composition time offsets present return fullBox('trun', 1, flags, [ u32(trackData.currentChunk.samples.length), // Sample count u32(trackData.currentChunk.offset! - trackData.currentChunk.moofOffset! || 0), // Data offset firstSampleFlagsPresent ? u32(allSampleFlags[0]!) : [], trackData.currentChunk.samples.map((_, i) => [ sampleDurationPresent ? u32(allSampleDurations[i]!) : [], // Sample duration sampleSizePresent ? u32(allSampleSizes[i]!) : [], // Sample size sampleFlagsPresent ? u32(allSampleFlags[i]!) : [], // Sample flags // Sample composition time offsets sampleCompositionTimeOffsetsPresent ? i32(allSampleCompositionTimeOffsets[i]!) : [], ]), ]); }; /** * Movie Fragment Random Access Box: For each track, provides pointers to sync samples within the file * for random access. */ export const mfra = (trackDatas: IsobmffTrackData[]) => { return box('mfra', undefined, [ ...trackDatas.map(tfra), mfro(), ]); }; /** Track Fragment Random Access Box: Provides pointers to sync samples within the file for random access. */ export const tfra = (trackData: IsobmffTrackData, trackIndex: number) => { const version = 1; // Using this version allows us to use 64-bit time and offset values return fullBox('tfra', version, 0, [ u32(trackData.track.id), // Track ID u32(0b111111), // This specifies that traf number, trun number and sample number are 32-bit ints u32(trackData.finalizedChunks.length), // Number of entries trackData.finalizedChunks.map(chunk => [ u64(intoTimescale(chunk.samples[0]!.timestamp, trackData.timescale)), // Time (in presentation time) u64(chunk.moofOffset!), // moof offset u32(trackIndex + 1), // traf number u32(1), // trun number u32(1), // Sample number ]), ]); }; /** * Movie Fragment Random Access Offset Box: Provides the size of the enclosing mfra box. This box can be used by readers * to quickly locate the mfra box by searching from the end of the file. */ export const mfro = () => { return fullBox('mfro', 0, 0, [ // This value needs to be overwritten manually from the outside, where the actual size of the enclosing mfra box // is known u32(0), // Size ]); }; /** VTT Empty Cue Box */ export const vtte = () => box('vtte'); /** VTT Cue Box */ export const vttc = ( payload: string, timestamp: number | null, identifier: string | null, settings: string | null, sourceId: number | null, ) => box('vttc', undefined, [ sourceId !== null ? box('vsid', [i32(sourceId)]) : null, identifier !== null ? box('iden', [...textEncoder.encode(identifier)]) : null, timestamp !== null ? box('ctim', [...textEncoder.encode(formatSubtitleTimestamp(timestamp))]) : null, settings !== null ? box('sttg', [...textEncoder.encode(settings)]) : null, box('payl', [...textEncoder.encode(payload)]), ]); /** VTT Additional Text Box */ export const vtta = (notes: string) => box('vtta', [...textEncoder.encode(notes)]); /** User Data Box */ const udta = (muxer: IsobmffMuxer) => { const boxes: Box[] = []; const metadataFormat = muxer.format._options.metadataFormat ?? 'auto'; const metadataTags = muxer.output._metadataTags; // Depending on the format, metadata tags are written differently if (metadataFormat === 'mdir' || (metadataFormat === 'auto' && !muxer.isQuickTime)) { const metaBox = metaMdir(metadataTags); if (metaBox) boxes.push(metaBox); } else if (metadataFormat === 'mdta') { const metaBox = metaMdta(metadataTags); if (metaBox) boxes.push(metaBox); } else if (metadataFormat === 'udta' || (metadataFormat === 'auto' && muxer.isQuickTime)) { addQuickTimeMetadataTagBoxes(boxes, muxer.output._metadataTags); } if (boxes.length === 0) { return null; } return box('udta', undefined, boxes); }; const addQuickTimeMetadataTagBoxes = (boxes: Box[], tags: MetadataTags) => { // https://exiftool.org/TagNames/QuickTime.html (QuickTime UserData Tags) // For QuickTime files, metadata tags are dumped into the udta box for (const { key, value } of keyValueIterator(tags)) { switch (key) { case 'title': { boxes.push(metadataTagStringBoxShort('©nam', value)); }; break; case 'description': { boxes.push(metadataTagStringBoxShort('©des', value)); }; break; case 'artist': { boxes.push(metadataTagStringBoxShort('©ART', value)); }; break; case 'album': { boxes.push(metadataTagStringBoxShort('©alb', value)); }; break; case 'albumArtist': { boxes.push(metadataTagStringBoxShort('albr', value)); }; break; case 'genre': { boxes.push(metadataTagStringBoxShort('©gen', value)); }; break; case 'date': { boxes.push(metadataTagStringBoxShort('©day', value.toISOString().slice(0, 10))); }; break; case 'comment': { boxes.push(metadataTagStringBoxShort('©cmt', value)); }; break; case 'lyrics': { boxes.push(metadataTagStringBoxShort('©lyr', value)); }; break; case 'raw': { // Handled later }; break; case 'discNumber': case 'discsTotal': case 'trackNumber': case 'tracksTotal': case 'images': { // Not written for QuickTime (common Apple L) }; break; default: assertNever(key); } } if (tags.raw) { for (const key in tags.raw) { const value = tags.raw[key]; if (value == null || key.length !== 4 || boxes.some(x => x.type === key)) { continue; } if (typeof value === 'string') { boxes.push(metadataTagStringBoxShort(key, value)); } else if (value instanceof Uint8Array) { boxes.push(box(key, Array.from(value))); } } } }; const metadataTagStringBoxShort = (name: string, value: string) => { const encoded = textEncoder.encode(value); return box(name, [ u16(encoded.length), u16(getLanguageCodeInt('und')), Array.from(encoded), ]); }; const DATA_BOX_MIME_TYPE_MAP: Record = { 'image/jpeg': 13, 'image/png': 14, 'image/bmp': 27, }; /** * Generates key-value metadata for inclusion in the "meta" box. */ const generateMetadataPairs = (tags: MetadataTags, isMdta: boolean) => { const pairs: { key: string; value: Box; }[] = []; // https://exiftool.org/TagNames/QuickTime.html (QuickTime ItemList Tags) // This is the metadata format used for MP4 files for (const { key, value } of keyValueIterator(tags)) { switch (key) { case 'title': { pairs.push({ key: isMdta ? 'title' : '©nam', value: dataStringBoxLong(value) }); }; break; case 'description': { pairs.push({ key: isMdta ? 'description' : '©des', value: dataStringBoxLong(value) }); }; break; case 'artist': { pairs.push({ key: isMdta ? 'artist' : '©ART', value: dataStringBoxLong(value) }); }; break; case 'album': { pairs.push({ key: isMdta ? 'album' : '©alb', value: dataStringBoxLong(value) }); }; break; case 'albumArtist': { pairs.push({ key: isMdta ? 'album_artist' : 'aART', value: dataStringBoxLong(value) }); }; break; case 'comment': { pairs.push({ key: isMdta ? 'comment' : '©cmt', value: dataStringBoxLong(value) }); }; break; case 'genre': { pairs.push({ key: isMdta ? 'genre' : '©gen', value: dataStringBoxLong(value) }); }; break; case 'lyrics': { pairs.push({ key: isMdta ? 'lyrics' : '©lyr', value: dataStringBoxLong(value) }); }; break; case 'date': { pairs.push({ key: isMdta ? 'date' : '©day', value: dataStringBoxLong(value.toISOString().slice(0, 10)), }); }; break; case 'images': { for (const image of value) { if (image.kind !== 'coverFront') { continue; } pairs.push({ key: 'covr', value: box('data', [ u32(DATA_BOX_MIME_TYPE_MAP[image.mimeType] ?? 0), // Type indicator u32(0), // Locale indicator Array.from(image.data), // Kinda slow, hopefully temp ]) }); } }; break; case 'trackNumber': { if (isMdta) { const string = tags.tracksTotal !== undefined ? `${value}/${tags.tracksTotal}` : value.toString(); pairs.push({ key: 'track', value: dataStringBoxLong(string) }); } else { pairs.push({ key: 'trkn', value: box('data', [ u32(0), // 8 bytes empty u32(0), u16(0), // Empty u16(value), u16(tags.tracksTotal ?? 0), u16(0), // Empty ]) }); } }; break; case 'discNumber': { if (!isMdta) { // Only written for mdir pairs.push({ key: 'disc', value: box('data', [ u32(0), // 8 bytes empty u32(0), u16(0), // Empty u16(value), u16(tags.discsTotal ?? 0), u16(0), // Empty ]) }); } }; break; case 'tracksTotal': case 'discsTotal':{ // These are included with 'trackNumber' and 'discNumber' respectively }; break; case 'raw': { // Handled later }; break; default: assertNever(key); } } if (tags.raw) { for (const key in tags.raw) { const value = tags.raw[key]; if (value == null || (!isMdta && key.length !== 4) || pairs.some(x => x.key === key)) { continue; } if (typeof value === 'string') { pairs.push({ key, value: dataStringBoxLong(value) }); } else if (value instanceof Uint8Array) { pairs.push({ key, value: box('data', [ u32(0), // Type indicator u32(0), // Locale indicator Array.from(value), ]) }); } else if (value instanceof RichImageData) { pairs.push({ key, value: box('data', [ u32(DATA_BOX_MIME_TYPE_MAP[value.mimeType] ?? 0), // Type indicator u32(0), // Locale indicator Array.from(value.data), // Kinda slow, hopefully temp ]) }); } } } return pairs; }; /** Metadata Box (mdir format) */ const metaMdir = (tags: MetadataTags) => { const pairs = generateMetadataPairs(tags, false); if (pairs.length === 0) { return null; } // fullBox format return fullBox('meta', 0, 0, undefined, [ hdlr(false, 'mdir', '', 'appl'), // mdir handler box('ilst', undefined, pairs.map(pair => box(pair.key, undefined, [pair.value]))), // Item list without keys box ]); }; /** Metadata Box (mdta format with keys box) */ const metaMdta = (tags: MetadataTags) => { const pairs = generateMetadataPairs(tags, true); if (pairs.length === 0) { return null; } // box without version and flags return box('meta', undefined, [ hdlr(false, 'mdta', ''), // mdta handler fullBox('keys', 0, 0, [ u32(pairs.length), ], pairs.map(pair => box('mdta', [ // Hacky since these aren't boxes technically, but if not box why box-shaped? ...textEncoder.encode(pair.key), ]))), box('ilst', undefined, pairs.map((pair, i) => { const boxName = String.fromCharCode(...u32(i + 1)); return box(boxName, undefined, [pair.value]); })), ]); }; const dataStringBoxLong = (value: string) => { return box('data', [ u32(1), // Type indicator (UTF-8) u32(0), // Locale indicator ...textEncoder.encode(value), ]); }; const videoCodecToBoxName = (codec: VideoCodec, fullCodecString: string) => { switch (codec) { case 'avc': return fullCodecString.startsWith('avc3') ? 'avc3' : 'avc1'; case 'hevc': return 'hvc1'; case 'vp8': return 'vp08'; case 'vp9': return 'vp09'; case 'av1': return 'av01'; } }; const VIDEO_CODEC_TO_CONFIGURATION_BOX: Record Box | null> = { avc: avcC, hevc: hvcC, vp8: vpcC, vp9: vpcC, av1: av1C, }; const audioCodecToBoxName = (codec: AudioCodec, isQuickTime: boolean): string => { switch (codec) { case 'aac': return 'mp4a'; case 'mp3': return 'mp4a'; case 'opus': return 'Opus'; case 'vorbis': return 'mp4a'; case 'flac': return 'fLaC'; case 'ulaw': return 'ulaw'; case 'alaw': return 'alaw'; case 'pcm-u8': return 'raw '; case 'pcm-s8': return 'sowt'; case 'ac3': return 'ac-3'; case 'eac3': return 'ec-3'; } // Logic diverges here if (isQuickTime) { switch (codec) { case 'pcm-s16': return 'sowt'; case 'pcm-s16be': return 'twos'; case 'pcm-s24': return 'in24'; case 'pcm-s24be': return 'in24'; case 'pcm-s32': return 'in32'; case 'pcm-s32be': return 'in32'; case 'pcm-f32': return 'fl32'; case 'pcm-f32be': return 'fl32'; case 'pcm-f64': return 'fl64'; case 'pcm-f64be': return 'fl64'; } } else { switch (codec) { case 'pcm-s16': return 'ipcm'; case 'pcm-s16be': return 'ipcm'; case 'pcm-s24': return 'ipcm'; case 'pcm-s24be': return 'ipcm'; case 'pcm-s32': return 'ipcm'; case 'pcm-s32be': return 'ipcm'; case 'pcm-f32': return 'fpcm'; case 'pcm-f32be': return 'fpcm'; case 'pcm-f64': return 'fpcm'; case 'pcm-f64be': return 'fpcm'; } } }; const audioCodecToConfigurationBox = (codec: AudioCodec, isQuickTime: boolean) => { switch (codec) { case 'aac': return esds; case 'mp3': return esds; case 'opus': return dOps; case 'vorbis': return esds; case 'flac': return dfLa; case 'ac3': return dac3; case 'eac3': return dec3; } // Logic diverges here if (isQuickTime) { switch (codec) { case 'pcm-s24': return wave; case 'pcm-s24be': return wave; case 'pcm-s32': return wave; case 'pcm-s32be': return wave; case 'pcm-f32': return wave; case 'pcm-f32be': return wave; case 'pcm-f64': return wave; case 'pcm-f64be': return wave; } } else { switch (codec) { case 'pcm-s16': return pcmC; case 'pcm-s16be': return pcmC; case 'pcm-s24': return pcmC; case 'pcm-s24be': return pcmC; case 'pcm-s32': return pcmC; case 'pcm-s32be': return pcmC; case 'pcm-f32': return pcmC; case 'pcm-f32be': return pcmC; case 'pcm-f64': return pcmC; case 'pcm-f64be': return pcmC; } } return null; }; const SUBTITLE_CODEC_TO_BOX_NAME: Record = { webvtt: 'wvtt', }; const SUBTITLE_CODEC_TO_CONFIGURATION_BOX: Record< SubtitleCodec, (trackData: IsobmffSubtitleTrackData) => Box | null > = { webvtt: vttC, }; const getLanguageCodeInt = (code: string) => { assert(code.length === 3); ; let language = 0; for (let i = 0; i < 3; i++) { language <<= 5; language += code.charCodeAt(i) - 0x60; } return language; }; ===== src/isobmff/isobmff-reader.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { RichImageData } from '../metadata'; import { textDecoder } from '../misc'; import { FileSlice, readAscii, readBytes, readI32Be, readU16Be, readU32Be, readU64Be, readU8 } from '../reader'; export const MIN_BOX_HEADER_SIZE = 8; export const MAX_BOX_HEADER_SIZE = 16; export const readBoxHeader = (slice: FileSlice) => { let totalSize = readU32Be(slice); const name = readAscii(slice, 4); let headerSize = 8; const hasLargeSize = totalSize === 1; if (hasLargeSize) { totalSize = readU64Be(slice); headerSize = 16; } const contentSize = totalSize - headerSize; if (contentSize < 0) { return null; // Hardly a box is it } return { name, totalSize, headerSize, contentSize }; }; export const readFixed_16_16 = (slice: FileSlice) => { return readI32Be(slice) / 0x10000; }; export const readFixed_2_30 = (slice: FileSlice) => { return readI32Be(slice) / 0x40000000; }; export const readIsomVariableInteger = (slice: FileSlice) => { let result = 0; for (let i = 0; i < 4; i++) { result <<= 7; const nextByte = readU8(slice); result |= nextByte & 0x7f; if ((nextByte & 0x80) === 0) { break; } } return result; }; export const readMetadataStringShort = (slice: FileSlice) => { let stringLength = readU16Be(slice); slice.skip(2); // Language stringLength = Math.min(stringLength, slice.remainingLength); return textDecoder.decode(readBytes(slice, stringLength)); }; export const readDataBox = (slice: FileSlice) => { const header = readBoxHeader(slice); if (!header || header.name !== 'data') { return null; } if (slice.remainingLength < 8) { // Box is too small return null; } const typeIndicator = readU32Be(slice); slice.skip(4); // Locale indicator const data = readBytes(slice, header.contentSize - 8); switch (typeIndicator) { case 1: return textDecoder.decode(data); // UTF-8 case 2: return new TextDecoder('utf-16be').decode(data); // UTF-16-BE case 13: return new RichImageData(data, 'image/jpeg'); case 14: return new RichImageData(data, 'image/png'); case 27: return new RichImageData(data, 'image/bmp'); default: return data; } }; ===== src/node.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ // This file contains Node.js-specific code that does not run in a browser. export * as fs from 'node:fs/promises'; ===== src/segmented-input.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { TrackType } from './output'; import { MediaCodec } from './codec'; import { DurationMetadataRequestOptions } from './demuxer'; import { Input } from './input'; import { InputAudioTrack, InputAudioTrackBacking, InputTrack, InputTrackBacking, InputVideoTrack, InputVideoTrackBacking, } from './input-track'; import { PacketRetrievalOptions } from './media-sink'; import { arrayCount, assert, MaybePromise, roundToDivisor } from './misc'; import { EncodedPacket } from './packet'; export type SegmentedInputMetadata = { name: string | null; bitrate: number | null; // doc block: this refers to the _peak_ bitrate averageBitrate: number | null; codecs: MediaCodec[]; codecStrings: string[]; resolution: { width: number; height: number } | null; frameRate: number | null; isKeyFrameOnly: boolean; }; export type AssociatedGroup = { id: string; type: 'video' | 'audio' | 'subtitles' | 'closed-captions'; }; export type Segment = { timestamp: number; duration: number; /** * The Unix time (in seconds) corresponding to this segment's start timestamp, or null if unknown. This is computed * whenever the source provides wall-clock information (e.g. HLS program date time), even if the segment timestamps * themselves are not shifted into Unix time space. */ unixEpochTimestamp: number | null; firstSegment: Segment | null; }; export type SegmentRetrievalOptions = { skipLiveWait?: boolean; }; export type SegmentedInputTrackDeclaration = { id: number; type: TrackType; }; export abstract class SegmentedInput { input: Input; path: string; trackDeclarations: SegmentedInputTrackDeclaration[] | null; nextInputCacheAge = 0; inputCache: { segment: Segment; input: Input; age: number; }[] = []; trackBackingsPromise: Promise | null = null; firstSegment: Segment | null = null; firstSegmentFirstTimestamps = new WeakMap(); firstTimestampCache = new WeakMap(); constructor(input: Input, path: string, trackDeclarations: SegmentedInputTrackDeclaration[] | null) { this.input = input; this.path = path; this.trackDeclarations = trackDeclarations; } abstract getFirstSegment(options: SegmentRetrievalOptions): Promise; abstract getSegmentAt(timestamp: number, options: SegmentRetrievalOptions): Promise; abstract getNextSegment(segment: Segment, options: SegmentRetrievalOptions): Promise; abstract getPreviousSegment(segment: Segment, options: SegmentRetrievalOptions): Promise; abstract getInputForSegment(segment: Segment): Input; abstract getLiveRefreshInterval(): Promise; async getDurationFromMetadata(options: DurationMetadataRequestOptions) { const lastSegment = await this.getSegmentAt(Infinity, { skipLiveWait: options.skipLiveWait, }); if (!lastSegment) { return null; } return lastSegment.timestamp + lastSegment.duration; } async getUnixTimeForTimestamp(timestamp: number): Promise { let segment = await this.getSegmentAt(timestamp, {}); segment ??= await this.getFirstSegment({}); if (!segment || segment.unixEpochTimestamp === null) { return null; } const elapsed = timestamp - segment.timestamp; return segment.unixEpochTimestamp + elapsed; } async getTrackBackings(): Promise { return this.trackBackingsPromise ??= (async () => { const backings: InputTrackBacking[] = []; if (this.trackDeclarations) { for (const decl of this.trackDeclarations) { if (decl.type === 'video') { const number = arrayCount(backings, x => x.getType() === 'video') + 1; backings.push( new SegmentedInputInputVideoTrackBacking(this, decl, number), ); } else if (decl.type === 'audio') { const number = arrayCount(backings, x => x.getType() === 'audio') + 1; backings.push( new SegmentedInputInputAudioTrackBacking(this, decl, number), ); } } } else { // There are no declarations, we must determine the tracks from the first segment this.firstSegment = await this.getFirstSegment({}); if (!this.firstSegment) { return []; } const input = this.getInputForSegment(this.firstSegment); const inputTracks = await input.getTracks(); for (const track of inputTracks) { if (track.type === 'video') { const number = arrayCount(backings, x => x.getType() === 'video') + 1; backings.push( new SegmentedInputInputVideoTrackBacking(this, { id: backings.length + 1, type: 'video', }, number), ); } else if (track.type === 'audio') { const number = arrayCount(backings, x => x.getType() === 'audio') + 1; backings.push( new SegmentedInputInputAudioTrackBacking(this, { id: backings.length + 1, type: 'audio', }, number), ); } } } return backings; })(); } // This operation is done a lot and can be semi-expensive, so it's good to have a cache for it async getFirstTimestampForInput(input: Input) { const existing = this.firstTimestampCache.get(input); if (existing !== undefined) { return existing; } const firstTimestamp = await input.getFirstTimestamp(); this.firstTimestampCache.set(input, firstTimestamp); return firstTimestamp; } async getMediaOffset(segment: Segment, input: Input) { const firstSegment = segment.firstSegment ?? segment; let firstSegmentFirstTimestamp: number; if (this.firstSegmentFirstTimestamps.has(firstSegment)) { firstSegmentFirstTimestamp = this.firstSegmentFirstTimestamps.get(firstSegment)!; } else { const firstInput = this.getInputForSegment(firstSegment); firstSegmentFirstTimestamp = await this.getFirstTimestampForInput(firstInput); this.firstSegmentFirstTimestamps.set(firstSegment, firstSegmentFirstTimestamp); } if (firstSegment === segment) { return firstSegment.timestamp - firstSegmentFirstTimestamp; } const segmentFirstTimestamp = await this.getFirstTimestampForInput(input); const segmentElapsed = segment.timestamp - firstSegment.timestamp; const inputElapsed = segmentFirstTimestamp - firstSegmentFirstTimestamp; const difference = inputElapsed - segmentElapsed; if (Math.abs(difference) <= Math.min(0.25, segmentElapsed)) { // Heuristic // We're close enough return firstSegment.timestamp - firstSegmentFirstTimestamp; } else { // Ideally, each segment has absolute timestamps that are relative to some outside clock which is // consistent across segments. This is often the case, but not always. Either the container format used is // not timestamped at all (like ADTS), or the segments are just fucky. In this case, use the segment's // relative timestamp to determine where we are, and completely offset out the segment's input start // timestamp. return segment.timestamp - segmentFirstTimestamp; } } dispose() { for (const entry of this.inputCache) { entry.input.dispose(); } this.inputCache.length = 0; } } type PacketInfo = { segment: Segment; track: InputTrack; sourcePacket: EncodedPacket; }; class SegmentedInputInputTrackBacking implements InputTrackBacking { segmentedInput: SegmentedInput; decl: SegmentedInputTrackDeclaration; number: number; packetInfos = new WeakMap(); hydrationPromise: Promise | null = null; firstInputTrack: InputTrack | null = null; constructor(segmentedInput: SegmentedInput, decl: SegmentedInputTrackDeclaration, number: number) { this.segmentedInput = segmentedInput; this.decl = decl; this.number = number; } hydrate() { return this.hydrationPromise ??= (async () => { this.segmentedInput.firstSegment ??= await this.segmentedInput.getFirstSegment({}); if (!this.segmentedInput.firstSegment) { throw new Error('Missing first segment, can\'t retrieve track.'); } const input = this.segmentedInput.getInputForSegment(this.segmentedInput.firstSegment); const inputTracks = await input.getTracks(); const track = inputTracks.find(x => x.type === this.decl.type && x.number === this.number); if (!track) { throw new Error('No matching track found in underlying media data.'); } this.firstInputTrack = track; })(); } getId(): number { return this.decl.id; } getType(): TrackType { return this.decl.type; } getNumber(): number { return this.number; } /** If the backing track is already present, delegate synchronously; otherwise, hydrate first. */ delegate(fn: () => MaybePromise): MaybePromise { if (this.firstInputTrack) { return fn(); } return this.hydrate().then(fn); } async getDecoderConfig() { return this.delegate(() => this.firstInputTrack!._backing.getDecoderConfig()); } getHasOnlyKeyPackets() { return this.delegate(() => this.firstInputTrack!._backing.getHasOnlyKeyPackets?.() ?? null); } getPairingMask() { return 1n; } getCodec() { return this.delegate(() => this.firstInputTrack!._backing.getCodec()); } getInternalCodecId() { return this.delegate(() => this.firstInputTrack!._backing.getInternalCodecId()); } getDisposition() { return this.delegate(() => this.firstInputTrack!._backing.getDisposition()); } getLanguageCode() { return this.delegate(() => this.firstInputTrack!._backing.getLanguageCode()); } getName() { return this.delegate(() => this.firstInputTrack!._backing.getName()); } getTimeResolution() { return this.delegate(() => this.firstInputTrack!._backing.getTimeResolution()); } async isRelativeToUnixEpoch() { await this.hydrate(); assert(this.segmentedInput.firstSegment); return this.segmentedInput.firstSegment.unixEpochTimestamp === this.segmentedInput.firstSegment.timestamp; } getUnixTimeForTimestamp(timestamp: number) { return this.segmentedInput.getUnixTimeForTimestamp(timestamp); } getBitrate() { return this.delegate(() => this.firstInputTrack!._backing.getBitrate()); } getAverageBitrate() { return this.delegate(() => this.firstInputTrack!._backing.getAverageBitrate()); } getDurationFromMetadata(options: DurationMetadataRequestOptions): Promise { return this.segmentedInput.getDurationFromMetadata(options); } getLiveRefreshInterval(): Promise { return this.segmentedInput.getLiveRefreshInterval(); } async createAdjustedPacket(packet: EncodedPacket, segment: Segment, track: InputTrack) { assert(packet.sequenceNumber >= 0); assert(this.segmentedInput.firstSegment); const mediaOffset = await this.segmentedInput.getMediaOffset(segment, track.input); // If we didn't do this then sequence numbers would exceed Number.MAX_SAFE_INTEGER for Unix-timestamped segments const segmentTimestampRelativeToFirst = segment.timestamp - this.segmentedInput.firstSegment.timestamp; const modified = packet.clone({ timestamp: roundToDivisor( packet.timestamp + mediaOffset, await track.getTimeResolution(), ), // The 1e8 assumes a max of 100 MB per second, highly unlikely to be hit, so this should guarantee // monotonically increasing sequence numbers across segments. sequenceNumber: Math.floor(1e8 * segmentTimestampRelativeToFirst) + packet.sequenceNumber, }); this.packetInfos.set(modified, { segment, track, sourcePacket: packet, }); return modified; } async getFirstPacket(options: PacketRetrievalOptions): Promise { await this.hydrate(); assert(this.segmentedInput.firstSegment); assert(this.firstInputTrack); const packet = await this.firstInputTrack._backing.getFirstPacket(options); if (!packet) { return null; } return this.createAdjustedPacket(packet, this.segmentedInput.firstSegment, this.firstInputTrack); } getNextPacket(packet: EncodedPacket, options: PacketRetrievalOptions): Promise { return this._getNextInternal(packet, options, false); } getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions): Promise { return this._getNextInternal(packet, options, true); } async _getNextInternal( packet: EncodedPacket, options: PacketRetrievalOptions, keyframesOnly: boolean, ): Promise { const info = this.packetInfos.get(packet); if (!info) { throw new Error('Packet was not created from this track.'); } const nextPacket = keyframesOnly ? await info.track._backing.getNextKeyPacket(info.sourcePacket, options) : await info.track._backing.getNextPacket(info.sourcePacket, options); if (nextPacket) { return this.createAdjustedPacket(nextPacket, info.segment, info.track); } let currentSegment: Segment | null = info.segment; while (true) { const nextSegment = await this.segmentedInput.getNextSegment(currentSegment, { skipLiveWait: options.skipLiveWait, }); if (!nextSegment) { return null; } const nextInput = this.segmentedInput.getInputForSegment(nextSegment); const nextTracks = await nextInput.getTracks(); const nextTrack = nextTracks.find(t => t.type === info.track.type && t.number === info.track.number); if (!nextTrack) { currentSegment = nextSegment; continue; } const firstPacket = await nextTrack._backing.getFirstPacket(options); if (!firstPacket) { return null; } return this.createAdjustedPacket(firstPacket, nextSegment, nextTrack); } } getPacket(timestamp: number, options: PacketRetrievalOptions): Promise { return this._getPacketInternal(timestamp, options, false); } getKeyPacket(timestamp: number, options: PacketRetrievalOptions): Promise { return this._getPacketInternal(timestamp, options, true); } async _getPacketInternal( timestamp: number, options: PacketRetrievalOptions, keyframesOnly: boolean, ): Promise { let currentSegment = await this.segmentedInput.getSegmentAt(timestamp, { skipLiveWait: options.skipLiveWait, }); if (!currentSegment) { return null; } await this.hydrate(); while (currentSegment) { const input = this.segmentedInput.getInputForSegment(currentSegment); const tracks = await input.getTracks(); const track = tracks.find(t => ( t.type === this.firstInputTrack!.type && t.number === this.firstInputTrack!.number )); if (!track) { // Search the previous segment currentSegment = await this.segmentedInput.getPreviousSegment(currentSegment, { skipLiveWait: options.skipLiveWait, }); continue; } const mediaOffset = await this.segmentedInput.getMediaOffset(currentSegment, input); const offsetTimestamp = timestamp - mediaOffset; const packet = keyframesOnly ? await track._backing.getKeyPacket(offsetTimestamp, options) : await track._backing.getPacket(offsetTimestamp, options); if (!packet) { // Search the previous segment currentSegment = await this.segmentedInput.getPreviousSegment(currentSegment, { skipLiveWait: options.skipLiveWait, }); continue; } return this.createAdjustedPacket(packet, currentSegment, track); } return null; } } class SegmentedInputInputVideoTrackBacking extends SegmentedInputInputTrackBacking implements InputVideoTrackBacking { override firstInputTrack!: InputVideoTrack | null; override getType() { return 'video' as const; } override getCodec() { return this.delegate(() => this.firstInputTrack!._backing.getCodec()); } getCodedWidth() { return this.delegate(() => this.firstInputTrack!._backing.getCodedWidth()); } getCodedHeight() { return this.delegate(() => this.firstInputTrack!._backing.getCodedHeight()); } getSquarePixelWidth() { return this.delegate(() => this.firstInputTrack!._backing.getSquarePixelWidth()); } getSquarePixelHeight() { return this.delegate(() => this.firstInputTrack!._backing.getSquarePixelHeight()); } getRotation() { return this.delegate(() => this.firstInputTrack!._backing.getRotation()); } async getColorSpace(): Promise { return this.delegate(() => this.firstInputTrack!._backing.getColorSpace()); } async canBeTransparent(): Promise { return this.delegate(() => this.firstInputTrack!._backing.canBeTransparent()); } override async getDecoderConfig(): Promise { return this.delegate(() => this.firstInputTrack!._backing.getDecoderConfig()); } } class SegmentedInputInputAudioTrackBacking extends SegmentedInputInputTrackBacking implements InputAudioTrackBacking { override firstInputTrack!: InputAudioTrack; override getType() { return 'audio' as const; } override getCodec() { return this.delegate(() => this.firstInputTrack._backing.getCodec()); } getNumberOfChannels() { return this.delegate(() => this.firstInputTrack._backing.getNumberOfChannels()); } getSampleRate() { return this.delegate(() => this.firstInputTrack._backing.getSampleRate()); } override async getDecoderConfig(): Promise { return this.delegate(() => this.firstInputTrack._backing.getDecoderConfig()); } } ===== src/packet.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { SECOND_TO_MICROSECOND_FACTOR } from './misc'; export const PLACEHOLDER_DATA = /* #__PURE__ */ new Uint8Array(0); /** * The type of a packet. Key packets can be decoded without previous packets, while delta packets depend on previous * packets. * @group Packets * @public */ export type PacketType = 'key' | 'delta'; /** * Holds additional data accompanying an {@link EncodedPacket}. * @group Packets * @public */ export type EncodedPacketSideData = { /** * An encoded alpha frame, encoded with the same codec as the packet. Typically used for transparent videos, where * the alpha information is stored separately from the color information. */ alpha?: Uint8Array; /** * The actual byte length of the alpha data. This field is useful for metadata-only packets where the * `alpha` field contains no bytes. */ alphaByteLength?: number; }; /** * Represents an encoded chunk of media. Mainly used as an expressive wrapper around WebCodecs API's * [`EncodedVideoChunk`](https://developer.mozilla.org/en-US/docs/Web/API/EncodedVideoChunk) and * [`EncodedAudioChunk`](https://developer.mozilla.org/en-US/docs/Web/API/EncodedAudioChunk), but can also be used * standalone. * @group Packets * @public */ export class EncodedPacket { /** * The actual byte length of the data in this packet. This field is useful for metadata-only packets where the * `data` field contains no bytes. */ readonly byteLength: number; /** Additional data carried with this packet. */ readonly sideData: EncodedPacketSideData; /** Creates a new {@link EncodedPacket} from raw bytes and timing information. */ constructor( /** * The encoded data of this packet. For any given codec, this data must adhere to the format specified in the * Mediabunny Codec Registry. */ public readonly data: Uint8Array, /** The type of this packet. */ public readonly type: PacketType, /** * The presentation timestamp of this packet in seconds. May be negative. Samples with negative end timestamps * should not be presented. */ public readonly timestamp: number, /** The duration of this packet in seconds. */ public readonly duration: number, /** * The sequence number indicates the decode order of the packets. Packet A must be decoded before packet B if A * has a lower sequence number than B. If two packets have the same sequence number, they are the same packet. * Otherwise, sequence numbers are arbitrary and are not guaranteed to have any meaning besides their relative * ordering. Negative sequence numbers mean the sequence number is undefined. */ public readonly sequenceNumber = -1, byteLength?: number, sideData?: EncodedPacketSideData, ) { if (data === PLACEHOLDER_DATA && byteLength === undefined) { throw new Error( 'Internal error: byteLength must be explicitly provided when constructing metadata-only packets.', ); } if (byteLength === undefined) { byteLength = data.byteLength; } if (!(data instanceof Uint8Array)) { throw new TypeError('data must be a Uint8Array.'); } if (type !== 'key' && type !== 'delta') { throw new TypeError('type must be either "key" or "delta".'); } if (!Number.isFinite(timestamp)) { throw new TypeError('timestamp must be a number.'); } if (!Number.isFinite(duration) || duration < 0) { throw new TypeError('duration must be a non-negative number.'); } if (!Number.isFinite(sequenceNumber)) { throw new TypeError('sequenceNumber must be a number.'); } if (!Number.isInteger(byteLength) || byteLength < 0) { throw new TypeError('byteLength must be a non-negative integer.'); } if (sideData !== undefined && (typeof sideData !== 'object' || !sideData)) { throw new TypeError('sideData, when provided, must be an object.'); } if (sideData?.alpha !== undefined && !(sideData.alpha instanceof Uint8Array)) { throw new TypeError('sideData.alpha, when provided, must be a Uint8Array.'); } if ( sideData?.alphaByteLength !== undefined && (!Number.isInteger(sideData.alphaByteLength) || sideData.alphaByteLength < 0) ) { throw new TypeError('sideData.alphaByteLength, when provided, must be a non-negative integer.'); } this.byteLength = byteLength; this.sideData = sideData ?? {}; if (this.sideData.alpha && this.sideData.alphaByteLength === undefined) { this.sideData.alphaByteLength = this.sideData.alpha.byteLength; } } /** * If this packet is a metadata-only packet. Metadata-only packets don't contain their packet data. They are the * result of retrieving packets with {@link PacketRetrievalOptions.metadataOnly} set to `true`. */ get isMetadataOnly() { return this.data === PLACEHOLDER_DATA; } /** The timestamp of this packet in microseconds. */ get microsecondTimestamp() { return Math.trunc(SECOND_TO_MICROSECOND_FACTOR * this.timestamp); } /** The duration of this packet in microseconds. */ get microsecondDuration() { return Math.trunc(SECOND_TO_MICROSECOND_FACTOR * this.duration); } /** Converts this packet to an * [`EncodedVideoChunk`](https://developer.mozilla.org/en-US/docs/Web/API/EncodedVideoChunk) for use with the * WebCodecs API. */ toEncodedVideoChunk() { if (this.isMetadataOnly) { throw new TypeError('Metadata-only packets cannot be converted to a video chunk.'); } if (typeof EncodedVideoChunk === 'undefined') { throw new Error('Your browser does not support EncodedVideoChunk.'); } return new EncodedVideoChunk({ data: this.data, type: this.type, timestamp: this.microsecondTimestamp, duration: this.microsecondDuration, }); } /** * Converts this packet to an * [`EncodedVideoChunk`](https://developer.mozilla.org/en-US/docs/Web/API/EncodedVideoChunk) for use with the * WebCodecs API, using the alpha side data instead of the color data. Throws if no alpha side data is defined. */ alphaToEncodedVideoChunk(type = this.type) { if (!this.sideData.alpha) { throw new TypeError('This packet does not contain alpha side data.'); } if (this.isMetadataOnly) { throw new TypeError('Metadata-only packets cannot be converted to a video chunk.'); } if (typeof EncodedVideoChunk === 'undefined') { throw new Error('Your browser does not support EncodedVideoChunk.'); } return new EncodedVideoChunk({ data: this.sideData.alpha, type, timestamp: this.microsecondTimestamp, duration: this.microsecondDuration, }); } /** Converts this packet to an * [`EncodedAudioChunk`](https://developer.mozilla.org/en-US/docs/Web/API/EncodedAudioChunk) for use with the * WebCodecs API. */ toEncodedAudioChunk() { if (this.isMetadataOnly) { throw new TypeError('Metadata-only packets cannot be converted to an audio chunk.'); } if (typeof EncodedAudioChunk === 'undefined') { throw new Error('Your browser does not support EncodedAudioChunk.'); } return new EncodedAudioChunk({ data: this.data, type: this.type, timestamp: this.microsecondTimestamp, duration: this.microsecondDuration, }); } /** * Creates an {@link EncodedPacket} from an * [`EncodedVideoChunk`](https://developer.mozilla.org/en-US/docs/Web/API/EncodedVideoChunk) or * [`EncodedAudioChunk`](https://developer.mozilla.org/en-US/docs/Web/API/EncodedAudioChunk). This method is useful * for converting chunks from the WebCodecs API to `EncodedPacket` instances. */ static fromEncodedChunk( chunk: EncodedVideoChunk | EncodedAudioChunk, sideData?: EncodedPacketSideData, ): EncodedPacket { if (!(chunk instanceof EncodedVideoChunk || chunk instanceof EncodedAudioChunk)) { throw new TypeError('chunk must be an EncodedVideoChunk or EncodedAudioChunk.'); } const data = new Uint8Array(chunk.byteLength); chunk.copyTo(data); return new EncodedPacket( data, chunk.type as PacketType, chunk.timestamp / 1e6, (chunk.duration ?? 0) / 1e6, undefined, undefined, sideData, ); } /** Clones this packet while optionally modifying the new packet's data. */ clone(options?: { /** The data of the cloned packet. */ data?: Uint8Array; /** The type of the cloned packet. */ type?: PacketType; /** The timestamp of the cloned packet in seconds. */ timestamp?: number; /** The duration of the cloned packet in seconds. */ duration?: number; /** The sequence number of the cloned packet. */ sequenceNumber?: number; /** The side data of the cloned packet. */ sideData?: EncodedPacketSideData; }): EncodedPacket { if (options !== undefined && (typeof options !== 'object' || options === null)) { throw new TypeError('options, when provided, must be an object.'); } if (options?.data !== undefined && !(options.data instanceof Uint8Array)) { throw new TypeError('options.data, when provided, must be a Uint8Array.'); } if (options?.type !== undefined && options.type !== 'key' && options.type !== 'delta') { throw new TypeError('options.type, when provided, must be either "key" or "delta".'); } if (options?.timestamp !== undefined && !Number.isFinite(options.timestamp)) { throw new TypeError('options.timestamp, when provided, must be a number.'); } if (options?.duration !== undefined && !Number.isFinite(options.duration)) { throw new TypeError('options.duration, when provided, must be a number.'); } if (options?.sequenceNumber !== undefined && !Number.isFinite(options.sequenceNumber)) { throw new TypeError('options.sequenceNumber, when provided, must be a number.'); } if (options?.sideData !== undefined && (typeof options.sideData !== 'object' || options.sideData === null)) { throw new TypeError('options.sideData, when provided, must be an object.'); } return new EncodedPacket( options?.data ?? this.data, options?.type ?? this.type, options?.timestamp ?? this.timestamp, options?.duration ?? this.duration, options?.sequenceNumber ?? this.sequenceNumber, this.byteLength, options?.sideData ?? this.sideData, ); } } ===== src/target.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import type { FileHandle } from 'node:fs/promises'; import * as nodeAlias from './node'; import { assert, EventEmitter, FilePath, MaybePromise } from './misc'; const node = typeof nodeAlias !== 'undefined' ? nodeAlias // Aliasing it prevents some bundler warnings : undefined!; /** * The events emitted by a {@link Target}. * @group Output targets * @public */ export type TargetEvents = { /** Emitted each time data is written to the target. */ write: { /** The start of the written range, inclusive. */ start: number; /** The end of the written range, exclusive. */ end: number; }; /** Emitted when the target is finalized. */ finalized: void; }; /** * Base class for targets, specifying where output files are written. * @group Output targets * @public */ export abstract class Target extends EventEmitter { /** @internal */ _writerAcquired = false; /** @internal */ _monotonicity: boolean | null = null; // null = unknown /** @internal */ abstract _start(): void; /** @internal */ abstract _write(data: Uint8Array, pos: number): void; /** @internal */ abstract _flush(): Promise; /** @internal */ abstract _finalize(): Promise; /** @internal */ abstract _close(): Promise; /** * Called each time data is written to the target. Will be called with the byte range into which data was written. * * Use this callback to track the size of the output file as it grows. But be warned, this function is chatty and * gets called *extremely* often. * * @deprecated Use `target.on('write', ({ start, end }) => ...)` instead. */ onwrite: ((start: number, end: number) => unknown) | null = null; /** @internal */ _setMonotonicity(monotonicity: boolean) { if (this._monotonicity !== false) { this._monotonicity = monotonicity; } else { // Once false, it's locked } } /** @internal */ _dispatchWrite(start: number, end: number) { // eslint-disable-next-line @typescript-eslint/no-deprecated this.onwrite?.(start, end); this._emit('write', { start, end }); } /** * Returns a new {@link RangedTarget} that writes data to this target using the given offset. * * Useful for writing a file into a section of a larger file. */ slice(offset: number) { if (!Number.isInteger(offset) || offset < 0) { throw new TypeError('offset must be a non-negative integer.'); } return new RangedTarget(this, offset); } } const ARRAY_BUFFER_INITIAL_SIZE = 2 ** 16; const ARRAY_BUFFER_MAX_SIZE = 2 ** 32; /** * Options for {@link BufferTarget}. * @group Output targets * @public */ export type BufferTargetOptions = { /** * Called once the target has been finalized, with the complete output buffer. If you return a promise, it will be * used to apply backpressure internally. * * One use for this callback is for uploading to a server where the full buffer must be known before * sending (e.g. S3 PutObject) and stream-uploading is not an option. */ onFinalize?: (buffer: ArrayBuffer) => MaybePromise; }; /** * A target that writes data directly into an ArrayBuffer in memory. Great for performance, but not suitable for very * large files. The buffer will be available once the output has been finalized. * @group Output targets * @public */ export class BufferTarget extends Target { /** Stores the final output buffer. Until the output is finalized, this will be `null`. */ buffer: ArrayBuffer | null = null; /** @internal */ _buffer: ArrayBuffer; /** @internal */ _bytes: Uint8Array; /** @internal */ _maxPos = 0; /** @internal */ _supportsResize: boolean; /** @internal */ _options: BufferTargetOptions; /** Creates a new {@link BufferTarget}. The buffer holding the data will be created and managed internally. */ constructor(options: BufferTargetOptions = {}) { super(); if (!options || typeof options !== 'object') { throw new TypeError('BufferTarget options, when provided, must be an object.'); } if (options.onFinalize !== undefined && typeof options.onFinalize !== 'function') { throw new TypeError('options.onFinalize, when provided, must be a function.'); } this._options = options; this._supportsResize = 'resize' in new ArrayBuffer(0); if (this._supportsResize) { try { // @ts-expect-error Don't want to bump "lib" in tsconfig this._buffer = new ArrayBuffer(ARRAY_BUFFER_INITIAL_SIZE, { maxByteLength: ARRAY_BUFFER_MAX_SIZE }); } catch { this._buffer = new ArrayBuffer(ARRAY_BUFFER_INITIAL_SIZE); this._supportsResize = false; } } else { this._buffer = new ArrayBuffer(ARRAY_BUFFER_INITIAL_SIZE); } this._bytes = new Uint8Array(this._buffer); } /** @internal */ _ensureSize(size: number) { let newLength = this._buffer.byteLength; while (newLength < size) newLength *= 2; if (newLength === this._buffer.byteLength) return; if (newLength > ARRAY_BUFFER_MAX_SIZE) { throw new Error( `ArrayBuffer exceeded maximum size of ${ARRAY_BUFFER_MAX_SIZE} bytes. Please consider using another` + ` target.`, ); } if (this._supportsResize) { // Use resize if it exists // @ts-expect-error Don't want to bump "lib" in tsconfig // eslint-disable-next-line @typescript-eslint/no-unsafe-call this._buffer.resize(newLength); // The Uint8Array scales automatically } else { const newBuffer = new ArrayBuffer(newLength); const newBytes = new Uint8Array(newBuffer); newBytes.set(this._bytes, 0); this._buffer = newBuffer; this._bytes = newBytes; } } /** @internal */ _start() {} /** @internal */ _write(data: Uint8Array, pos: number) { this._ensureSize(pos + data.byteLength); this._bytes.set(data, pos); this._maxPos = Math.max(this._maxPos, pos + data.byteLength); this._dispatchWrite(pos, pos + data.byteLength); } /** @internal */ async _flush() {} /** @internal */ async _finalize() { this.buffer = this._buffer.slice(0, this._maxPos); if (this._options.onFinalize) { await this._options.onFinalize(this.buffer); } this._emit('finalized'); } /** @internal */ async _close() {} /** @internal */ _getSlice(start: number, end: number) { return this._bytes.slice(start, end); } } /** * A data chunk for {@link StreamTarget}. * @group Output targets * @public */ export type StreamTargetChunk = { /** The operation type. */ type: 'write'; // This ensures automatic compatibility with FileSystemWritableFileStream /** The data to write. */ data: Uint8Array; /** The byte offset in the output file at which to write the data. */ position: number; }; /** * Options for {@link StreamTarget}. * @group Output targets * @public */ export type StreamTargetOptions = { /** * When setting this to true, data created by the output will first be accumulated and only written out * once it has reached sufficient size, using a default chunk size of 16 MiB. This is useful for reducing the total * amount of writes, at the cost of latency. */ chunked?: boolean; /** When using `chunked: true`, this specifies the maximum size of each chunk. Defaults to 16 MiB. */ chunkSize?: number; }; const DEFAULT_CHUNK_SIZE = 2 ** 24; const MAX_CHUNKS_AT_ONCE = 2; type Chunk = { start: number; written: ChunkSection[]; data: Uint8Array; shouldFlush: boolean; }; type ChunkSection = { start: number; end: number; }; /** * This target writes data to a [`WritableStream`](https://developer.mozilla.org/en-US/docs/Web/API/WritableStream), * making it a general-purpose target for writing data anywhere. It is also compatible with * [`FileSystemWritableFileStream`](https://developer.mozilla.org/en-US/docs/Web/API/FileSystemWritableFileStream) for * use with the [File System Access API](https://developer.mozilla.org/en-US/docs/Web/API/File_System_API). The * `WritableStream` can also apply backpressure, which will propagate to the output and throttle the encoders. * @group Output targets * @public */ export class StreamTarget extends Target { /** @internal */ _writable: WritableStream; /** @internal */ _options: StreamTargetOptions; /** @internal */ _sections: { data: Uint8Array; start: number; }[] = []; /** @internal */ _lastWriteEnd = 0; /** @internal */ _lastFlushEnd = 0; /** @internal */ _streamWriter: WritableStreamDefaultWriter | null = null; /** @internal */ _writeError: unknown = null; // These variables regard chunked mode: /** @internal */ _chunked: boolean; /** @internal */ _chunkSize: number; /** * The data is divided up into fixed-size chunks, whose contents are first filled in RAM and then flushed out. * A chunk is flushed if all of its contents have been written. */ /** @internal */ _chunks: Chunk[] = []; /** Creates a new {@link StreamTarget} which writes to the specified `writable`. */ constructor( writable: WritableStream, options: StreamTargetOptions = {}, ) { super(); if (!(writable instanceof WritableStream)) { throw new TypeError('StreamTarget requires a WritableStream instance.'); } if (options != null && typeof options !== 'object') { throw new TypeError('StreamTarget options, when provided, must be an object.'); } if (options.chunked !== undefined && typeof options.chunked !== 'boolean') { throw new TypeError('options.chunked, when provided, must be a boolean.'); } if (options.chunkSize !== undefined && (!Number.isInteger(options.chunkSize) || options.chunkSize < 1024)) { throw new TypeError('options.chunkSize, when provided, must be an integer and not smaller than 1024.'); } this._writable = writable; this._options = options; this._chunked = options.chunked ?? false; this._chunkSize = options.chunkSize ?? DEFAULT_CHUNK_SIZE; } /** @internal */ _start() { this._streamWriter = this._writable.getWriter(); } /** @internal */ _write(data: Uint8Array, pos: number) { if (pos > this._lastWriteEnd) { const paddingBytesNeeded = pos - this._lastWriteEnd; this._write(new Uint8Array(paddingBytesNeeded), this._lastWriteEnd); } this._sections.push({ data: data.slice(), start: pos, }); this._lastWriteEnd = Math.max(this._lastWriteEnd, pos + data.byteLength); this._dispatchWrite(pos, pos + data.byteLength); } /** @internal */ async _flush() { if (this._writeError !== null) { // eslint-disable-next-line @typescript-eslint/only-throw-error throw this._writeError; } assert(this._streamWriter); if (this._sections.length === 0) { return; } const chunks: { start: number; size: number; data?: Uint8Array; }[] = []; const sorted = [...this._sections].sort((a, b) => a.start - b.start); chunks.push({ start: sorted[0]!.start, size: sorted[0]!.data.byteLength, }); // Figure out how many contiguous chunks we have for (let i = 1; i < sorted.length; i++) { const lastChunk = chunks[chunks.length - 1]!; const section = sorted[i]!; if (section.start <= lastChunk.start + lastChunk.size) { lastChunk.size = Math.max(lastChunk.size, section.start + section.data.byteLength - lastChunk.start); } else { chunks.push({ start: section.start, size: section.data.byteLength, }); } } for (const chunk of chunks) { chunk.data = new Uint8Array(chunk.size); // Make sure to write the data in the correct order for correct overwriting for (const section of this._sections) { // Check if the section is in the chunk if (chunk.start <= section.start && section.start < chunk.start + chunk.size) { chunk.data.set(section.data, section.start - chunk.start); } } if (this._streamWriter.desiredSize !== null && this._streamWriter.desiredSize <= 0) { await this._streamWriter.ready; // Allow the writer to apply backpressure } if (this._chunked) { // Let's first gather the data into bigger chunks before writing it this._writeDataIntoChunks(chunk.data, chunk.start); this._tryToFlushChunks(); } else { if (this._monotonicity === true && chunk.start !== this._lastFlushEnd) { throw new Error('Internal error: Monotonicity violation.'); } void this._streamWriter.write({ type: 'write', data: chunk.data, position: chunk.start, }).catch((error) => { this._writeError ??= error; }); this._lastFlushEnd = chunk.start + chunk.data.byteLength; } } this._sections.length = 0; } /** @internal */ _writeDataIntoChunks(data: Uint8Array, position: number) { // First, find the chunk to write the data into, or create one if none exists let chunkIndex = this._chunks.findIndex(x => x.start <= position && position < x.start + this._chunkSize); if (chunkIndex === -1) chunkIndex = this._createChunk(position); const chunk = this._chunks[chunkIndex]!; // Figure out how much to write to the chunk, and then write to the chunk const relativePosition = position - chunk.start; const toWrite = data.subarray(0, Math.min(this._chunkSize - relativePosition, data.byteLength)); chunk.data.set(toWrite, relativePosition); // Create a section describing the region of data that was just written to const section: ChunkSection = { start: relativePosition, end: relativePosition + toWrite.byteLength, }; this._insertSectionIntoChunk(chunk, section); // Queue chunk for flushing to target if it has been fully written to if (chunk.written[0]!.start === 0 && chunk.written[0]!.end === this._chunkSize) { chunk.shouldFlush = true; } // Make sure we don't hold too many chunks in memory at once to keep memory usage down if (this._chunks.length > MAX_CHUNKS_AT_ONCE) { // Flush all but the last chunk for (let i = 0; i < this._chunks.length - 1; i++) { this._chunks[i]!.shouldFlush = true; } this._tryToFlushChunks(); } // If the data didn't fit in one chunk, recurse with the remaining data if (toWrite.byteLength < data.byteLength) { this._writeDataIntoChunks(data.subarray(toWrite.byteLength), position + toWrite.byteLength); } } /** @internal */ _insertSectionIntoChunk(chunk: Chunk, section: ChunkSection) { let low = 0; let high = chunk.written.length - 1; let index = -1; // Do a binary search to find the last section with a start not larger than `section`'s start while (low <= high) { const mid = Math.floor(low + (high - low + 1) / 2); if (chunk.written[mid]!.start <= section.start) { low = mid + 1; index = mid; } else { high = mid - 1; } } // Insert the new section chunk.written.splice(index + 1, 0, section); if (index === -1 || chunk.written[index]!.end < section.start) index++; // Merge overlapping sections while (index < chunk.written.length - 1 && chunk.written[index]!.end >= chunk.written[index + 1]!.start) { chunk.written[index]!.end = Math.max(chunk.written[index]!.end, chunk.written[index + 1]!.end); chunk.written.splice(index + 1, 1); } } /** @internal */ _createChunk(includesPosition: number) { const start = Math.floor(includesPosition / this._chunkSize) * this._chunkSize; const chunk: Chunk = { start, data: new Uint8Array(this._chunkSize), written: [], shouldFlush: false, }; this._chunks.push(chunk); this._chunks.sort((a, b) => a.start - b.start); return this._chunks.indexOf(chunk); } /** @internal */ _tryToFlushChunks(force = false) { assert(this._streamWriter); for (let i = 0; i < this._chunks.length; i++) { const chunk = this._chunks[i]!; if (!chunk.shouldFlush && !force) continue; for (const section of chunk.written) { const position = chunk.start + section.start; if (this._monotonicity === true && position !== this._lastFlushEnd) { throw new Error('Internal error: Monotonicity violation.'); } void this._streamWriter.write({ type: 'write', data: chunk.data.subarray(section.start, section.end), position, }).catch((error) => { this._writeError ??= error; }); this._lastFlushEnd = chunk.start + section.end; } this._chunks.splice(i--, 1); } } /** @internal */ async _finalize() { if (this._chunked) { this._tryToFlushChunks(true); } if (this._writeError !== null) { // eslint-disable-next-line @typescript-eslint/only-throw-error throw this._writeError; } assert(this._streamWriter); await this._streamWriter.ready; await this._streamWriter.close(); this._emit('finalized'); } /** @internal */ async _close() { return this._streamWriter?.close(); } } /** * This target writes to a `WritableStream`, meaning all writes are necessarily append-only and involve no * seeking. Great for streaming data to a source that can only accept sequential data, like an HTTP server processing * an incoming upload. * * Note that using this target *requires* that the underlying format write data sequentially. Not all formats do this, * and this target will throw for the formats that don't. Check the guide for more. * * @group Output targets * @public */ export class AppendOnlyStreamTarget extends Target { /** @internal */ _writable: WritableStream; /** @internal */ _streamTarget: StreamTarget; /** @internal */ _writer: WritableStreamDefaultWriter | null = null; /** @internal */ _nextWritePos = 0; constructor(writable: WritableStream) { super(); this._writable = writable; this._streamTarget = new StreamTarget(new WritableStream({ start: () => { this._writer = this._writable.getWriter(); }, write: (chunk) => { if (this._monotonicity !== true) { throw new Error( 'AppendOnlyStreamTarget requires that data be written monotonically (always appended to the' + ' end). You must use a format that guarantees this behavior.', ); } assert(chunk.position === this._nextWritePos); this._nextWritePos += chunk.data.byteLength; assert(this._writer); return this._writer.write(chunk.data); }, close: () => { return this._writer?.close(); }, })); } /** @internal */ _start(): void { this._streamTarget._start(); } /** @internal */ _write(data: Uint8Array, pos: number): void { this._streamTarget._write(data, pos); } /** @internal */ _flush(): Promise { return this._streamTarget._flush(); } /** @internal */ _finalize(): Promise { return this._streamTarget._finalize(); } /** @internal */ _close(): Promise { return this._streamTarget._close(); } /** @internal */ override _setMonotonicity(monotonicity: boolean): void { super._setMonotonicity(monotonicity); this._streamTarget._setMonotonicity(monotonicity); } } /** * Options for {@link FilePathTarget}. * @group Output targets * @public */ export type FilePathTargetOptions = StreamTargetOptions; /** * A target that writes to a file at the specified path. Intended for server-side usage in Node, Bun, or Deno. * * Writing is chunked by default. The internally held file handle will be closed when `.finalize()` or `.cancel()` are * called on the corresponding {@link Output}. * @group Output targets * @public */ export class FilePathTarget extends Target { /** @internal */ _streamTarget: StreamTarget; /** @internal */ _fileHandle: FileHandle | null = null; /** Creates a new {@link FilePathTarget} that writes to the file at the specified file path. */ constructor(filePath: string, options: FilePathTargetOptions = {}) { if (typeof filePath !== 'string') { throw new TypeError('filePath must be a string.'); } if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (!node.fs) { throw new Error( 'FilePathTarget is only available in server-side environments (Node.js, Bun, Deno).', ); } super(); // Let's back this target with a StreamTarget, makes the implementation very simple const writable = new WritableStream({ start: async () => { this._fileHandle = await node.fs.open(filePath, 'w'); }, write: async (chunk) => { assert(this._fileHandle); await this._fileHandle.write(chunk.data, 0, chunk.data.byteLength, chunk.position); }, close: async () => { if (this._fileHandle) { await this._fileHandle.close(); this._fileHandle = null; } }, }); this._streamTarget = new StreamTarget(writable, { chunked: true, ...options, }); } /** @internal */ _start() { this._streamTarget._start(); } /** @internal */ _write(data: Uint8Array, pos: number) { this._streamTarget._write(data, pos); this._dispatchWrite(pos, pos + data.byteLength); } /** @internal */ async _flush() { return this._streamTarget._flush(); } /** @internal */ async _finalize() { await this._streamTarget._finalize(); this._emit('finalized'); } /** @internal */ async _close() { return this._streamTarget._close(); } /** @internal */ override _setMonotonicity(monotonicity: boolean): void { super._setMonotonicity(monotonicity); this._streamTarget._setMonotonicity(monotonicity); } } /** * This target just discards all incoming data. It is useful for when you need an {@link Output} but extract data from * it differently, for example through format-specific callbacks (`onMoof`, `onMdat`, ...) or encoder events. * @group Output targets * @public */ export class NullTarget extends Target { /** @internal */ _start() {} /** @internal */ _write(data: Uint8Array, pos: number) { this._dispatchWrite(pos, pos + data.byteLength); } /** @internal */ async _flush() {} /** @internal */ async _finalize() { this._emit('finalized'); } /** @internal */ async _close() {} } /** * A target that writes to a subrange (defined by an offset) of another, underlying target. Useful for writing a file * into a section of a larger file. * @group Output targets * @public */ export class RangedTarget extends Target { /** @internal */ _baseTarget: Target; /** @internal */ _offset: number; /** @internal */ constructor(baseTarget: Target, offset: number) { super(); this._baseTarget = baseTarget; this._offset = offset; } /** @internal */ _start() {} /** @internal */ _write(data: Uint8Array, pos: number): void { this._baseTarget._write(data, this._offset + pos); this._dispatchWrite(pos, pos + data.byteLength); } /** @internal */ _flush() { return this._baseTarget._flush(); } /** @internal */ async _finalize() { this._emit('finalized'); } /** @internal */ async _close() {} /** @internal */ override _setMonotonicity(monotonicity: boolean): void { super._setMonotonicity(monotonicity); this._baseTarget._setMonotonicity(monotonicity); } } /** * A special target for writing multi-file media where each file is uniquely identified by a path. * @group Output targets * @public */ export class PathedTarget { /** Creates a new {@link PathedTarget} from a root path and a callback. */ constructor( /** The path that points to the root file; the entry file of the media. */ public readonly rootPath: FilePath, /** The callback that is called for each file that needs to be written; must return a {@link Target}. */ public readonly getTarget: (request: TargetRequest) => MaybePromise, ) { if (typeof rootPath !== 'string') { throw new TypeError('rootPath must be a string.'); } if (typeof getTarget !== 'function') { throw new TypeError('getTarget must be a function.'); } } } /** * A request for a {@link Target} at the given path. * @group Output targets * @public */ export type TargetRequest = { /** The requested file path. */ path: FilePath; /** Whether the to-be-written file will be the root file. */ isRoot: boolean; /** The MIME type of the to-be-written file. */ mimeType: string; }; ===== src/writer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { assert } from './misc'; import { Target } from './target'; export class Writer { target: Target; finalized = false; started = false; private pos = 0; constructor(target: Target, isMonotonic: boolean) { if (target._writerAcquired) { throw new Error('Can\'t have multiple Writers for the same Target.'); } this.target = target; target._setMonotonicity(isMonotonic); target._writerAcquired = true; } start() { assert(!this.started); this.target._start(); this.started = true; } /** Writes the given data to the target, at the current position. */ write(data: Uint8Array) { assert(this.started && !this.finalized); this.maybeTrackWrites(data); this.target._write(data, this.pos); this.pos += data.byteLength; } /** Sets the current position for future writes to a new one. */ seek(newPos: number) { this.pos = newPos; } /** Returns the current position. */ getPos() { return this.pos; } /** Signals to the writer that it may be time to flush. */ async flush() { assert(this.started && !this.finalized); return this.target._flush(); } /** Called after muxing has finished. */ async finalize() { assert(this.started && !this.finalized); await this.target._finalize(); this.finalized = true; } private trackedWrites: Uint8Array | null = null; private trackedStart = -1; private trackedEnd = -1; private maybeTrackWrites(data: Uint8Array) { if (!this.trackedWrites) { return; } // Handle negative relative write positions let pos = this.getPos(); if (pos < this.trackedStart) { if (pos + data.byteLength <= this.trackedStart) { return; } data = data.subarray(this.trackedStart - pos); pos = 0; } const neededSize = pos + data.byteLength - this.trackedStart; let newLength = this.trackedWrites.byteLength; while (newLength < neededSize) { newLength *= 2; } // Check if we need to resize the buffer if (newLength !== this.trackedWrites.byteLength) { const copy = new Uint8Array(newLength); copy.set(this.trackedWrites, 0); this.trackedWrites = copy; } this.trackedWrites.set(data, pos - this.trackedStart); this.trackedEnd = Math.max(this.trackedEnd, pos + data.byteLength); } startTrackingWrites() { this.trackedWrites = new Uint8Array(2 ** 10); this.trackedStart = this.getPos(); this.trackedEnd = this.trackedStart; } stopTrackingWrites() { if (!this.trackedWrites) { throw new Error('Internal error: Can\'t get tracked writes since nothing was tracked.'); } const slice = this.trackedWrites.subarray(0, this.trackedEnd - this.trackedStart); const result = { data: slice, start: this.trackedStart, end: this.trackedEnd, }; this.trackedWrites = null; return result; } } ===== src/sample.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { assert, clamp, COLOR_PRIMARIES_MAP, isAllowSharedBufferSource, MATRIX_COEFFICIENTS_MAP, Rotation, SECOND_TO_MICROSECOND_FACTOR, toDataView, toUint8Array, SetRequired, TRANSFER_CHARACTERISTICS_MAP, isFirefox, polyfillSymbolDispose, assertNever, isWebKit, Rational, simplifyRational, Rectangle, validateRectangle, normalizeRotation, roundToMultiple, arrayArgmin, MaybePromise, DeepReadonly, } from './misc'; import { Logging } from './logging'; polyfillSymbolDispose(); type FinalizationRegistryValue = { type: 'video'; data: VideoFrame | OffscreenCanvas | Uint8Array | VideoSampleResource; } | { type: 'audio'; data: AudioData | Uint8Array | AudioSampleResource; }; // Let's manually handle logging the garbage collection errors that are typically logged by the browser. This way, they // also kick for audio samples (which is normally not the case), making sure any incorrect code is quickly caught. let lastVideoGcErrorLog = -Infinity; let lastAudioGcErrorLog = -Infinity; let finalizationRegistry: FinalizationRegistry | null = null; if (typeof FinalizationRegistry !== 'undefined') { finalizationRegistry = new FinalizationRegistry((value) => { const now = performance.now(); if (value.type === 'video') { if (now - lastVideoGcErrorLog >= 1000) { // This error is annoying but oh so important Logging._error( `A VideoSample was garbage collected without first being closed. For proper resource management,` + ` make sure to call close() on all your VideoSamples as soon as you're done using them.`, ); lastVideoGcErrorLog = now; } if (typeof VideoFrame !== 'undefined' && value.data instanceof VideoFrame) { value.data.close(); // Prevent the browser error since we're logging our own } } else { if (now - lastAudioGcErrorLog >= 1000) { Logging._error( `An AudioSample was garbage collected without first being closed. For proper resource management,` + ` make sure to call close() on all your AudioSamples as soon as you're done using them.`, ); lastAudioGcErrorLog = now; } if (typeof AudioData !== 'undefined' && value.data instanceof AudioData) { value.data.close(); } } }); } /** * Abstract base class for custom video sample resources. Implement this class to provide custom backing * for VideoSample instances. * @group Samples * @public */ export abstract class VideoSampleResource { /** @internal */ _referenceCount: number = 0; /** @internal */ _lastAllocationBuffer: ArrayBuffer | null = null; /** * Returns the internal pixel format in which the frame is stored. * [See pixel formats](https://developer.mozilla.org/en-US/docs/Web/API/VideoFrame/format) */ abstract getFormat(): VideoSamplePixelFormat | null; /** Returns the width of the frame in pixels. */ abstract getCodedWidth(): number; /** Returns the height of the frame in pixels. */ abstract getCodedHeight(): number; /** Returns the width of the frame in square pixels, respecting pixel aspect ratio. */ abstract getSquarePixelWidth(): number; /** Returns the height of the frame in square pixels, respecting pixel aspect ratio. */ abstract getSquarePixelHeight(): number; /** Returns the color space of the frame. */ abstract getColorSpace(): VideoSampleColorSpace; /** * Closes this resource, releasing held resources. Called automatically when the last {@link VideoSample} using this * resource is closed. */ abstract close(): void; /** * Returns the data planes that hold the video data for this sample. The returned planes and data must be in the * format returned by `getFormat()`. */ abstract getDataPlanes(): MaybePromise; /** * Returns a new RGB {@link VideoSample} that contains the same content as this sample. The provided `init` object * must be used to set the metadata of this new video sample. When converting from a non-RGB format to RGB, the * conversion must respect `colorSpace`. */ abstract toRgbSample( init: SetRequired, colorSpace: PredefinedColorSpace, ): MaybePromise; } /** * Describes a single data plane of a video frame. * @group Samples * @public */ export type VideoDataPlane = { /** The data of the plane. */ data: Uint8Array; /** The stride of the plane, in bytes. This is the distance in bytes between the start of each row of pixels. */ stride: number; }; /** * The list of {@link VideoSample} pixel formats. * @group Samples * @public */ export const VIDEO_SAMPLE_PIXEL_FORMATS = [ // 4:2:0 Y, U, V 'I420', 'I420P10', 'I420P12', // 4:2:0 Y, U, V, A 'I420A', 'I420AP10', 'I420AP12', // 4:2:2 Y, U, V 'I422', 'I422P10', 'I422P12', // 4:2:2 Y, U, V, A 'I422A', 'I422AP10', 'I422AP12', // 4:4:4 Y, U, V 'I444', 'I444P10', 'I444P12', // 4:4:4 Y, U, V, A 'I444A', 'I444AP10', 'I444AP12', // 4:2:0 Y, UV 'NV12', // 4:4:4 RGBA 'RGBA', // 4:4:4 RGBX (opaque) 'RGBX', // 4:4:4 BGRA 'BGRA', // 4:4:4 BGRX (opaque) 'BGRX', ] as const; const VIDEO_SAMPLE_PIXEL_FORMATS_SET = new Set(VIDEO_SAMPLE_PIXEL_FORMATS); /** * The internal pixel format with which a {@link VideoSample} is stored. * [See pixel formats](https://www.w3.org/TR/webcodecs/#pixel-format) for more. * @group Samples * @public */ export type VideoSamplePixelFormat = typeof VIDEO_SAMPLE_PIXEL_FORMATS[number]; /** * Metadata used for VideoSample initialization. * @group Samples * @public */ export type VideoSampleInit = { /** * The internal pixel format in which the frame is stored. * [See pixel formats](https://www.w3.org/TR/webcodecs/#pixel-format) */ format?: VideoSamplePixelFormat; /** The width of the frame in pixels. */ codedWidth?: number; /** The height of the frame in pixels. */ codedHeight?: number; /** The rotation of the frame in degrees, clockwise. */ rotation?: Rotation; /** The presentation timestamp of the frame in seconds. */ timestamp?: number; /** The duration of the frame in seconds. */ duration?: number; /** The color space of the frame. */ colorSpace?: VideoColorSpaceInit; /** The byte layout of the planes of the frame. */ layout?: PlaneLayout[]; /** Visible region in the coded frame. When omitted, the rect defaults to `(0, 0, codedWidth, codedHeight)`. */ visibleRect?: Rectangle | undefined; /** Width of the frame in pixels after applying aspect ratio adjustments and rotation. */ displayWidth?: number | undefined; /** Height of the frame in pixels after applying aspect ratio adjustments and rotation. */ displayHeight?: number | undefined; /** The encode options to use when this sample is passed to an encoder. */ encodeOptions?: DeepReadonly; /** @internal */ _doNotCopy?: boolean; }; /** * Represents a raw, unencoded video sample (frame). Mainly used as an expressive wrapper around WebCodecs API's * [`VideoFrame`](https://developer.mozilla.org/en-US/docs/Web/API/VideoFrame), but can also be used standalone. * @group Samples * @public */ export class VideoSample implements Disposable { /** @internal */ _data!: VideoFrame | OffscreenCanvas | Uint8Array | VideoSampleResource | null; /** * Used for the ArrayBuffer-backed case. * @internal */ _layout!: PlaneLayout[] | null; /** @internal */ _closed: boolean = false; /** * The internal pixel format in which the frame is stored. Will be `null` if it's using an arbitrary internal * format not representable by `VideoSamplePixelFormat`. * [See pixel formats](https://www.w3.org/TR/webcodecs/#pixel-format) */ readonly format!: VideoSamplePixelFormat | null; /** The visible region of the frame in the coded pixel grid. */ readonly visibleRect!: Rectangle; /** The width of the frame in square pixels (respecting pixel aspect ratio), before rotation is applied. */ readonly squarePixelWidth!: number; /** The height of the frame in square pixels (respecting pixel aspect ratio), before rotation is applied. */ readonly squarePixelHeight!: number; /** The rotation of the frame in degrees, clockwise. */ readonly rotation!: Rotation; /** * The pixel aspect ratio of the frame, as a rational number in its reduced form. Most videos use * square pixels (1:1). */ readonly pixelAspectRatio!: Rational; /** * The presentation timestamp of the frame in seconds. May be negative. Frames with negative end timestamps should * not be presented. */ readonly timestamp!: number; /** The duration of the frame in seconds. */ readonly duration!: number; /** The color space of the frame. */ readonly colorSpace!: VideoSampleColorSpace; /** The encode options to use when this sample is passed to an encoder. */ readonly encodeOptions!: DeepReadonly; /** The width of the frame in pixels. */ get codedWidth() { // This is wrong, but the fix is a v2 thing return this.visibleRect.width; } /** The height of the frame in pixels. */ get codedHeight() { // Same here return this.visibleRect.height; } /** The display width of the frame in pixels, after aspect ratio adjustment and rotation. */ get displayWidth() { return this.rotation % 180 === 0 ? this.squarePixelWidth : this.squarePixelHeight; } /** The display height of the frame in pixels, after aspect ratio adjustment and rotation. */ get displayHeight() { return this.rotation % 180 === 0 ? this.squarePixelHeight : this.squarePixelWidth; } /** The presentation timestamp of the frame in microseconds. */ get microsecondTimestamp() { return Math.trunc(SECOND_TO_MICROSECOND_FACTOR * this.timestamp); } /** The duration of the frame in microseconds. */ get microsecondDuration() { return Math.trunc(SECOND_TO_MICROSECOND_FACTOR * this.duration); } /** * Whether this sample uses a pixel format that can hold transparency data. Note that this doesn't necessarily mean * that the sample is transparent. */ get hasAlpha() { return this.format && this.format.includes('A'); } /** * Creates a new {@link VideoSample} from a * [`VideoFrame`](https://developer.mozilla.org/en-US/docs/Web/API/VideoFrame). This is essentially a near zero-cost * wrapper around `VideoFrame`. The sample's metadata is optionally refined using the data specified in `init`. */ constructor(data: VideoFrame, init?: VideoSampleInit); /** * Creates a new {@link VideoSample} from a * [`CanvasImageSource`](https://udn.realityripple.com/docs/Web/API/CanvasImageSource), similar to the * [`VideoFrame`](https://developer.mozilla.org/en-US/docs/Web/API/VideoFrame) constructor. When `VideoFrame` is * available, this is simply a wrapper around its constructor. If not, it will copy the source's image data to an * internal canvas for later use. */ constructor(data: CanvasImageSource, init: SetRequired); /** * Creates a new {@link VideoSample} from raw pixel data specified in `data`. Additional metadata must be provided * in `init`. */ constructor( data: AllowSharedBufferSource, init: SetRequired ); /** * Creates a new {@link VideoSample} backed by a custom {@link VideoSampleResource}. */ constructor(resource: VideoSampleResource, init: SetRequired); constructor( data: VideoFrame | CanvasImageSource | AllowSharedBufferSource | VideoSampleResource, init?: VideoSampleInit, ) { if ( data instanceof ArrayBuffer || (typeof SharedArrayBuffer !== 'undefined' && data instanceof SharedArrayBuffer) || ArrayBuffer.isView(data) ) { if (!init || typeof init !== 'object') { throw new TypeError('init must be an object.'); } if (init.format === undefined || !VIDEO_SAMPLE_PIXEL_FORMATS_SET.has(init.format)) { throw new TypeError('init.format must be one of: ' + VIDEO_SAMPLE_PIXEL_FORMATS.join(', ')); } if (!Number.isInteger(init.codedWidth) || init.codedWidth! <= 0) { throw new TypeError('init.codedWidth must be a positive integer.'); } if (!Number.isInteger(init.codedHeight) || init.codedHeight! <= 0) { throw new TypeError('init.codedHeight must be a positive integer.'); } if (init.rotation !== undefined && ![0, 90, 180, 270].includes(init.rotation)) { throw new TypeError('init.rotation, when provided, must be 0, 90, 180, or 270.'); } if (!Number.isFinite(init.timestamp)) { throw new TypeError('init.timestamp must be a number.'); } if (init.duration !== undefined && (!Number.isFinite(init.duration) || init.duration < 0)) { throw new TypeError('init.duration, when provided, must be a non-negative number.'); } if (init.layout !== undefined) { if (!Array.isArray(init.layout)) { throw new TypeError('init.layout, when provided, must be an array.'); } for (const plane of init.layout) { if (!plane || typeof plane !== 'object' || Array.isArray(plane)) { throw new TypeError('Each entry in init.layout must be an object.'); } if (!Number.isInteger(plane.offset) || plane.offset < 0) { throw new TypeError('plane.offset must be a non-negative integer.'); } if (!Number.isInteger(plane.stride) || plane.stride < 0) { throw new TypeError('plane.stride must be a non-negative integer.'); } } } if (init.visibleRect !== undefined) { validateRectangle(init.visibleRect, 'init.visibleRect'); } if ( init.displayWidth !== undefined && (!Number.isInteger(init.displayWidth) || init.displayWidth <= 0) ) { throw new TypeError('init.displayWidth, when provided, must be a positive integer.'); } if ( init.displayHeight !== undefined && (!Number.isInteger(init.displayHeight) || init.displayHeight <= 0) ) { throw new TypeError('init.displayHeight, when provided, must be a positive integer.'); } if ((init.displayWidth !== undefined) !== (init.displayHeight !== undefined)) { throw new TypeError( 'init.displayWidth and init.displayHeight must be either both provided or both omitted.', ); } this._data = init._doNotCopy ? toUint8Array(data) : toUint8Array(data).slice(); // Copy it this._layout = init.layout ?? createDefaultPlaneLayout(init.format, init.codedWidth!, init.codedHeight!); this.format = init.format; this.rotation = init.rotation ?? 0; this.timestamp = init.timestamp!; this.duration = init.duration ?? 0; let colorSpaceInit = init.colorSpace ?? null; if (colorSpaceInit === null) { if ( this.format === 'RGBA' || this.format === 'RGBX' || this.format === 'BGRA' || this.format === 'BGRX' ) { // sRGB Color Space colorSpaceInit = { primaries: 'bt709', transfer: 'iec61966-2-1', matrix: 'rgb', fullRange: true, }; } else { // REC709 Color Space colorSpaceInit = { primaries: 'bt709', transfer: 'bt709', matrix: 'bt709', fullRange: false, }; } } this.colorSpace = new VideoSampleColorSpace(colorSpaceInit); this.visibleRect = { left: init.visibleRect?.left ?? 0, top: init.visibleRect?.top ?? 0, width: init.visibleRect?.width ?? init.codedWidth!, height: init.visibleRect?.height ?? init.codedHeight!, }; if (init.displayWidth !== undefined) { this.squarePixelWidth = this.rotation % 180 === 0 ? init.displayWidth : init.displayHeight!; this.squarePixelHeight = this.rotation % 180 === 0 ? init.displayHeight! : init.displayWidth; } else { this.squarePixelWidth = this.visibleRect.width; this.squarePixelHeight = this.visibleRect.height; } } else if (typeof VideoFrame !== 'undefined' && data instanceof VideoFrame) { if (init?.rotation !== undefined && ![0, 90, 180, 270].includes(init.rotation)) { throw new TypeError('init.rotation, when provided, must be 0, 90, 180, or 270.'); } if (init?.timestamp !== undefined && !Number.isFinite(init?.timestamp)) { throw new TypeError('init.timestamp, when provided, must be a number.'); } if (init?.duration !== undefined && (!Number.isFinite(init.duration) || init.duration < 0)) { throw new TypeError('init.duration, when provided, must be a non-negative number.'); } if (init?.visibleRect !== undefined) { validateRectangle(init.visibleRect, 'init.visibleRect'); } this._data = data; this._layout = null; this.format = data.format; this.visibleRect = { left: data.visibleRect?.x ?? 0, top: data.visibleRect?.y ?? 0, width: data.visibleRect?.width ?? data.codedWidth, height: data.visibleRect?.height ?? data.codedHeight, }; // The VideoFrame's rotation is ignored here. It's still a new field, and I'm not sure of any application // where the browser makes use of it. If a case gets found, I'll add it. this.rotation = init?.rotation ?? 0; // Assuming no innate VideoFrame rotation here this.squarePixelWidth = data.displayWidth; this.squarePixelHeight = data.displayHeight; this.timestamp = init?.timestamp ?? data.timestamp / 1e6; this.duration = init?.duration ?? (data.duration ?? 0) / 1e6; this.colorSpace = new VideoSampleColorSpace(data.colorSpace); } else if ( (typeof HTMLImageElement !== 'undefined' && data instanceof HTMLImageElement) || (typeof SVGImageElement !== 'undefined' && data instanceof SVGImageElement) || (typeof ImageBitmap !== 'undefined' && data instanceof ImageBitmap) || (typeof HTMLVideoElement !== 'undefined' && data instanceof HTMLVideoElement) || (typeof HTMLCanvasElement !== 'undefined' && data instanceof HTMLCanvasElement) || (typeof OffscreenCanvas !== 'undefined' && data instanceof OffscreenCanvas) ) { if (!init || typeof init !== 'object') { throw new TypeError('init must be an object.'); } if (init.rotation !== undefined && ![0, 90, 180, 270].includes(init.rotation)) { throw new TypeError('init.rotation, when provided, must be 0, 90, 180, or 270.'); } if (!Number.isFinite(init.timestamp)) { throw new TypeError('init.timestamp must be a number.'); } if (init.duration !== undefined && (!Number.isFinite(init.duration) || init.duration < 0)) { throw new TypeError('init.duration, when provided, must be a non-negative number.'); } if (typeof VideoFrame !== 'undefined') { return new VideoSample( new VideoFrame(data, { timestamp: Math.trunc(init.timestamp! * SECOND_TO_MICROSECOND_FACTOR), // Drag 0 to undefined duration: Math.trunc((init.duration ?? 0) * SECOND_TO_MICROSECOND_FACTOR) || undefined, }), init, ); } let width = 0; let height = 0; // Determine the dimensions of the thing if ('naturalWidth' in data) { width = data.naturalWidth; height = data.naturalHeight; } else if ('videoWidth' in data) { width = data.videoWidth; height = data.videoHeight; } else if ('width' in data) { width = Number(data.width); height = Number(data.height); } if (!width || !height) { throw new TypeError('Could not determine dimensions.'); } const canvas = new OffscreenCanvas(width, height); const context = canvas.getContext('2d', { alpha: isFirefox(), // Firefox has VideoFrame glitches with opaque canvases willReadFrequently: true, }); if (!context) { throw new Error( 'OffscreenCanvas must have support for the \'2d\' context in order to create a VideoSample from' + ' this data.', ); } // Draw it to a canvas context.drawImage(data, 0, 0); this._data = canvas; this._layout = null; this.format = 'RGBX'; this.visibleRect = { left: 0, top: 0, width, height }; this.squarePixelWidth = width; this.squarePixelHeight = height; this.rotation = init.rotation ?? 0; this.timestamp = init.timestamp!; this.duration = init.duration ?? 0; this.colorSpace = new VideoSampleColorSpace({ matrix: 'rgb', primaries: 'bt709', transfer: 'iec61966-2-1', fullRange: true, }); } else if (data instanceof VideoSampleResource) { if (!init || typeof init !== 'object') { throw new TypeError('init must be an object.'); } if (init.rotation !== undefined && ![0, 90, 180, 270].includes(init.rotation)) { throw new TypeError('init.rotation, when provided, must be 0, 90, 180, or 270.'); } if (!Number.isFinite(init.timestamp)) { throw new TypeError('init.timestamp must be a number.'); } if (init.duration !== undefined && (!Number.isFinite(init.duration) || init.duration < 0)) { throw new TypeError('init.duration, when provided, must be a non-negative number.'); } this._data = data; data._referenceCount++; this.format = data.getFormat(); if (this.format !== null && !VIDEO_SAMPLE_PIXEL_FORMATS.includes(this.format)) { throw new TypeError('getFormat() must return a VideoSamplePixelFormat or null.'); } this.visibleRect = { left: 0, top: 0, width: data.getCodedWidth(), height: data.getCodedHeight(), }; if (!Number.isInteger(this.visibleRect.width) || this.visibleRect.width <= 0) { throw new TypeError('getCodedWidth() must return a positive integer.'); } if (!Number.isInteger(this.visibleRect.height) || this.visibleRect.height <= 0) { throw new TypeError('getCodedHeight() must return a positive integer.'); } this.squarePixelWidth = data.getSquarePixelWidth(); if (!Number.isInteger(this.squarePixelWidth) || this.squarePixelWidth <= 0) { throw new TypeError('getSquarePixelWidth() must return a positive integer.'); } this.squarePixelHeight = data.getSquarePixelHeight(); if (!Number.isInteger(this.squarePixelHeight) || this.squarePixelHeight <= 0) { throw new TypeError('getSquarePixelHeight() must return a positive integer.'); } this.rotation = init.rotation ?? 0; this.timestamp = init.timestamp!; this.duration = init.duration ?? 0; this.colorSpace = data.getColorSpace(); } else { throw new TypeError( 'Invalid data type: Must be a BufferSource, CanvasImageSource, or VideoSampleResource.', ); } this.encodeOptions = init?.encodeOptions ?? {}; this.pixelAspectRatio = simplifyRational({ num: this.squarePixelWidth * this.codedHeight, den: this.squarePixelHeight * this.codedWidth, }); finalizationRegistry?.register(this, { type: 'video', data: this._data }, this); } /** Clones this video sample. */ clone() { if (this._closed) { throw new Error('VideoSample is closed.'); } assert(this._data !== null); if (this._data instanceof VideoSampleResource) { return new VideoSample(this._data, { timestamp: this.timestamp, duration: this.duration, rotation: this.rotation, encodeOptions: this.encodeOptions, }); } else if (isVideoFrame(this._data)) { return new VideoSample(this._data.clone(), { timestamp: this.timestamp, duration: this.duration, rotation: this.rotation, encodeOptions: this.encodeOptions, }); } else if (this._data instanceof Uint8Array) { assert(this._layout); return new VideoSample(this._data, { format: this.format!, layout: this._layout, codedWidth: this.codedWidth, codedHeight: this.codedHeight, timestamp: this.timestamp, duration: this.duration, colorSpace: this.colorSpace, rotation: this.rotation, visibleRect: this.visibleRect, displayWidth: this.displayWidth, displayHeight: this.displayHeight, encodeOptions: this.encodeOptions, // It's already been copied, if we copy it again we make the clone unnecessarily expensive _doNotCopy: true, }); } else { return new VideoSample(this._data, { format: this.format!, codedWidth: this.codedWidth, codedHeight: this.codedHeight, timestamp: this.timestamp, duration: this.duration, colorSpace: this.colorSpace, rotation: this.rotation, visibleRect: this.visibleRect, displayWidth: this.displayWidth, displayHeight: this.displayHeight, encodeOptions: this.encodeOptions, }); } } /** * Closes this video sample, releasing held resources. Video samples should be closed as soon as they are not * needed anymore. */ close() { if (this._closed) { return; } finalizationRegistry?.unregister(this); if (this._data instanceof VideoSampleResource) { this._data._referenceCount--; if (this._data._referenceCount === 0) { this._data.close(); } } else if (isVideoFrame(this._data)) { this._data.close(); } else { this._data = null; // GC that shit } this._closed = true; } /** * Returns the number of bytes required to hold this video sample's pixel data. */ allocationSize(options: VideoFrameCopyToOptions = {}): number { validateVideoFrameCopyToOptions(options); if (this._closed) { throw new Error('VideoSample is closed.'); } if ((options.format ?? this.format) == null) { // https://github.com/Vanilagy/mediabunny/issues/267 // https://github.com/w3c/webcodecs/issues/920 throw new Error('Cannot get allocation size when format is null.'); } if (isVideoFrame(this._data)) { // Call the native method purely for performance return this._data.allocationSize(options); } const combinedLayout = ParseVideoFrameCopyToOptions(this, options); return combinedLayout.allocationSize; } /** * Copies this video sample's pixel data to an ArrayBuffer or ArrayBufferView. * @returns The byte layout of the planes of the copied data. */ async copyTo(destination: AllowSharedBufferSource, options: VideoFrameCopyToOptions = {}): Promise { if (!isAllowSharedBufferSource(destination)) { throw new TypeError('destination must be an ArrayBuffer or an ArrayBuffer view.'); } validateVideoFrameCopyToOptions(options); if (this._closed) { throw new Error('VideoSample is closed.'); } if ((options.format ?? this.format) == null) { throw new Error('Cannot copy video sample data when format is null.'); } assert(this._data !== null); if (isVideoFrame(this._data)) { return this._data.copyTo(destination, options); } // Detect non-RGB to RGB conversion if ( options.format && !['RGBA', 'RGBX', 'BGRA', 'BGRX'].includes(this.format!) && ['RGBA', 'RGBX', 'BGRA', 'BGRX'].includes(options.format) ) { // RGB conversion for custom VideoSampleResource if (this._data instanceof VideoSampleResource) { using rgbSample = await this._data.toRgbSample( { timestamp: this.timestamp, duration: this.duration, rotation: this.rotation, }, options.colorSpace ?? 'srgb', ); if (!(rgbSample instanceof VideoSample)) { throw new TypeError('toRgbSample() must return a VideoSample.'); } if (!['RGBA', 'RGBX', 'BGRA', 'BGRX'].includes(rgbSample.format!)) { throw new Error( `Sample returned by toRgbSample was expected to have an RGB format, got` + ` '${rgbSample.format}' instead.`, ); } // Note that we DON'T force the RGB format to be exactly what was requested; any RGB format will do return await rgbSample.copyTo(destination, options); // 'await' is intentional here cuz of using } else { if (typeof VideoFrame === 'undefined') { throw new Error( 'For this sample, converting from a non-RGB to an RGB format requires VideoFrame to' + ' be defined.', ); } const tempFrame = this.toVideoFrame(); const result = await tempFrame.copyTo(destination, options); tempFrame.close(); return result; } } const combinedLayout = ParseVideoFrameCopyToOptions(this, options); assert(this.format); // 4. If destination.byteLength is less than combinedLayout’s allocationSize, return a promise rejected with const destBytes = toUint8Array(destination); if (destBytes.byteLength < combinedLayout.allocationSize) { throw new TypeError( `Destination buffer too small. Required: ${combinedLayout.allocationSize},` + ` Available: ${destBytes.byteLength}`, ); } const planeConfigs = getPlaneConfigs(this.format); let dataPlanes: VideoDataPlane[]; if (this._data instanceof VideoSampleResource) { let result = this._data.getDataPlanes(); if (result instanceof Promise) result = await result; if ( !Array.isArray(result) || result.some(x => !(x.data instanceof Uint8Array) || !Number.isInteger(x.stride) || x.stride < 0) ) { throw new TypeError( 'getDataPlanes() must return an array of objects with a Uint8Array "data" property and a' + ' non-negative integer "stride" property.', ); } dataPlanes = result; } else if (this._data instanceof Uint8Array) { assert(this._layout); assert(this._layout.length === planeConfigs.length); dataPlanes = this._layout.map((planeLayout, i) => { const height = Math.ceil(this.codedHeight / planeConfigs[i]!.heightDivisor); return { data: (this._data as Uint8Array).subarray( planeLayout.offset, planeLayout.offset + planeLayout.stride * height, ), stride: planeLayout.stride, }; }); } else { const canvas = this._data; const context = canvas.getContext('2d'); assert(context); // We already got it earlier so it's definitely available const imageData = context.getImageData(0, 0, this.codedWidth, this.codedHeight); dataPlanes = [{ data: toUint8Array(imageData.data), stride: 4 * this.codedWidth, }]; } // Algo taken from WebCodecs spec: // 6. Let p be a new Promise. (Implicit) // 7. Let copyStepsQueue be the result of starting a new parallel queue. (Implicit) // 8. Let planeLayouts be a new list. const planeLayouts: PlaneLayout[] = []; // Enqueue the following steps to copyStepsQueue: (fuck the queuing part) // Let resource be the media resource referenced by [[resource reference]]. // (this.data) // Let numPlanes be the number of planes as defined by [[format]]. const numPlanes = planeConfigs.length; // Let planeIndex be 0. // While planeIndex is less than combinedLayout’s numPlanes: for (let planeIndex = 0; planeIndex < numPlanes; planeIndex++) { const computedLayout = combinedLayout.computedLayouts[planeIndex]!; // Let sourceStride be the stride of the plane in resource as identified by planeIndex. const sourceStride = dataPlanes[planeIndex]!.stride; const sourceData = dataPlanes[planeIndex]!.data; // Let sourceOffset be the product of multiplying computedLayout’s sourceTop by sourceStride let sourceOffset = computedLayout.sourceTop * sourceStride; // Add computedLayout’s sourceLeftBytes to sourceOffset. sourceOffset += computedLayout.sourceLeftBytes; // Let destinationOffset be computedLayout’s destinationOffset. let destinationOffset = computedLayout.destinationOffset; // Let rowBytes be computedLayout’s sourceWidthBytes. const rowBytes = computedLayout.sourceWidthBytes; // Let layout be a new PlaneLayout, with offset set to destinationOffset and stride set to rowBytes. // This is a spec error actually (https://github.com/w3c/webcodecs/issues/918) const layout: PlaneLayout = { offset: destinationOffset, stride: computedLayout.destinationStride, }; // Let row be 0. // While row is less than computedLayout’s sourceHeight: for (let row = 0; row < computedLayout.sourceHeight; row++) { // Copy rowBytes bytes from resource starting at sourceOffset to destination starting // at destinationOffset. if (sourceOffset + rowBytes > sourceData.byteLength) { throw new Error(`Source buffer OOB read.`); } if (destinationOffset + rowBytes > destBytes.byteLength) { throw new Error(`Destination buffer OOB write.`); } const srcSub = sourceData.subarray(sourceOffset, sourceOffset + rowBytes); destBytes.set(srcSub, destinationOffset); // Increment sourceOffset by sourceStride. sourceOffset += sourceStride; // Increment destinationOffset by computedLayout’s destinationStride. destinationOffset += computedLayout.destinationStride; } // Append layout to planeLayouts. planeLayouts.push(layout); } // Now, handle converting between different RGB formats if (options.format !== undefined) { const needsRgbConversion = this.format.startsWith('RGB') !== options.format.startsWith('RGB'); // Going X->A requires setting the alpha to 255, going the other way doesn't since the value of X is w/e const needsAlphaConversion = this.format.includes('X') && options.format.includes('A'); if (needsRgbConversion || needsAlphaConversion) { // Loop over the destination bytes for (let i = 0; i < combinedLayout.allocationSize; i += 4) { if (needsRgbConversion) { // Swap R with B const r = destBytes[i]!; const b = destBytes[i + 2]!; destBytes[i] = b; destBytes[i + 2] = r; } if (needsAlphaConversion) { destBytes[i + 3] = 255; } } } } // Queue a task to resolve p with planeLayouts. return planeLayouts; } /** * Converts this video sample to a VideoFrame for use with the WebCodecs API. The VideoFrame returned by this * method *must* be closed separately from this video sample. */ toVideoFrame(): VideoFrame { if (this._closed) { throw new Error('VideoSample is closed.'); } assert(this._data !== null); if (this._data instanceof VideoSampleResource) { if (this.format === null) { throw new Error( 'Cannot convert a VideoSampleResource-backed VideoSample to VideoFrame if format is null.', ); } const planes = this._data.getDataPlanes(); if (planes instanceof Promise) { throw new Error( 'Cannot convert a VideoSampleResource-backed VideoSample to VideoFrame if getDataPlanes() returns' + ' a promise.', ); } // We can't use allocationSize since that method assumes a tight packing const size = planes.reduce((a, b) => a + b.data.byteLength, 0); const buffer = new Uint8Array(size); let offset = 0; const offsets: number[] = []; for (const plane of planes) { buffer.set(plane.data, offset); offsets.push(offset); offset += plane.data.byteLength; } return new VideoFrame(buffer, { format: this.format as VideoPixelFormat, layout: planes.map((x, i) => ({ offset: offsets[i]!, stride: x.stride, })), codedWidth: this.codedWidth, codedHeight: this.codedHeight, timestamp: this.microsecondTimestamp, duration: this.microsecondDuration, colorSpace: this.colorSpace, displayWidth: this.squarePixelWidth, // Not display* since we're not passing rotation displayHeight: this.squarePixelHeight, }); } else if (isVideoFrame(this._data)) { return new VideoFrame(this._data, { timestamp: this.microsecondTimestamp, duration: this.microsecondDuration || undefined, // Drag 0 duration to undefined, glitches some codecs }); } else if (this._data instanceof Uint8Array) { return new VideoFrame(this._data, { format: this.format! as VideoPixelFormat, codedWidth: this.codedWidth, codedHeight: this.codedHeight, timestamp: this.microsecondTimestamp, duration: this.microsecondDuration || undefined, colorSpace: this.colorSpace, displayWidth: this.squarePixelWidth, // Not display* since we're not passing rotation displayHeight: this.squarePixelHeight, }); } else { return new VideoFrame(this._data, { timestamp: this.microsecondTimestamp, duration: this.microsecondDuration || undefined, }); } } /** * Draws the video sample to a 2D canvas context. Rotation metadata will be taken into account. * * @param dx - The x-coordinate in the destination canvas at which to place the top-left corner of the source image. * @param dy - The y-coordinate in the destination canvas at which to place the top-left corner of the source image. * @param dWidth - The width in pixels with which to draw the image in the destination canvas. * @param dHeight - The height in pixels with which to draw the image in the destination canvas. */ draw( context: CanvasRenderingContext2D | OffscreenCanvasRenderingContext2D, dx: number, dy: number, dWidth?: number, dHeight?: number, ): void; /** * Draws the video sample to a 2D canvas context. Rotation metadata will be taken into account. * * @param sx - The x-coordinate of the top left corner of the sub-rectangle of the source image to draw into the * destination context. * @param sy - The y-coordinate of the top left corner of the sub-rectangle of the source image to draw into the * destination context. * @param sWidth - The width of the sub-rectangle of the source image to draw into the destination context. * @param sHeight - The height of the sub-rectangle of the source image to draw into the destination context. * @param dx - The x-coordinate in the destination canvas at which to place the top-left corner of the source image. * @param dy - The y-coordinate in the destination canvas at which to place the top-left corner of the source image. * @param dWidth - The width in pixels with which to draw the image in the destination canvas. * @param dHeight - The height in pixels with which to draw the image in the destination canvas. */ draw( context: CanvasRenderingContext2D | OffscreenCanvasRenderingContext2D, sx: number, sy: number, sWidth: number, sHeight: number, dx: number, dy: number, dWidth?: number, dHeight?: number, ): void; draw( context: CanvasRenderingContext2D | OffscreenCanvasRenderingContext2D, arg1: number, arg2: number, arg3?: number, arg4?: number, arg5?: number, arg6?: number, arg7?: number, arg8?: number, ) { let sx = 0; let sy = 0; let sWidth = this.displayWidth; let sHeight = this.displayHeight; let dx = 0; let dy = 0; let dWidth = this.displayWidth; let dHeight = this.displayHeight; if (arg5 !== undefined) { sx = arg1!; sy = arg2!; sWidth = arg3!; sHeight = arg4!; dx = arg5; dy = arg6!; if (arg7 !== undefined) { dWidth = arg7; dHeight = arg8!; } else { dWidth = sWidth; dHeight = sHeight; } } else { dx = arg1; dy = arg2; if (arg3 !== undefined) { dWidth = arg3; dHeight = arg4!; } } if (!( (typeof CanvasRenderingContext2D !== 'undefined' && context instanceof CanvasRenderingContext2D) || ( typeof OffscreenCanvasRenderingContext2D !== 'undefined' && context instanceof OffscreenCanvasRenderingContext2D ) )) { throw new TypeError('context must be a CanvasRenderingContext2D or OffscreenCanvasRenderingContext2D.'); } if (!Number.isFinite(sx)) { throw new TypeError('sx must be a number.'); } if (!Number.isFinite(sy)) { throw new TypeError('sy must be a number.'); } if (!Number.isFinite(sWidth) || sWidth < 0) { throw new TypeError('sWidth must be a non-negative number.'); } if (!Number.isFinite(sHeight) || sHeight < 0) { throw new TypeError('sHeight must be a non-negative number.'); } if (!Number.isFinite(dx)) { throw new TypeError('dx must be a number.'); } if (!Number.isFinite(dy)) { throw new TypeError('dy must be a number.'); } if (!Number.isFinite(dWidth) || dWidth < 0) { throw new TypeError('dWidth must be a non-negative number.'); } if (!Number.isFinite(dHeight) || dHeight < 0) { throw new TypeError('dHeight must be a non-negative number.'); } if (this._closed) { throw new Error('VideoSample is closed.'); } ({ sx, sy, sWidth, sHeight } = this._rotateSourceRegion(sx, sy, sWidth, sHeight, this.rotation)); const source = this.toCanvasImageSource(); context.save(); const centerX = dx + dWidth / 2; const centerY = dy + dHeight / 2; context.translate(centerX, centerY); context.rotate(this.rotation * Math.PI / 180); const aspectRatioChange = this.rotation % 180 === 0 ? 1 : dWidth / dHeight; // Scale to compensate for aspect ratio changes when rotated context.scale(1 / aspectRatioChange, aspectRatioChange); context.drawImage( source, sx, sy, sWidth, sHeight, -dWidth / 2, -dHeight / 2, dWidth, dHeight, ); context.restore(); } /** * Draws the sample in the middle of the canvas corresponding to the context with the specified fit behavior. */ drawWithFit(context: CanvasRenderingContext2D | OffscreenCanvasRenderingContext2D, options: { /** * Controls the fitting algorithm. * * - `'fill'` will stretch the image to fill the entire box, potentially altering aspect ratio. * - `'contain'` will contain the entire image within the box while preserving aspect ratio. This may lead to * letterboxing. * - `'cover'` will scale the image until the entire box is filled, while preserving aspect ratio. */ fit: 'fill' | 'contain' | 'cover'; /** A way to override rotation. Defaults to the rotation of the sample. */ rotation?: Rotation; /** * Specifies the rectangular region of the video sample to crop to. The crop region will automatically be * clamped to the dimensions of the video sample. Cropping is performed after rotation but before resizing. * The crop region is in the _display pixel space_ of the underlying video data. */ crop?: CropRectangle; }) { if (!( (typeof CanvasRenderingContext2D !== 'undefined' && context instanceof CanvasRenderingContext2D) || ( typeof OffscreenCanvasRenderingContext2D !== 'undefined' && context instanceof OffscreenCanvasRenderingContext2D ) )) { throw new TypeError('context must be a CanvasRenderingContext2D or OffscreenCanvasRenderingContext2D.'); } if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (!['fill', 'contain', 'cover'].includes(options.fit)) { throw new TypeError('options.fit must be \'fill\', \'contain\', or \'cover\'.'); } if (options.rotation !== undefined && ![0, 90, 180, 270].includes(options.rotation)) { throw new TypeError('options.rotation, when provided, must be 0, 90, 180, or 270.'); } if (options.crop !== undefined) { validateCropRectangle(options.crop, 'options.'); } const canvasWidth = context.canvas.width; const canvasHeight = context.canvas.height; const rotation = options.rotation ?? this.rotation; const [rotatedWidth, rotatedHeight] = rotation % 180 === 0 ? [this.squarePixelWidth, this.squarePixelHeight] : [this.squarePixelHeight, this.squarePixelWidth]; let finalCrop = options.crop; if (finalCrop) { finalCrop = clampCropRectangle(finalCrop, rotatedWidth, rotatedHeight); } // These variables specify where the final sample will be drawn on the canvas let dx: number; let dy: number; let newWidth: number; let newHeight: number; const { sx, sy, sWidth, sHeight } = this._rotateSourceRegion( options.crop?.left ?? 0, options.crop?.top ?? 0, options.crop?.width ?? rotatedWidth, options.crop?.height ?? rotatedHeight, rotation, ); if (options.fit === 'fill') { dx = 0; dy = 0; newWidth = canvasWidth; newHeight = canvasHeight; } else { const [sampleWidth, sampleHeight] = options.crop ? [options.crop.width, options.crop.height] : [rotatedWidth, rotatedHeight]; const scale = options.fit === 'contain' ? Math.min(canvasWidth / sampleWidth, canvasHeight / sampleHeight) : Math.max(canvasWidth / sampleWidth, canvasHeight / sampleHeight); newWidth = sampleWidth * scale; newHeight = sampleHeight * scale; dx = (canvasWidth - newWidth) / 2; dy = (canvasHeight - newHeight) / 2; } context.save(); const aspectRatioChange = rotation % 180 === 0 ? 1 : newWidth / newHeight; context.translate(canvasWidth / 2, canvasHeight / 2); context.rotate(rotation * Math.PI / 180); // This aspect ratio compensation is done so that we can draw the sample with the intended dimensions and // don't need to think about how those dimensions change after the rotation context.scale(1 / aspectRatioChange, aspectRatioChange); context.translate(-canvasWidth / 2, -canvasHeight / 2); // Important that we don't use .draw() here since that would take rotation into account, but we wanna handle it // ourselves here context.drawImage(this.toCanvasImageSource(), sx, sy, sWidth, sHeight, dx, dy, newWidth, newHeight); context.restore(); } /** @internal */ _rotateSourceRegion(sx: number, sy: number, sWidth: number, sHeight: number, rotation: number) { // The provided sx,sy,sWidth,sHeight refer to the final rotated image, but that's not actually how the image is // stored. Therefore, we must map these back onto the original, pre-rotation image. if (rotation === 90) { [sx, sy, sWidth, sHeight] = [ sy, this.squarePixelHeight - sx - sWidth, sHeight, sWidth, ]; } else if (rotation === 180) { [sx, sy] = [ this.squarePixelWidth - sx - sWidth, this.squarePixelHeight - sy - sHeight, ]; } else if (rotation === 270) { [sx, sy, sWidth, sHeight] = [ this.squarePixelWidth - sy - sHeight, sx, sHeight, sWidth, ]; } return { sx, sy, sWidth, sHeight }; } /** * Converts this video sample to a * [`CanvasImageSource`](https://udn.realityripple.com/docs/Web/API/CanvasImageSource) for drawing to a canvas. * * You must use the value returned by this method immediately, as any VideoFrame created internally may * automatically be closed in the next microtask. */ toCanvasImageSource() { if (this._closed) { throw new Error('VideoSample is closed.'); } assert(this._data !== null); if (this._data instanceof VideoSampleResource || this._data instanceof Uint8Array) { // Requires VideoFrame to be defined const videoFrame = this.toVideoFrame(); queueMicrotask(() => videoFrame.close()); // Let's automatically close the frame in the next microtask return videoFrame; } else { return this._data; } } /** * Transform this video sample to a new video sample given the options. Can be used to resize, rotate, and crop * the sample. * * In non-browser environments, this method will not work by default. To make it work, register a custom * transformer function via {@link registerVideoSampleTransformer}. */ async transform(options: VideoSampleTransformOptions) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.width !== undefined && (!Number.isInteger(options.width) || options.width <= 0)) { throw new TypeError('options.width, when provided, must be a positive integer.'); } if (options.height !== undefined && (!Number.isInteger(options.height) || options.height <= 0)) { throw new TypeError('options.height, when provided, must be a positive integer.'); } if ( options.roundDimensionsTo !== undefined && (!Number.isInteger(options.roundDimensionsTo) || options.roundDimensionsTo <= 0) ) { throw new TypeError('options.roundDimensionsTo, when provided, must be a positive integer.'); } if (options.fit !== undefined && !['fill', 'contain', 'cover'].includes(options.fit)) { throw new TypeError('options.fit, when provided, must be one of "fill", "contain", or "cover".'); } if ( options.width !== undefined && options.height !== undefined && options.fit === undefined ) { throw new TypeError( 'When both options.width and options.height are provided, options.fit must also be provided.', ); } if (options.rotate !== undefined && ![0, 90, 180, 270].includes(options.rotate)) { throw new TypeError('options.rotate, when provided, must be 0, 90, 180 or 270.'); } if (options.crop !== undefined) { validateCropRectangle(options.crop, 'options.'); } if (options.alpha !== undefined && !['keep', 'discard'].includes(options.alpha)) { throw new TypeError('options.alpha, when provided, must be \'keep\' or \'discard\'.'); } const rotation = normalizeRotation(this.rotation + (options.rotate ?? 0)); const [rotatedWidth, rotatedHeight] = rotation % 180 === 0 ? [this.squarePixelWidth, this.squarePixelHeight] : [this.squarePixelHeight, this.squarePixelWidth]; // Clamp crop rectangle to the rotated video dimensions let finalCrop = options.crop; if (finalCrop) { finalCrop = clampCropRectangle(finalCrop, rotatedWidth, rotatedHeight); } const cropWidth = finalCrop ? finalCrop.width : rotatedWidth; const cropHeight = finalCrop ? finalCrop.height : rotatedHeight; const originalAspectRatio = cropWidth / cropHeight; let targetWidth: number; let targetHeight: number; if (options.width !== undefined && options.height === undefined) { targetWidth = options.width; targetHeight = targetWidth / originalAspectRatio; } else if (options.width === undefined && options.height !== undefined) { targetHeight = options.height; targetWidth = targetHeight * originalAspectRatio; } else if (options.width !== undefined && options.height !== undefined) { targetWidth = options.width; targetHeight = options.height; } else { targetWidth = cropWidth; targetHeight = cropHeight; } targetWidth = roundToMultiple(targetWidth, options.roundDimensionsTo ?? 1); targetHeight = roundToMultiple(targetHeight, options.roundDimensionsTo ?? 1); const description: VideoSampleTransformationDescription = { width: targetWidth, height: targetHeight, fit: options.fit ?? 'fill', rotation, crop: finalCrop ?? { left: 0, top: 0, width: rotatedWidth, height: rotatedHeight, }, alpha: options.alpha ?? 'keep', }; // Description's finalized; let's see if a registered transformer wants to handle it for (const transformer of registeredVideoSampleTransformers) { let result = transformer(this, description); if (result instanceof Promise) result = await result; if (result !== null) { return result; } } // We need to handle it ourselves, and we use canvases to do it let canvas: HTMLCanvasElement | OffscreenCanvas | null = null; let canvasIsNew = false; for (const entry of transformationCanvasCache) { if (entry.canvas.width === description.width && entry.canvas.height === description.height) { canvas = entry.canvas; entry.age = transformationCanvasCacheNextAge++; break; } } if (canvas === null) { if (typeof OffscreenCanvas !== 'undefined') { canvas = new OffscreenCanvas(description.width, description.height); } else { if (typeof window === 'undefined' || typeof document === 'undefined') { throw new Error( 'Cannot transform VideoSamples in this environment. Either run in an environment with' + ' OffscreenCanvas or HTMLCanvasElement, or supply a custom VideoSample transformer using' + ' registerVideoSampleTransformer().', ); } canvas = document.createElement('canvas'); canvas.width = description.width; canvas.height = description.height; } canvasIsNew = true; if (transformationCanvasCache.length >= TRANSFORMATION_CANVAS_CACHE_MAX_SIZE) { transformationCanvasCache.splice(arrayArgmin(transformationCanvasCache, x => x.age), 1); } transformationCanvasCache.push({ canvas, age: transformationCanvasCacheNextAge++, }); } const context = canvas.getContext('2d', { alpha: true, }) as CanvasRenderingContext2D | OffscreenCanvasRenderingContext2D; if (!context) { throw new Error( 'The \'2d\' canvas context is required to transform VideoSamples. Register a custom transformer using' + ' registerVideoSampleTransformer to work around this limitation.', ); } if (description.alpha === 'discard') { context.fillStyle = 'black'; context.fillRect(0, 0, description.width, description.height); } else if (!canvasIsNew) { // Cached canvases carry stale pixels from a prior draw context.clearRect(0, 0, description.width, description.height); } this.drawWithFit(context, { fit: description.fit, rotation: description.rotation, crop: description.crop, }); return new VideoSample(canvas, { timestamp: this.timestamp, duration: this.duration, rotation: 0, // Any previous rotation is now baked in }); } /** Sets the rotation metadata of this video sample. */ setRotation(newRotation: Rotation) { if (![0, 90, 180, 270].includes(newRotation)) { throw new TypeError('newRotation must be 0, 90, 180, or 270.'); } // eslint-disable-next-line @typescript-eslint/no-unnecessary-type-assertion (this.rotation as Rotation) = newRotation; } /** Sets the presentation timestamp of this video sample, in seconds. */ setTimestamp(newTimestamp: number) { if (!Number.isFinite(newTimestamp)) { throw new TypeError('newTimestamp must be a number.'); } // eslint-disable-next-line @typescript-eslint/no-unnecessary-type-assertion (this.timestamp as number) = newTimestamp; } /** Sets the duration of this video sample, in seconds. */ setDuration(newDuration: number) { if (!Number.isFinite(newDuration) || newDuration < 0) { throw new TypeError('newDuration must be a non-negative number.'); } // eslint-disable-next-line @typescript-eslint/no-unnecessary-type-assertion (this.duration as number) = newDuration; } /** Sets the encode options used when this sample is passed to an encoder. */ setEncodeOptions(newEncodeOptions: VideoEncoderEncodeOptions) { if (!newEncodeOptions || typeof newEncodeOptions !== 'object') { throw new TypeError('newEncodeOptions must be an object.'); } // eslint-disable-next-line @typescript-eslint/no-unnecessary-type-assertion (this.encodeOptions as DeepReadonly) = newEncodeOptions; } /** Calls `.close()`. */ [Symbol.dispose]() { this.close(); } } /** * Options for transforming a {@link VideoSample}. The order of operations are: * * 1. Pixel aspect ratio normalization (always applied) * 2. Rotation * 3. Crop * 4. Resize using fit * @group Samples * @public */ export type VideoSampleTransformOptions = { /** * The width in pixels to resize the frames to. If height is not set, it will be deduced * automatically based on aspect ratio. */ width?: number; /** * The height in pixels to resize the frames to. If width is not set, it will be deduced * automatically based on aspect ratio. */ height?: number; /** * A positive integer. When provided, both the width and height will be rounded to the nearest multiple of * this number. */ roundDimensionsTo?: number; /** * The fitting algorithm in case both width and height are set. * * - `'fill'` will stretch the image to fill the entire box, potentially altering aspect ratio. * - `'contain'` will contain the entire image within the box while preserving aspect ratio. This may lead to * letterboxing. * - `'cover'` will scale the image until the entire box is filled, while preserving aspect ratio. */ fit?: 'fill' | 'contain' | 'cover'; /** * The clockwise rotation by which to rotate the frames. Rotation is applied before resizing. */ rotate?: Rotation; /** * Specifies the rectangular region of the frames to crop to. The crop region will automatically be * clamped to the dimensions of the frame. Cropping is performed after rotation but before resizing. */ crop?: CropRectangle; /** * Whether to discard or keep the transparency information of the video sample. The default is `'keep'`. */ alpha?: 'keep' | 'discard'; }; /** * A fully-resolved description of a video sample transformation, with all defaults and constraints baked in. * * The order of operations must be: * 1. Pixel aspect ratio normalization (always applied) * 2. Rotation * 3. Crop * 4. Resize using fit * @group Samples * @public */ export type VideoSampleTransformationDescription = { /** The width in pixels to resize the frames to. */ width: number; /** The height in pixels to resize the frames to. */ height: number; /** * The fitting algorithm. * * - `'fill'` will stretch the image to fill the entire box, potentially altering aspect ratio. * - `'contain'` will contain the entire image within the box while preserving aspect ratio. This may lead to * letterboxing. * - `'cover'` will scale the image until the entire box is filled, while preserving aspect ratio. */ fit: 'fill' | 'contain' | 'cover'; /** The clockwise rotation by which to rotate the frames. Rotation is applied before resizing. */ rotation: Rotation; /** * The rectangular region of the frames to crop to, clamped to the dimensions of the frame. Cropping is * performed after rotation but before resizing. */ crop: CropRectangle; /** Whether to discard or keep the transparency information of the video sample. */ alpha: 'keep' | 'discard'; }; const registeredVideoSampleTransformers: (( sample: VideoSample, description: VideoSampleTransformationDescription, ) => MaybePromise)[] = []; /** * Registers a callback to handle the transformation of {@link VideoSample} instances. The callback can either return * the transformed sample, or `null` to indicate that it doesn't want to handle the given transformation task. * @group Samples * @public */ export const registerVideoSampleTransformer = ( transformer: ( sample: VideoSample, description: VideoSampleTransformationDescription, ) => MaybePromise, ) => { if (registeredVideoSampleTransformers.includes(transformer)) { return; // Already in there } registeredVideoSampleTransformers.push(transformer); }; const TRANSFORMATION_CANVAS_CACHE_MAX_SIZE = 3; const transformationCanvasCache: { canvas: HTMLCanvasElement | OffscreenCanvas; age: number; }[] = []; let transformationCanvasCacheNextAge = 0; /** * Describes the color space of a {@link VideoSample}. Corresponds to the WebCodecs API's VideoColorSpace. * @group Samples * @public */ export class VideoSampleColorSpace { /** The color primaries standard used. */ readonly primaries: VideoColorPrimaries | null; /** The transfer characteristics used. */ readonly transfer: VideoTransferCharacteristics | null; /** The color matrix coefficients used. */ readonly matrix: VideoMatrixCoefficients | null; /** Whether the color values use the full range or limited range. */ readonly fullRange: boolean | null; /** Creates a new VideoSampleColorSpace. */ constructor(init?: VideoColorSpaceInit) { if (init !== undefined) { if (!init || typeof init !== 'object') { throw new TypeError('init.colorSpace, when provided, must be an object.'); } const primariesValues = Object.keys(COLOR_PRIMARIES_MAP); if (init.primaries != null && !primariesValues.includes(init.primaries)) { throw new TypeError( `init.colorSpace.primaries, when provided, must be one of ${primariesValues.join(', ')}.`, ); } const transferValues = Object.keys(TRANSFER_CHARACTERISTICS_MAP); if (init.transfer != null && !transferValues.includes(init.transfer)) { throw new TypeError( `init.colorSpace.transfer, when provided, must be one of ${transferValues.join(', ')}.`, ); } const matrixValues = Object.keys(MATRIX_COEFFICIENTS_MAP); if (init.matrix != null && !matrixValues.includes(init.matrix)) { throw new TypeError( `init.colorSpace.matrix, when provided, must be one of ${matrixValues.join(', ')}.`, ); } if (init.fullRange != null && typeof init.fullRange !== 'boolean') { throw new TypeError('init.colorSpace.fullRange, when provided, must be a boolean.'); } } this.primaries = init?.primaries ?? null; this.transfer = init?.transfer ?? null; this.matrix = init?.matrix ?? null; this.fullRange = init?.fullRange ?? null; } /** Serializes the color space to a JSON object. */ toJSON(): VideoColorSpaceInit { return { primaries: this.primaries, transfer: this.transfer, matrix: this.matrix, fullRange: this.fullRange, }; } } const isVideoFrame = (x: unknown): x is VideoFrame => { return typeof VideoFrame !== 'undefined' && x instanceof VideoFrame; }; /** * Specifies the rectangular cropping region. * @group Miscellaneous * @public */ export type CropRectangle = { /** The distance in pixels from the left edge of the source frame to the left edge of the crop rectangle. */ left: number; /** The distance in pixels from the top edge of the source frame to the top edge of the crop rectangle. */ top: number; /** The width in pixels of the crop rectangle. */ width: number; /** The height in pixels of the crop rectangle. */ height: number; }; export const clampCropRectangle = (crop: CropRectangle, outerWidth: number, outerHeight: number): CropRectangle => { const left = Math.min(crop.left, outerWidth); const top = Math.min(crop.top, outerHeight); const width = Math.min(crop.width, outerWidth - left); const height = Math.min(crop.height, outerHeight - top); assert(width >= 0); assert(height >= 0); return { left, top, width, height }; }; export const validateCropRectangle = (crop: CropRectangle, prefix: string) => { if (!crop || typeof crop !== 'object') { throw new TypeError(prefix + 'crop, when provided, must be an object.'); } if (!Number.isInteger(crop.left) || crop.left < 0) { throw new TypeError(prefix + 'crop.left must be a non-negative integer.'); } if (!Number.isInteger(crop.top) || crop.top < 0) { throw new TypeError(prefix + 'crop.top must be a non-negative integer.'); } if (!Number.isInteger(crop.width) || crop.width < 0) { throw new TypeError(prefix + 'crop.width must be a non-negative integer.'); } if (!Number.isInteger(crop.height) || crop.height < 0) { throw new TypeError(prefix + 'crop.height must be a non-negative integer.'); } }; const validateVideoFrameCopyToOptions = (options: VideoFrameCopyToOptions) => { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.colorSpace !== undefined && !['display-p3', 'srgb'].includes(options.colorSpace)) { throw new TypeError('options.colorSpace, when provided, must be \'display-p3\' or \'srgb\'.'); } if (options.format !== undefined && typeof options.format !== 'string') { throw new TypeError('options.format, when provided, must be a string.'); } if (options.layout !== undefined) { if (!Array.isArray(options.layout)) { throw new TypeError('options.layout, when provided, must be an array.'); } for (const plane of options.layout) { if (!plane || typeof plane !== 'object') { throw new TypeError('Each entry in options.layout must be an object.'); } if (!Number.isInteger(plane.offset) || plane.offset < 0) { throw new TypeError('plane.offset must be a non-negative integer.'); } if (!Number.isInteger(plane.stride) || plane.stride < 0) { throw new TypeError('plane.stride must be a non-negative integer.'); } } } if (options.rect !== undefined) { if (!options.rect || typeof options.rect !== 'object') { throw new TypeError('options.rect, when provided, must be an object.'); } if (options.rect.x !== undefined && (!Number.isInteger(options.rect.x) || options.rect.x < 0)) { throw new TypeError('options.rect.x, when provided, must be a non-negative integer.'); } if (options.rect.y !== undefined && (!Number.isInteger(options.rect.y) || options.rect.y < 0)) { throw new TypeError('options.rect.y, when provided, must be a non-negative integer.'); } if (options.rect.width !== undefined && (!Number.isInteger(options.rect.width) || options.rect.width < 0)) { throw new TypeError('options.rect.width, when provided, must be a non-negative integer.'); } if (options.rect.height !== undefined && (!Number.isInteger(options.rect.height) || options.rect.height < 0)) { throw new TypeError('options.rect.height, when provided, must be a non-negative integer.'); } } }; /** Implements logic from WebCodecs § 9.4.6 "Compute Layout and Allocation Size" */ const createDefaultPlaneLayout = ( format: VideoSamplePixelFormat, codedWidth: number, codedHeight: number, ): PlaneLayout[] => { const planes = getPlaneConfigs(format); const layouts: PlaneLayout[] = []; let currentOffset = 0; for (const plane of planes) { // Per § 9.8, dimensions are usually "rounded up to the nearest integer". const planeWidth = Math.ceil(codedWidth / plane.widthDivisor); const planeHeight = Math.ceil(codedHeight / plane.heightDivisor); const stride = planeWidth * plane.sampleBytes; // Tight packing const planeSize = stride * planeHeight; layouts.push({ offset: currentOffset, stride: stride, }); currentOffset += planeSize; } return layouts; }; type PlaneConfig = { sampleBytes: number; widthDivisor: number; // Horizontal sub-sampling factor heightDivisor: number; // Vertical sub-sampling factor }; /** Helper to retrieve plane configurations based on WebCodecs § 9.8 Pixel Format definitions. */ export const getPlaneConfigs = (format: VideoSamplePixelFormat): PlaneConfig[] => { // Helper for standard YUV planes const yuv = ( yBytes: number, uvBytes: number, subX: number, subY: number, hasAlpha: boolean, ): PlaneConfig[] => { const configs: PlaneConfig[] = [ { sampleBytes: yBytes, widthDivisor: 1, heightDivisor: 1 }, { sampleBytes: uvBytes, widthDivisor: subX, heightDivisor: subY }, { sampleBytes: uvBytes, widthDivisor: subX, heightDivisor: subY }, ]; if (hasAlpha) { // Match luma dimensions configs.push({ sampleBytes: yBytes, widthDivisor: 1, heightDivisor: 1 }); } return configs; }; switch (format) { case 'I420': return yuv(1, 1, 2, 2, false); case 'I420P10': case 'I420P12': return yuv(2, 2, 2, 2, false); case 'I420A': return yuv(1, 1, 2, 2, true); case 'I420AP10': case 'I420AP12': return yuv(2, 2, 2, 2, true); case 'I422': return yuv(1, 1, 2, 1, false); case 'I422P10': case 'I422P12': return yuv(2, 2, 2, 1, false); case 'I422A': return yuv(1, 1, 2, 1, true); case 'I422AP10': case 'I422AP12': return yuv(2, 2, 2, 1, true); case 'I444': return yuv(1, 1, 1, 1, false); case 'I444P10': case 'I444P12': return yuv(2, 2, 1, 1, false); case 'I444A': return yuv(1, 1, 1, 1, true); case 'I444AP10': case 'I444AP12': return yuv(2, 2, 1, 1, true); case 'NV12': return [ { sampleBytes: 1, widthDivisor: 1, heightDivisor: 1 }, { sampleBytes: 2, widthDivisor: 2, heightDivisor: 2 }, // Interleaved U and V ]; case 'RGBA': case 'RGBX': case 'BGRA': case 'BGRX': return [ { sampleBytes: 4, widthDivisor: 1, heightDivisor: 1 }, ]; default: assertNever(format); assert(false); } }; type CombinedBufferLayout = { allocationSize: number; computedLayouts: ComputedPlaneLayout[]; }; type ComputedPlaneLayout = { destinationOffset: number; destinationStride: number; sourceTop: number; sourceHeight: number; sourceLeftBytes: number; sourceWidthBytes: number; }; /** Taken from the WebCodecs spec. */ const ParseVideoFrameCopyToOptions = ( sample: VideoSample, options: VideoFrameCopyToOptions, ): CombinedBufferLayout => { // 1. Let defaultRect be the result of performing the getter steps for visibleRect. const defaultRect: Rectangle = { left: 0, top: 0, width: sample.codedWidth, height: sample.codedHeight, }; // 2. Let overrideRect be undefined. // 3. If options.rect exists, assign the value of options.rect to overrideRect. const overrideRect = options.rect; // 4. Let parsedRect be the result of running the Parse Visible Rect algorithm... const parsedRect = ParseVisibleRect( defaultRect, overrideRect, sample.codedWidth, sample.codedHeight, sample.format, ); // 5. If parsedRect is an exception, return parsedRect. (Handled by throw) // 6. Let optLayout be undefined. // 7. If options.layout exists, assign its value to optLayout. const optLayout = options.layout; // 8. Let format be undefined. let format: VideoSamplePixelFormat | undefined; // 9. If options.format does not exist, assign [[format]] to format. if (!options.format || options.format === sample.format) { format = sample.format!; } else if (['RGBA', 'RGBX', 'BGRA', 'BGRX'].includes(options.format)) { // 10. Otherwise, if options.format is equal to one of RGBA, RGBX, BGRA, BGRX, then assign options.format // to format... format = options.format; } else { throw new Error('NotSupportedError: Invalid destination format.'); } // 11. Let combinedLayout be the result of running the Compute Layout and Allocation Size algorithm... return ComputeLayoutAndAllocationSize(parsedRect, format, optLayout); }; /** Taken from the WebCodecs spec. */ const ParseVisibleRect = ( defaultRect: DOMRectInit, overrideRect: DOMRectInit | undefined, codedWidth: number, codedHeight: number, format: VideoSamplePixelFormat | null, ): DOMRectInit => { // 1. Let sourceRect be defaultRect const sourceRect = { ...defaultRect }; // 2. If overrideRect is not undefined: if (overrideRect !== undefined) { // If either of overrideRect.width or height is 0, return a TypeError. if (overrideRect.width === 0 || overrideRect.height === 0) { throw new TypeError('visibleRect dimensions cannot be zero.'); } // If the sum of overrideRect.x and overrideRect.width is greater than codedWidth, return a TypeError. if ((overrideRect.x || 0) + (overrideRect.width || 0) > codedWidth) { throw new TypeError('visibleRect exceeds codedWidth.'); } // If the sum of overrideRect.y and overrideRect.height is greater than codedHeight, return a TypeError. if ((overrideRect.y || 0) + (overrideRect.height || 0) > codedHeight) { throw new TypeError('visibleRect exceeds codedHeight.'); } // Assign overrideRect to sourceRect. sourceRect.x = overrideRect.x || 0; sourceRect.y = overrideRect.y || 0; sourceRect.width = overrideRect.width || 0; sourceRect.height = overrideRect.height || 0; } // 3. Let validAlignment be the result of running the Verify Rect Offset Alignment algorithm. const validAlignment = VerifyRectOffsetAlignment(format, sourceRect); // 4. If validAlignment is false, throw a TypeError. if (!validAlignment) { throw new TypeError('visibleRect alignment is invalid for the format.'); } // 5. Return sourceRect. return sourceRect; }; /** Taken from the WebCodecs spec. */ const VerifyRectOffsetAlignment = (format: VideoSamplePixelFormat | null, rect: DOMRectInit): boolean => { // 1. If format is null, return true. if (format === null) return true; const planes = getPlaneConfigs(format); // 2. Let planeIndex be 0. // 3. Let numPlanes be the number of planes as defined by format. // 4. While planeIndex is less than numPlanes: for (let planeIndex = 0; planeIndex < planes.length; planeIndex++) { const plane = planes[planeIndex]!; const sampleWidth = plane.widthDivisor; const sampleHeight = plane.heightDivisor; // If rect.x is not a multiple of sampleWidth, return false. if ((rect.x || 0) % sampleWidth !== 0) return false; // If rect.y is not a multiple of sampleHeight, return false. if ((rect.y || 0) % sampleHeight !== 0) return false; } return true; }; /** Taken from the WebCodecs spec. */ const ComputeLayoutAndAllocationSize = ( parsedRect: DOMRectInit, format: VideoSamplePixelFormat, layout?: PlaneLayout[], ): CombinedBufferLayout => { const planes = getPlaneConfigs(format); // 1. Let numPlanes be the number of planes as defined by format. const numPlanes = planes.length; // 2. If layout is not undefined and its length does not equal numPlanes, throw a TypeError. if (layout !== undefined && layout.length !== numPlanes) { throw new TypeError(`Layout must have ${numPlanes} planes.`); } // 3. Let minAllocationSize be 0. let minAllocationSize = 0; // 4. Let computedLayouts be a new list. const computedLayouts: ComputedPlaneLayout[] = []; // 5. Let endOffsets be a new list. const endOffsets: number[] = []; // 6. Let planeIndex be 0. // 7. While planeIndex < numPlanes: for (let planeIndex = 0; planeIndex < numPlanes; planeIndex++) { const plane = planes[planeIndex]!; const sampleBytes = plane.sampleBytes; const sampleWidth = plane.widthDivisor; const sampleHeight = plane.heightDivisor; // Let computedLayout be a new computed plane layout. const computedLayout: ComputedPlaneLayout = { destinationOffset: 0, destinationStride: 0, sourceTop: 0, sourceHeight: 0, sourceLeftBytes: 0, sourceWidthBytes: 0, }; // Set computedLayout’s sourceTop... computedLayout.sourceTop = Math.ceil(Math.trunc(parsedRect.y || 0) / sampleHeight); // Set computedLayout’s sourceHeight... computedLayout.sourceHeight = Math.ceil(Math.trunc(parsedRect.height || 0) / sampleHeight); // Set computedLayout’s sourceLeftBytes... computedLayout.sourceLeftBytes = Math.floor(Math.trunc(parsedRect.x || 0) / sampleWidth) * sampleBytes; // Set computedLayout’s sourceWidthBytes... computedLayout.sourceWidthBytes = Math.floor(Math.trunc(parsedRect.width || 0) / sampleWidth) * sampleBytes; // If layout is not undefined: if (layout !== undefined) { const planeLayout = layout[planeIndex]!; // If planeLayout.stride is less than computedLayout’s sourceWidthBytes, return a TypeError. if (planeLayout.stride < computedLayout.sourceWidthBytes) { throw new TypeError(`Stride for plane ${planeIndex} is too small.`); } // Assign planeLayout.offset to computedLayout’s destinationOffset. computedLayout.destinationOffset = planeLayout.offset; // Assign planeLayout.stride to computedLayout’s destinationStride. computedLayout.destinationStride = planeLayout.stride; } else { // Otherwise: // Assign minAllocationSize to computedLayout’s destinationOffset. computedLayout.destinationOffset = minAllocationSize; // Assign computedLayout’s sourceWidthBytes to computedLayout’s destinationStride. computedLayout.destinationStride = computedLayout.sourceWidthBytes; } // Let planeSize be the product of multiplying computedLayout’s destinationStride and sourceHeight. const planeSize = computedLayout.destinationStride * computedLayout.sourceHeight; // Let planeEnd be the sum of planeSize and computedLayout’s destinationOffset. const planeEnd = planeSize + computedLayout.destinationOffset; // If planeSize or planeEnd is greater than maximum range of unsigned long, return a TypeError. if (planeEnd > 4294967295) { throw new TypeError('Allocation size exceeds limit.'); } // Append planeEnd to endOffsets. endOffsets.push(planeEnd); // Assign the maximum of minAllocationSize and planeEnd to minAllocationSize. minAllocationSize = Math.max(minAllocationSize, planeEnd); // Check for overlap for (let earlierPlaneIndex = 0; earlierPlaneIndex < planeIndex; earlierPlaneIndex++) { const earlierLayout = computedLayouts[earlierPlaneIndex]!; // If plane A ends before plane B starts, they do not overlap. if ( endOffsets[planeIndex]! <= earlierLayout.destinationOffset || endOffsets[earlierPlaneIndex]! <= computedLayout.destinationOffset ) { continue; } throw new TypeError('Planes overlap.'); } computedLayouts.push(computedLayout); } // 12. Return combinedLayout. return { allocationSize: minAllocationSize, computedLayouts: computedLayouts, }; }; const AUDIO_SAMPLE_FORMATS = new Set( ['f32', 'f32-planar', 's16', 's16-planar', 's32', 's32-planar', 'u8', 'u8-planar'], ); /** * Abstract base class for custom audio sample resources. Implement this class to provide custom backing * for AudioSample instances. * @group Samples * @public */ export abstract class AudioSampleResource { /** @internal */ _referenceCount: number = 0; /** * Returns the audio sample format. * [See sample formats](https://developer.mozilla.org/en-US/docs/Web/API/AudioData/format) */ abstract getFormat(): AudioSampleFormat; /** Returns the audio sample rate in hertz. */ abstract getSampleRate(): number; /** Returns the number of audio frames in the sample, per channel. */ abstract getNumberOfFrames(): number; /** Returns the number of audio channels. */ abstract getNumberOfChannels(): number; /** Returns the presentation timestamp of the sample in seconds. */ abstract getTimestamp(): number; /** * Closes this resource, releasing held resources. Called automatically when the last {@link AudioSample} using this * resource is closed. */ abstract close(): void; /** * Returns the audio sample data for the plane given by `planeIndex`. The audio data must be in the format returned * by `getFormat()`. For interleaved formats, there is only one plane. */ abstract getDataPlane(planeIndex: number): Uint8Array; } /** * Metadata used for AudioSample initialization. * @group Samples * @public */ export type AudioSampleInit = { /** The audio data for this sample. */ data: AllowSharedBufferSource; /** * The audio sample format. [See sample formats](https://developer.mozilla.org/en-US/docs/Web/API/AudioData/format) */ format: AudioSampleFormat; /** The number of audio channels. */ numberOfChannels: number; /** The audio sample rate in hertz. */ sampleRate: number; /** The presentation timestamp of the sample in seconds. */ timestamp: number; }; /** * Options used for copying audio sample data. * @group Samples * @public */ export type AudioSampleCopyToOptions = { /** * The index identifying the plane to copy from. This must be 0 if using a non-planar (interleaved) output format. */ planeIndex: number; /** * The output format for the destination data. Defaults to the AudioSample's format. * [See sample formats](https://developer.mozilla.org/en-US/docs/Web/API/AudioData/format) */ format?: AudioSampleFormat; /** An offset into the source plane data indicating which frame to begin copying from. Defaults to 0. */ frameOffset?: number; /** * The number of frames to copy. If not provided, the copy will include all frames in the plane beginning * with frameOffset. */ frameCount?: number; }; /** * Represents a raw, unencoded audio sample. Mainly used as an expressive wrapper around WebCodecs API's * [`AudioData`](https://developer.mozilla.org/en-US/docs/Web/API/AudioData), but can also be used standalone. * @group Samples * @public */ export class AudioSample implements Disposable { /** @internal */ _data: AudioData | Uint8Array | AudioSampleResource; /** @internal */ _closed: boolean = false; /** * The audio sample format. * [See sample formats](https://developer.mozilla.org/en-US/docs/Web/API/AudioData/format) */ readonly format: AudioSampleFormat; /** The audio sample rate in hertz. */ readonly sampleRate: number; /** * The number of audio frames in the sample, per channel. In other words, the length of this audio sample in frames. */ readonly numberOfFrames: number; /** The number of audio channels. */ readonly numberOfChannels: number; /** The duration of the sample in seconds. */ readonly duration: number; /** * The presentation timestamp of the sample in seconds. May be negative. Samples with negative end timestamps should * not be presented. */ readonly timestamp: number; /** The presentation timestamp of the sample in microseconds. */ get microsecondTimestamp() { return Math.trunc(SECOND_TO_MICROSECOND_FACTOR * this.timestamp); } /** The duration of the sample in microseconds. */ get microsecondDuration() { return Math.trunc(SECOND_TO_MICROSECOND_FACTOR * this.duration); } /** * Creates a new {@link AudioSample}, either from an existing * [`AudioData`](https://developer.mozilla.org/en-US/docs/Web/API/AudioData) or from raw bytes specified in * {@link AudioSampleInit}. */ constructor(init: AudioData | AudioSampleInit | AudioSampleResource); constructor(init: AudioData | AudioSampleInit | AudioSampleResource) { if (isAudioData(init)) { if (init.format === null) { throw new TypeError('AudioData with null format is not supported.'); } this._data = init; this.format = init.format; this.sampleRate = init.sampleRate; this.numberOfFrames = init.numberOfFrames; this.numberOfChannels = init.numberOfChannels; this.timestamp = init.timestamp / 1e6; this.duration = init.numberOfFrames / init.sampleRate; } else if (init instanceof AudioSampleResource) { this._data = init; init._referenceCount++; this.format = init.getFormat(); if (!AUDIO_SAMPLE_FORMATS.has(this.format)) { throw new TypeError('getFormat() must return an AudioSampleFormat.'); } this.sampleRate = init.getSampleRate(); if (!Number.isInteger(this.sampleRate) || this.sampleRate <= 0) { throw new TypeError('getSampleRate() must return a positive integer.'); } this.numberOfFrames = init.getNumberOfFrames(); if (!Number.isInteger(this.numberOfFrames) || this.numberOfFrames < 0) { throw new TypeError('getNumberOfFrames() must return a non-negative integer.'); } this.numberOfChannels = init.getNumberOfChannels(); if (!Number.isInteger(this.numberOfChannels) || this.numberOfChannels <= 0) { throw new TypeError('getNumberOfChannels() must return a positive integer.'); } this.timestamp = init.getTimestamp(); if (!Number.isFinite(this.timestamp)) { throw new TypeError('getTimestamp() must return a finite number.'); } this.duration = this.numberOfFrames / this.sampleRate; } else { if (!init || typeof init !== 'object') { throw new TypeError('Invalid AudioDataInit: must be an object.'); } if (!AUDIO_SAMPLE_FORMATS.has(init.format)) { throw new TypeError('Invalid AudioDataInit: invalid format.'); } if (!Number.isFinite(init.sampleRate) || init.sampleRate <= 0) { throw new TypeError('Invalid AudioDataInit: sampleRate must be > 0.'); } if (!Number.isInteger(init.numberOfChannels) || init.numberOfChannels === 0) { throw new TypeError('Invalid AudioDataInit: numberOfChannels must be an integer > 0.'); } if (!Number.isFinite(init?.timestamp)) { throw new TypeError('init.timestamp must be a number.'); } const numberOfFrames = init.data.byteLength / (getBytesPerSample(init.format) * init.numberOfChannels); if (!Number.isInteger(numberOfFrames)) { throw new TypeError('Invalid AudioDataInit: data size is not a multiple of frame size.'); } this.format = init.format; this.sampleRate = init.sampleRate; this.numberOfFrames = numberOfFrames; this.numberOfChannels = init.numberOfChannels; this.timestamp = init.timestamp; this.duration = numberOfFrames / init.sampleRate; let dataBuffer: Uint8Array; if (init.data instanceof ArrayBuffer) { dataBuffer = new Uint8Array(init.data); } else if (ArrayBuffer.isView(init.data)) { dataBuffer = new Uint8Array(init.data.buffer, init.data.byteOffset, init.data.byteLength); } else { throw new TypeError('Invalid AudioDataInit: data is not a BufferSource.'); } const expectedSize = this.numberOfFrames * this.numberOfChannels * getBytesPerSample(this.format); if (dataBuffer.byteLength < expectedSize) { throw new TypeError('Invalid AudioDataInit: insufficient data size.'); } this._data = dataBuffer; } finalizationRegistry?.register(this, { type: 'audio', data: this._data }, this); } /** Returns the number of bytes required to hold the audio sample's data as specified by the given options. */ allocationSize(options: AudioSampleCopyToOptions) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (!Number.isInteger(options.planeIndex) || options.planeIndex < 0) { throw new TypeError('planeIndex must be a non-negative integer.'); } if (options.format !== undefined && !AUDIO_SAMPLE_FORMATS.has(options.format)) { throw new TypeError('Invalid format.'); } if (options.frameOffset !== undefined && (!Number.isInteger(options.frameOffset) || options.frameOffset < 0)) { throw new TypeError('frameOffset must be a non-negative integer.'); } if (options.frameCount !== undefined && (!Number.isInteger(options.frameCount) || options.frameCount < 0)) { throw new TypeError('frameCount must be a non-negative integer.'); } if (this._closed) { throw new Error('AudioSample is closed.'); } const destFormat = options.format ?? this.format; const frameOffset = options.frameOffset ?? 0; if (frameOffset >= this.numberOfFrames) { throw new RangeError('frameOffset out of range'); } const copyFrameCount = options.frameCount !== undefined ? options.frameCount : (this.numberOfFrames - frameOffset); if (copyFrameCount > (this.numberOfFrames - frameOffset)) { throw new RangeError('frameCount out of range'); } const bytesPerSample = getBytesPerSample(destFormat); const isPlanar = formatIsPlanar(destFormat); if (isPlanar && options.planeIndex >= this.numberOfChannels) { throw new RangeError('planeIndex out of range'); } if (!isPlanar && options.planeIndex !== 0) { throw new RangeError('planeIndex out of range'); } const elementCount = isPlanar ? copyFrameCount : copyFrameCount * this.numberOfChannels; return elementCount * bytesPerSample; } /** Copies the audio sample's data to an ArrayBuffer or ArrayBufferView as specified by the given options. */ copyTo(destination: AllowSharedBufferSource, options: AudioSampleCopyToOptions) { if (!isAllowSharedBufferSource(destination)) { throw new TypeError('destination must be an ArrayBuffer or an ArrayBuffer view.'); } if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (!Number.isInteger(options.planeIndex) || options.planeIndex < 0) { throw new TypeError('planeIndex must be a non-negative integer.'); } if (options.format !== undefined && !AUDIO_SAMPLE_FORMATS.has(options.format)) { throw new TypeError('Invalid format.'); } if (options.frameOffset !== undefined && (!Number.isInteger(options.frameOffset) || options.frameOffset < 0)) { throw new TypeError('frameOffset must be a non-negative integer.'); } if (options.frameCount !== undefined && (!Number.isInteger(options.frameCount) || options.frameCount < 0)) { throw new TypeError('frameCount must be a non-negative integer.'); } if (this._closed) { throw new Error('AudioSample is closed.'); } const { format, frameCount: optFrameCount, frameOffset: optFrameOffset } = options; let { planeIndex } = options; const srcFormat = this.format; const destFormat = format ?? this.format; if (!destFormat) throw new Error('Destination format not determined'); const numFrames = this.numberOfFrames; const numChannels = this.numberOfChannels; const frameOffset = optFrameOffset ?? 0; if (frameOffset >= numFrames) { throw new RangeError('frameOffset out of range'); } const copyFrameCount = optFrameCount !== undefined ? optFrameCount : (numFrames - frameOffset); if (copyFrameCount > (numFrames - frameOffset)) { throw new RangeError('frameCount out of range'); } const destBytesPerSample = getBytesPerSample(destFormat); const destIsPlanar = formatIsPlanar(destFormat); if (destIsPlanar && planeIndex >= numChannels) { throw new RangeError('planeIndex out of range'); } if (!destIsPlanar && planeIndex !== 0) { throw new RangeError('planeIndex out of range'); } const destElementCount = destIsPlanar ? copyFrameCount : copyFrameCount * numChannels; const requiredSize = destElementCount * destBytesPerSample; if (destination.byteLength < requiredSize) { throw new RangeError('Destination buffer is too small'); } const destView = toDataView(destination); const writeFn = getWriteFunction(destFormat); if (isAudioData(this._data)) { if (isWebKit() && numChannels > 2 && destFormat !== srcFormat) { // WebKit bug workaround doAudioDataCopyToWebKitWorkaround( this._data, destView, srcFormat, destFormat, numChannels, planeIndex, frameOffset, copyFrameCount, ); } else { // Per spec, only f32-planar conversion must be supported, but in practice, all browsers support all // destination formats, so let's just delegate here: this._data.copyTo(destination, { planeIndex, frameOffset, frameCount: copyFrameCount, format: destFormat, }); } } else { const readFn = getReadFunction(srcFormat); const srcBytesPerSample = getBytesPerSample(srcFormat); const srcIsPlanar = formatIsPlanar(srcFormat); let uint8Data: Uint8Array; if (this._data instanceof AudioSampleResource) { const getDataPlaneValidated = (index: number) => { const result = (this._data as AudioSampleResource).getDataPlane(index); if (!(result instanceof Uint8Array)) { throw new TypeError('getDataPlane() must return a Uint8Array.'); } const expectedSize = numFrames * srcBytesPerSample * (srcIsPlanar ? 1 : numChannels); if (result.byteLength !== expectedSize) { throw new TypeError( `Data plane ${index} has invalid size. Expected exactly ${expectedSize} bytes, got` + ` ${result.byteLength} bytes.`, ); } return result; }; if (srcIsPlanar) { if (destIsPlanar) { // Only one source plane will be extracted, so let's fetch only that one uint8Data = getDataPlaneValidated(planeIndex); planeIndex = 0; // To fix the subsequent access } else { // Pack all planes tightly together uint8Data = new Uint8Array(numFrames * srcBytesPerSample * numChannels); for (let ch = 0; ch < numChannels; ch++) { const planeData = getDataPlaneValidated(ch); uint8Data.set(planeData, ch * numFrames * srcBytesPerSample); } } } else { uint8Data = getDataPlaneValidated(0); // That's the only plane there is } } else { uint8Data = this._data; } const srcView = toDataView(uint8Data); for (let i = 0; i < copyFrameCount; i++) { if (destIsPlanar) { const destOffset = i * destBytesPerSample; let srcOffset: number; if (srcIsPlanar) { srcOffset = (planeIndex * numFrames + (i + frameOffset)) * srcBytesPerSample; } else { srcOffset = (((i + frameOffset) * numChannels) + planeIndex) * srcBytesPerSample; } const normalized = readFn(srcView, srcOffset); writeFn(destView, destOffset, normalized); } else { for (let ch = 0; ch < numChannels; ch++) { const destIndex = i * numChannels + ch; const destOffset = destIndex * destBytesPerSample; let srcOffset: number; if (srcIsPlanar) { srcOffset = (ch * numFrames + (i + frameOffset)) * srcBytesPerSample; } else { srcOffset = (((i + frameOffset) * numChannels) + ch) * srcBytesPerSample; } const normalized = readFn(srcView, srcOffset); writeFn(destView, destOffset, normalized); } } } } } /** Clones this audio sample. */ clone(): AudioSample { if (this._closed) { throw new Error('AudioSample is closed.'); } if (this._data instanceof AudioSampleResource) { const sample = new AudioSample(this._data); sample.setTimestamp(this.timestamp); // Make sure the timestamp is correct return sample; } else if (isAudioData(this._data)) { const sample = new AudioSample(this._data.clone()); sample.setTimestamp(this.timestamp); return sample; } else { return new AudioSample({ format: this.format, sampleRate: this.sampleRate, numberOfFrames: this.numberOfFrames, numberOfChannels: this.numberOfChannels, timestamp: this.timestamp, data: this._data, }); } } /** * Returns a new {@link AudioSample} containing only the frames in the range [startSample, endSample). Both bounds * must lie within this sample's range of frames. The returned sample's timestamp is shifted to match the start of * the trimmed section. */ trim(startSample: number, endSample = this.numberOfFrames) { if (!Number.isInteger(startSample) || startSample < 0) { throw new TypeError('startSample must be a non-negative integer.'); } if (!Number.isInteger(endSample) || endSample < 0) { throw new TypeError('endSample must be a non-negative integer.'); } if (startSample > this.numberOfFrames) { throw new RangeError('startSample out of range.'); } if (endSample > this.numberOfFrames) { throw new RangeError('endSample out of range.'); } if (endSample < startSample) { throw new RangeError('endSample must not be less than startSample.'); } if (this._closed) { throw new Error('AudioSample is closed.'); } const frameCount = endSample - startSample; const bytesPerSample = getBytesPerSample(this.format); let data: Uint8Array; if (formatIsPlanar(this.format)) { const planeSize = frameCount * bytesPerSample; data = new Uint8Array(planeSize * this.numberOfChannels); if (frameCount > 0) { // Copy plane-by-plane for (let i = 0; i < this.numberOfChannels; i++) { this.copyTo(data.subarray(i * planeSize, (i + 1) * planeSize), { planeIndex: i, format: this.format, frameOffset: startSample, frameCount, }); } } } else { // Trivial data = new Uint8Array(frameCount * this.numberOfChannels * bytesPerSample); if (frameCount > 0) { this.copyTo(data, { planeIndex: 0, format: this.format, frameOffset: startSample, frameCount, }); } } return new AudioSample({ data, format: this.format, sampleRate: this.sampleRate, numberOfChannels: this.numberOfChannels, timestamp: this.timestamp + startSample / this.sampleRate, }); } /** * Closes this audio sample, releasing held resources. Audio samples should be closed as soon as they are not * needed anymore. */ close(): void { if (this._closed) { return; } finalizationRegistry?.unregister(this); if (this._data instanceof AudioSampleResource) { this._data._referenceCount--; if (this._data._referenceCount === 0) { this._data.close(); } } else if (isAudioData(this._data)) { this._data.close(); } else { this._data = new Uint8Array(0); } this._closed = true; } /** * Converts this audio sample to an AudioData for use with the WebCodecs API. The AudioData returned by this * method *must* be closed separately from this audio sample. */ toAudioData() { if (this._closed) { throw new Error('AudioSample is closed.'); } if (this._data instanceof AudioSampleResource) { return this._createAudioDataFromData(); } else if (isAudioData(this._data)) { if (this._data.timestamp === this.microsecondTimestamp) { // Timestamp matches, let's just return the data (but cloned) return this._data.clone(); } else { // It's impossible to simply change an AudioData's timestamp, so we'll need to create a new one return this._createAudioDataFromData(); } } else { return new AudioData({ format: this.format, sampleRate: this.sampleRate, numberOfFrames: this.numberOfFrames, numberOfChannels: this.numberOfChannels, timestamp: this.microsecondTimestamp, data: this._data.buffer instanceof ArrayBuffer ? this._data.buffer : this._data.slice(), // In the case of SharedArrayBuffer, convert to ArrayBuffer }); } } /** @internal */ _createAudioDataFromData() { if (formatIsPlanar(this.format)) { const size = this.allocationSize({ planeIndex: 0, format: this.format }); const data = new ArrayBuffer(size * this.numberOfChannels); // We gotta read out each plane individually for (let i = 0; i < this.numberOfChannels; i++) { this.copyTo(new Uint8Array(data, i * size, size), { planeIndex: i, format: this.format }); } return new AudioData({ format: this.format, sampleRate: this.sampleRate, numberOfFrames: this.numberOfFrames, numberOfChannels: this.numberOfChannels, timestamp: this.microsecondTimestamp, data, }); } else { const data = new ArrayBuffer(this.allocationSize({ planeIndex: 0, format: this.format })); this.copyTo(data, { planeIndex: 0, format: this.format }); return new AudioData({ format: this.format, sampleRate: this.sampleRate, numberOfFrames: this.numberOfFrames, numberOfChannels: this.numberOfChannels, timestamp: this.microsecondTimestamp, data, }); } } /** Convert this audio sample to an AudioBuffer for use with the Web Audio API. */ toAudioBuffer() { if (this._closed) { throw new Error('AudioSample is closed.'); } const audioBuffer = new AudioBuffer({ numberOfChannels: this.numberOfChannels, length: this.numberOfFrames, sampleRate: this.sampleRate, }); const dataBytes = new Float32Array(this.allocationSize({ planeIndex: 0, format: 'f32-planar' }) / 4); for (let i = 0; i < this.numberOfChannels; i++) { this.copyTo(dataBytes, { planeIndex: i, format: 'f32-planar' }); audioBuffer.copyToChannel(dataBytes, i); } return audioBuffer; } /** Sets the presentation timestamp of this audio sample, in seconds. */ setTimestamp(newTimestamp: number) { if (!Number.isFinite(newTimestamp)) { throw new TypeError('newTimestamp must be a number.'); } // eslint-disable-next-line @typescript-eslint/no-unnecessary-type-assertion (this.timestamp as number) = newTimestamp; } /** Calls `.close()`. */ [Symbol.dispose]() { this.close(); } /** @internal */ static* _fromAudioBuffer(audioBuffer: AudioBuffer, timestamp: number) { if (!(audioBuffer instanceof AudioBuffer)) { throw new TypeError('audioBuffer must be an AudioBuffer.'); } const MAX_FLOAT_COUNT = 48000 * 5; // 5 seconds of mono 48 kHz audio per sample const numberOfChannels = audioBuffer.numberOfChannels; const sampleRate = audioBuffer.sampleRate; const totalFrames = audioBuffer.length; const maxFramesPerChunk = Math.floor(MAX_FLOAT_COUNT / numberOfChannels); let currentRelativeFrame = 0; let remainingFrames = totalFrames; // Create AudioSamples in a chunked fashion so we don't create huge Float32Arrays while (remainingFrames > 0) { const framesToCopy = Math.min(maxFramesPerChunk, remainingFrames); const chunkData = new Float32Array(numberOfChannels * framesToCopy); for (let channel = 0; channel < numberOfChannels; channel++) { audioBuffer.copyFromChannel( chunkData.subarray(channel * framesToCopy, (channel + 1) * framesToCopy), channel, currentRelativeFrame, ); } yield new AudioSample({ format: 'f32-planar', sampleRate, numberOfFrames: framesToCopy, numberOfChannels, timestamp: timestamp + currentRelativeFrame / sampleRate, data: chunkData, }); currentRelativeFrame += framesToCopy; remainingFrames -= framesToCopy; } } /** * Creates AudioSamples from an AudioBuffer, starting at the given timestamp in seconds. Typically creates exactly * one sample, but may create multiple if the AudioBuffer is exceedingly large. */ static fromAudioBuffer(audioBuffer: AudioBuffer, timestamp: number) { if (!(audioBuffer instanceof AudioBuffer)) { throw new TypeError('audioBuffer must be an AudioBuffer.'); } const MAX_FLOAT_COUNT = 48000 * 5; // 5 seconds of mono 48 kHz audio per sample const numberOfChannels = audioBuffer.numberOfChannels; const sampleRate = audioBuffer.sampleRate; const totalFrames = audioBuffer.length; const maxFramesPerChunk = Math.floor(MAX_FLOAT_COUNT / numberOfChannels); let currentRelativeFrame = 0; let remainingFrames = totalFrames; const result: AudioSample[] = []; // Create AudioSamples in a chunked fashion so we don't create huge Float32Arrays while (remainingFrames > 0) { const framesToCopy = Math.min(maxFramesPerChunk, remainingFrames); const chunkData = new Float32Array(numberOfChannels * framesToCopy); for (let channel = 0; channel < numberOfChannels; channel++) { audioBuffer.copyFromChannel( chunkData.subarray(channel * framesToCopy, (channel + 1) * framesToCopy), channel, currentRelativeFrame, ); } const audioSample = new AudioSample({ format: 'f32-planar', sampleRate, numberOfFrames: framesToCopy, numberOfChannels, timestamp: timestamp + currentRelativeFrame / sampleRate, data: chunkData, }); result.push(audioSample); currentRelativeFrame += framesToCopy; remainingFrames -= framesToCopy; } return result; } } const getBytesPerSample = (format: AudioSampleFormat): number => { switch (format) { case 'u8': case 'u8-planar': return 1; case 's16': case 's16-planar': return 2; case 's32': case 's32-planar': return 4; case 'f32': case 'f32-planar': return 4; default: throw new Error('Unknown AudioSampleFormat'); } }; const formatIsPlanar = (format: AudioSampleFormat): boolean => { switch (format) { case 'u8-planar': case 's16-planar': case 's32-planar': case 'f32-planar': return true; default: return false; } }; const getReadFunction = (format: AudioSampleFormat): (view: DataView, offset: number) => number => { switch (format) { case 'u8': case 'u8-planar': return (view, offset) => (view.getUint8(offset) - 128) / 128; case 's16': case 's16-planar': return (view, offset) => view.getInt16(offset, true) / 32768; case 's32': case 's32-planar': return (view, offset) => view.getInt32(offset, true) / 2147483648; case 'f32': case 'f32-planar': return (view, offset) => view.getFloat32(offset, true); } }; const getWriteFunction = (format: AudioSampleFormat): (view: DataView, offset: number, value: number) => void => { switch (format) { case 'u8': case 'u8-planar': return (view, offset, value) => view.setUint8(offset, clamp((value + 1) * 127.5, 0, 255)); case 's16': case 's16-planar': return (view, offset, value) => view.setInt16(offset, clamp(Math.round(value * 32767), -32768, 32767), true); case 's32': case 's32-planar': return (view, offset, value) => view.setInt32(offset, clamp(Math.round(value * 2147483647), -2147483648, 2147483647), true); case 'f32': case 'f32-planar': return (view, offset, value) => view.setFloat32(offset, value, true); } }; const isAudioData = (x: unknown): x is AudioData => { return typeof AudioData !== 'undefined' && x instanceof AudioData; }; export const toInterleavedAudioFormat = (format: AudioSampleFormat): 'u8' | 's16' | 's32' | 'f32' => { switch (format) { case 'u8-planar': return 'u8'; case 's16-planar': return 's16'; case 's32-planar': return 's32'; case 'f32-planar': return 'f32'; default: return format; } }; /** * WebKit has a bug where calling AudioData.copyTo with a format different from the source format * crashes the tab when there are more than 2 channels. This function works around that by always * copying with the source format and then manually converting to the destination format. * * See https://bugs.webkit.org/show_bug.cgi?id=302521. */ const doAudioDataCopyToWebKitWorkaround = ( audioData: AudioData, destView: DataView, srcFormat: AudioSampleFormat, destFormat: AudioSampleFormat, numChannels: number, planeIndex: number, frameOffset: number, copyFrameCount: number, ) => { const readFn = getReadFunction(srcFormat); const writeFn = getWriteFunction(destFormat); const srcBytesPerSample = getBytesPerSample(srcFormat); const destBytesPerSample = getBytesPerSample(destFormat); const srcIsPlanar = formatIsPlanar(srcFormat); const destIsPlanar = formatIsPlanar(destFormat); if (destIsPlanar) { if (srcIsPlanar) { // src planar -> dest planar: copy single plane and convert const data = new ArrayBuffer(copyFrameCount * srcBytesPerSample); const dataView = toDataView(data); audioData.copyTo(data, { planeIndex, frameOffset, frameCount: copyFrameCount, format: srcFormat, }); for (let i = 0; i < copyFrameCount; i++) { const srcOffset = i * srcBytesPerSample; const destOffset = i * destBytesPerSample; const sample = readFn(dataView, srcOffset); writeFn(destView, destOffset, sample); } } else { // src interleaved -> dest planar: copy all interleaved data, extract one channel const data = new ArrayBuffer(copyFrameCount * numChannels * srcBytesPerSample); const dataView = toDataView(data); audioData.copyTo(data, { planeIndex: 0, frameOffset, frameCount: copyFrameCount, format: srcFormat, }); for (let i = 0; i < copyFrameCount; i++) { const srcOffset = (i * numChannels + planeIndex) * srcBytesPerSample; const destOffset = i * destBytesPerSample; const sample = readFn(dataView, srcOffset); writeFn(destView, destOffset, sample); } } } else { if (srcIsPlanar) { // src planar -> dest interleaved: copy each plane and interleave const planeSize = copyFrameCount * srcBytesPerSample; const data = new ArrayBuffer(planeSize); const dataView = toDataView(data); for (let ch = 0; ch < numChannels; ch++) { audioData.copyTo(data, { planeIndex: ch, frameOffset, frameCount: copyFrameCount, format: srcFormat, }); for (let i = 0; i < copyFrameCount; i++) { const srcOffset = i * srcBytesPerSample; const destOffset = (i * numChannels + ch) * destBytesPerSample; const sample = readFn(dataView, srcOffset); writeFn(destView, destOffset, sample); } } } else { // src interleaved -> dest interleaved: copy all and convert const data = new ArrayBuffer(copyFrameCount * numChannels * srcBytesPerSample); const dataView = toDataView(data); audioData.copyTo(data, { planeIndex: 0, frameOffset, frameCount: copyFrameCount, format: srcFormat, }); for (let i = 0; i < copyFrameCount; i++) { for (let ch = 0; ch < numChannels; ch++) { const idx = i * numChannels + ch; const srcOffset = idx * srcBytesPerSample; const destOffset = idx * destBytesPerSample; const sample = readFn(dataView, srcOffset); writeFn(destView, destOffset, sample); } } } } }; export const audioSampleToInterleavedFormat = (sample: AudioSample, format: 'u8' | 's16' | 's32' | 'f32') => { const size = sample.allocationSize({ format, planeIndex: 0 }); const buffer = new ArrayBuffer(size); sample.copyTo(buffer, { format, planeIndex: 0 }); return new AudioSample({ data: buffer, format, numberOfChannels: sample.numberOfChannels, sampleRate: sample.sampleRate, timestamp: sample.timestamp, duration: sample.duration, }); }; ===== src/input.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Demuxer, DurationMetadataRequestOptions } from './demuxer'; import { InputFormat, InputFormatOptions, validateInputFormatOptions } from './input-format'; import { InputAudioTrack, InputAudioTrackBacking, InputTrack, InputTrackBacking, InputVideoTrack, InputVideoTrackBacking, InputTrackQuery, mergeInputTrackQueries, queryInputTracks, toValidatedInputTrackQuery, prefer, desc, } from './input-track'; import { PacketRetrievalOptions } from './media-sink'; import { arrayArgmin, arrayCount, assert, EventEmitter, polyfillSymbolDispose, removeItem, } from './misc'; import { Reader } from './reader'; import { PathedSource, Source, SourceRef, SourceRequest, sourceRequestsAreEqual, } from './source'; polyfillSymbolDispose(); export const DEFAULT_SOURCE_CACHE_GROUP = 1; export const ENCRYPTION_KEY_CACHE_GROUP = 2; /** * The options for creating an Input object. * @group Input files & tracks * @public */ export type InputOptions = { /** A list of supported formats. If the source file is not of one of these formats, then it cannot be read. */ formats: InputFormat[]; /** The source from which data will be read. */ source: S | SourceRef; /** * An optional, second {@link Input} instance that contains the necessary metadata to initialize the tracks of * this input. This is necessary in cases where track initialization info and media data are carried in separate * files, like is the case with segmented MP4 (CMAF) files. * * The use of this field depends on the input format. */ initInput?: Input; /** Can be used to specify additional per-format configuration. */ formatOptions?: InputFormatOptions; }; type SourceCacheEntry = { request: SourceRequest; sourceRef: SourceRef; age: number; cacheGroup: number; }; /** * Describes the events that an {@link Input} emits, with each key being an event name and its value being the * event data. * * @group Input files & tracks * @public */ export type InputEvents = { /** Emitted whenever a {@link Source} is loaded by the input. Useful to track reads. */ source: { /** The loaded source. */ source: Source; /** The request that led to loading this source, or `null` if the input is not pathed. */ request: SourceRequest | null; /** Whether the source is the root file of the media. */ isRoot: boolean; }; }; /** * Represents input media, backed by a single file or multiple files depending on the format. * * This is the root object from which all media read operations start. * @group Input files & tracks * @public */ export class Input extends EventEmitter implements Disposable { /** @internal */ _rootRef: SourceRef; /** @internal */ _formats: InputFormat[]; /** @internal */ _initInput: Input | null; /** @internal */ _demuxerPromise: Promise | null = null; /** @internal */ _format: InputFormat | null = null; /** @internal */ _reader!: Reader; /** @internal */ _trackBackingsCache: InputTrackBacking[] | null = null; /** @internal */ _backingToTrack = new Map(); /** @internal */ _disposed = false; /** @internal */ _nextSourceCacheAge = 0; /** @internal */ _sourceRefs: SourceRef[] = []; /** @internal */ _sourceCache: SourceCacheEntry[] = []; /** @internal */ _sourceCachePromises: { request: SourceRequest; cacheGroup: number; promise: Promise; }[] = []; /** @internal */ _formatOptions: InputFormatOptions; /** @internal */ _onFormatDetermined: ((format: InputFormat) => void) | null = null; /** True if the input has been disposed. */ get disposed() { return this._disposed; } /** * Creates a new input file from the specified options. No reading operations will be performed until methods are * called on this instance. */ constructor(options: InputOptions) { super(); if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (!Array.isArray(options.formats) || options.formats.some(x => !(x instanceof InputFormat))) { throw new TypeError('options.formats must be an array of InputFormat.'); } if (!(options.source instanceof Source || options.source instanceof SourceRef)) { throw new TypeError('options.source must be a Source or SourceRef.'); } if (options.source instanceof Source && options.source._disposed) { throw new TypeError('options.source must not be a disposed Source.'); } if (options.initInput !== undefined && !(options.initInput instanceof Input)) { throw new TypeError('options.initInput, when provided, must be an Input.'); } if (options.formatOptions !== undefined) { validateInputFormatOptions(options.formatOptions, 'formatOptions'); } this._formats = options.formats; this._initInput = options.initInput ?? null; this._formatOptions = options.formatOptions ?? {}; if (options.source instanceof Source) { this._rootRef = options.source.ref(); } else { this._rootRef = options.source; } this._sourceRefs.push(this._rootRef); } /** @internal */ get _rootSource() { return this._rootRef.source; } /** @internal */ async _getSourceUncached(request: SourceRequest) { assert(this._rootSource instanceof PathedSource); const ref = await this._rootSource._resolveRequest(request); this._emit('source', { source: ref.source, request, isRoot: request.isRoot }); return ref; } /** @internal */ _getSourceCached(request: SourceRequest, cacheGroup = DEFAULT_SOURCE_CACHE_GROUP): Promise { const cachedEntry = this._sourceCache.find(x => x.cacheGroup === cacheGroup && sourceRequestsAreEqual(x.request, request), ); if (cachedEntry) { cachedEntry.age++; return Promise.resolve(cachedEntry.sourceRef.source.ref()); } const cachedPromiseEntry = this._sourceCachePromises.find(x => x.cacheGroup === cacheGroup && sourceRequestsAreEqual(x.request, request), ); if (cachedPromiseEntry) { return cachedPromiseEntry.promise.then(x => x.sourceRef.source.ref()); } const promise = (async () => { const sourceRef = await this._getSourceUncached(request); const MAX_SOURCE_CACHE_SIZE = 4; const count = arrayCount( this._sourceCache, x => x.cacheGroup === cacheGroup && x.sourceRef.source._refCount === 1, ); if (count >= MAX_SOURCE_CACHE_SIZE) { const minAgeIndex = arrayArgmin( this._sourceCache, x => x.cacheGroup === cacheGroup && x.sourceRef.source._refCount === 1 ? x.age : Infinity, ); assert(minAgeIndex !== -1); const entry = this._sourceCache[minAgeIndex]!; this._sourceCache.splice(minAgeIndex, 1); entry.sourceRef.free(); removeItem(this._sourceRefs, entry.sourceRef); } this._sourceRefs.push(sourceRef); const promiseIndex = this._sourceCachePromises.findIndex(x => x.request === request); assert(promiseIndex !== -1); this._sourceCachePromises.splice(promiseIndex, 1); const cacheEntry: SourceCacheEntry = { request, sourceRef, age: this._nextSourceCacheAge++, cacheGroup, }; return cacheEntry; })(); this._sourceCachePromises.push({ request, cacheGroup, promise, }); return promise.then((entry) => { const ref = entry.sourceRef.source.ref(); // We need to add it to the cache this late to avoid the ref being freed prematurely due to race conditions this._sourceCache.push(entry); return ref; }); } /** @internal */ _getDemuxer() { return this._demuxerPromise ??= (async () => { this._reader = new Reader(this._rootSource); this._emit('source', { source: this._rootSource, request: null, isRoot: true }); for (const format of this._formats) { const canRead = await format._canReadInput(this); if (canRead) { this._format = format; this._onFormatDetermined?.(format); return format._createDemuxer(this); } } throw new UnsupportedInputFormatError(); })(); } /** * Returns the source from which this input file reads data for the root path. */ get source(): S { return this._rootSource; } /** * Returns the format of the input file. You can compare this result directly to the {@link InputFormat} singletons * or use `instanceof` checks for subset-aware logic (for example, `format instanceof MatroskaInputFormat` is true * for both MKV and WebM). */ async getFormat() { await this._getDemuxer(); assert(this._format!); return this._format; } /** Returns `true` if the format of the input file is known and the file can be read, `false` otherwise. */ async canRead(): Promise { try { await this._getDemuxer(); return true; } catch (error) { if (error instanceof UnsupportedInputFormatError) { return false; } throw error; } } /** * Returns the timestamp at which the input file starts. More precisely, returns the smallest starting timestamp * among all tracks. * * Optionally, you can pass in the list of tracks for which you want to compute the starting timestamp. * * Note that this method is potentially expensive for inputs with many tracks (such as HLS manifests), since it * probes every track. */ async getFirstTimestamp(tracks?: InputTrack[]) { tracks ??= await this.getTracks(); const filtered = tracks.filter(x => x !== null); if (filtered.length === 0) { return 0; } const firstTimestamps = await Promise.all(filtered.map(x => x.getFirstTimestamp())); return Math.min(...firstTimestamps); } /** * Computes the duration of the input file, in seconds. More precisely, returns the largest end timestamp among * all tracks. * * Optionally, you can pass in the list of tracks for which you want to compute the duration. * * This method can be potentially expensive depending on the underlying file format, because it returns the most * accurate duration possible and must check all tracks. Use {@link Input.getDurationFromMetadata} for a faster but * less accurate estimate of duration. * * By default, when any track in the underlying media is live, this method will only resolve once the live stream * ends. If you want to query the current duration of the media, set {@link PacketRetrievalOptions.skipLiveWait} * to `true` in the options. */ async computeDuration(tracks?: InputTrack[], options?: PacketRetrievalOptions) { tracks ??= await this.getTracks(); const filtered = tracks.filter(x => x !== null); if (filtered.length === 0) { return 0; } const tracksDurations = await Promise.all(filtered.map(x => x.computeDuration(options))); return Math.max(...tracksDurations); } /** * Gets the duration (end timestamp) in seconds of the input file from metadata stored in the file. This value may * be approximate or diverge from the actual, precise duration returned by `.computeDuration()`, but compared to * that method, this method is cheaper. When the duration cannot be determined from the file metadata, `null` * is returned. * * Optionally, you can pass in the list of tracks for which you want to get the duration from metadata. * * By default, when the underlying media is live, this method will only resolve once the live stream * ends. If you want to query the current duration of the media, set * {@link DurationMetadataRequestOptions.skipLiveWait} to `true` in the options. */ async getDurationFromMetadata(tracks?: InputTrack[], options?: DurationMetadataRequestOptions) { tracks ??= await this.getTracks(); const filtered = tracks.filter(x => x !== null); const tracksDurations = await Promise.all(filtered.map(x => x.getDurationFromMetadata(options))); const nonNullDurations = tracksDurations.filter(x => x !== null); if (nonNullDurations.length === 0) { return null; } return Math.max(...nonNullDurations); } /** * Returns the list of all tracks of this input file in the order in which they appear in the file. An optional * query can be provided. */ async getTracks(query?: InputTrackQuery): Promise { query &&= toValidatedInputTrackQuery(query); const backings = await this._getTrackBackings(); const tracks = backings.map(backing => this._wrapBackingAsTrack(backing)); return queryInputTracks(tracks, query); } /** Returns the list of all video tracks of this input file. An optional query can be provided. */ async getVideoTracks(query?: InputTrackQuery): Promise { query &&= toValidatedInputTrackQuery(query); const tracks = await this.getTracks(); const videoTracks = tracks.filter((x): x is InputVideoTrack => x.isVideoTrack()); return queryInputTracks(videoTracks, query); } /** Returns the list of all audio tracks of this input file. An optional query can be provided. */ async getAudioTracks(query?: InputTrackQuery): Promise { query &&= toValidatedInputTrackQuery(query); const tracks = await this.getTracks(); const audioTracks = tracks.filter((x): x is InputAudioTrack => x.isAudioTrack()); return queryInputTracks(audioTracks, query); } /** * Returns the primary video track of this input file, or null if there are no video tracks. * * Multiple factors determine which track is considered primary, including its position in the file, disposition, * bitrate (higher bitrate is preferred), and if it can be paired with an audio track. */ async getPrimaryVideoTrack( query?: InputTrackQuery, ): Promise { query &&= toValidatedInputTrackQuery(query); const merged = mergeInputTrackQueries(query, { sortBy: async t => [ prefer((await t.getDisposition()).default), prefer(await t.hasPairableAudioTrack()), prefer(!(await t.hasOnlyKeyPackets())), desc(await t.getBitrate()), ], }); const sorted = await this.getVideoTracks(merged); return sorted[0] ?? null; } /** * Returns the primary audio track of this input file, or null if there are no audio tracks. * * Multiple factors determine which track is considered primary, including its position in the file, disposition, * bitrate (higher bitrate is preferred), and if it can be paired with the primary video track. */ async getPrimaryAudioTrack( query?: InputTrackQuery, ): Promise { query &&= toValidatedInputTrackQuery(query); const primaryVideoTrack = await this.getPrimaryVideoTrack(); const merged = mergeInputTrackQueries(query, { sortBy: async t => [ prefer(!primaryVideoTrack || t.canBePairedWith(primaryVideoTrack)), prefer((await t.getDisposition()).default), desc(await t.getBitrate()), ], }); const sorted = await this.getAudioTracks(merged); return sorted[0] ?? null; } /** @internal */ async _getTrackBackings() { const demuxer = await this._getDemuxer(); return this._trackBackingsCache ??= await demuxer.getTrackBackings(); } /** @internal */ _wrapBackingAsTrack(backing: InputTrackBacking): InputTrack { const existing = this._backingToTrack.get(backing); if (existing) { return existing; } const type = backing.getType(); const track = type === 'video' ? new InputVideoTrack(this, backing as InputVideoTrackBacking) : new InputAudioTrack(this, backing as InputAudioTrackBacking); this._backingToTrack.set(backing, track); return track; } /** Returns the full MIME type of this input file, including track codecs. */ async getMimeType() { const demuxer = await this._getDemuxer(); return demuxer.getMimeType(); } /** * Returns descriptive metadata tags about the media file, such as title, author, date, cover art, or other * attached files. */ async getMetadataTags() { const demuxer = await this._getDemuxer(); return demuxer.getMetadataTags(); } /** * Disposes this input and frees connected resources. When an input is disposed, ongoing read operations will be * canceled, all future read operations will fail, any open decoders will be closed, and all ongoing media sink * operations will be canceled. Disallowed and canceled operations will throw an {@link InputDisposedError}. * * You are expected not to use an input after disposing it. While some operations may still work, it is not * specified and may change in any future update. */ dispose() { if (this._disposed) { return; } this._disposed = true; for (const ref of this._sourceRefs) { ref.free(); } this._sourceRefs.length = 0; if (this._demuxerPromise) { // The demuxer promise may already be rejected after failed format detection. void this._demuxerPromise .then(demuxer => demuxer.dispose()) .catch(() => {}); } } /** * Calls `.dispose()` on the input, implementing the `Disposable` interface for use with * JavaScript Explicit Resource Management features. */ [Symbol.dispose]() { this.dispose(); } } /** * Thrown when trying to operate on an input that has an unsupported or unrecognizable format. * @group Input files & tracks * @public */ export class UnsupportedInputFormatError extends Error { /** Creates a new {@link UnsupportedInputFormatError}. */ constructor(message = 'Input has an unsupported or unrecognizable format.') { super(message); this.name = 'UnsupportedInputFormatError'; } } /** * Thrown when an operation was prevented because the corresponding {@link Input} has been disposed. * @group Input files & tracks * @public */ export class InputDisposedError extends Error { /** Creates a new {@link InputDisposedError}. */ constructor(message = 'Input has been disposed.') { super(message); this.name = 'InputDisposedError'; } } ===== src/id3.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { decodeSynchsafe, encodeSynchsafe } from '../shared/mp3-misc'; import { Logging } from './logging'; import { MetadataTags } from './metadata'; import { coalesceIndex, textDecoder, textEncoder, isIso88591Compatible, assertNever, keyValueIterator, toDataView, isRecordStringString, } from './misc'; import { FileSlice, readAscii, readBytes, readU32Be, readU8 } from './reader'; import { Writer } from './writer'; export type Id3V2Header = { majorVersion: number; revision: number; flags: number; size: number; }; export enum Id3V2HeaderFlags { Unsynchronisation = 1 << 7, ExtendedHeader = 1 << 6, ExperimentalIndicator = 1 << 5, Footer = 1 << 4, } export enum Id3V2TextEncoding { ISO_8859_1, UTF_16_WITH_BOM, UTF_16_BE_NO_BOM, UTF_8, } export const ID3_V1_TAG_SIZE = 128; export const ID3_V2_HEADER_SIZE = 10; export const ID3_V1_GENRES = [ 'Blues', 'Classic rock', 'Country', 'Dance', 'Disco', 'Funk', 'Grunge', 'Hip-hop', 'Jazz', 'Metal', 'New age', 'Oldies', 'Other', 'Pop', 'Rhythm and blues', 'Rap', 'Reggae', 'Rock', 'Techno', 'Industrial', 'Alternative', 'Ska', 'Death metal', 'Pranks', 'Soundtrack', 'Euro-techno', 'Ambient', 'Trip-hop', 'Vocal', 'Jazz & funk', 'Fusion', 'Trance', 'Classical', 'Instrumental', 'Acid', 'House', 'Game', 'Sound clip', 'Gospel', 'Noise', 'Alternative rock', 'Bass', 'Soul', 'Punk', 'Space', 'Meditative', 'Instrumental pop', 'Instrumental rock', 'Ethnic', 'Gothic', 'Darkwave', 'Techno-industrial', 'Electronic', 'Pop-folk', 'Eurodance', 'Dream', 'Southern rock', 'Comedy', 'Cult', 'Gangsta', 'Top 40', 'Christian rap', 'Pop/funk', 'Jungle music', 'Native US', 'Cabaret', 'New wave', 'Psychedelic', 'Rave', 'Showtunes', 'Trailer', 'Lo-fi', 'Tribal', 'Acid punk', 'Acid jazz', 'Polka', 'Retro', 'Musical', 'Rock \'n\' roll', 'Hard rock', 'Folk', 'Folk rock', 'National folk', 'Swing', 'Fast fusion', 'Bebop', 'Latin', 'Revival', 'Celtic', 'Bluegrass', 'Avantgarde', 'Gothic rock', 'Progressive rock', 'Psychedelic rock', 'Symphonic rock', 'Slow rock', 'Big band', 'Chorus', 'Easy listening', 'Acoustic', 'Humour', 'Speech', 'Chanson', 'Opera', 'Chamber music', 'Sonata', 'Symphony', 'Booty bass', 'Primus', 'Porn groove', 'Satire', 'Slow jam', 'Club', 'Tango', 'Samba', 'Folklore', 'Ballad', 'Power ballad', 'Rhythmic Soul', 'Freestyle', 'Duet', 'Punk rock', 'Drum solo', 'A cappella', 'Euro-house', 'Dance hall', 'Goa music', 'Drum & bass', 'Club-house', 'Hardcore techno', 'Terror', 'Indie', 'Britpop', 'Negerpunk', 'Polsk punk', 'Beat', 'Christian gangsta rap', 'Heavy metal', 'Black metal', 'Crossover', 'Contemporary Christian', 'Christian rock', 'Merengue', 'Salsa', 'Thrash metal', 'Anime', 'Jpop', 'Synthpop', 'Christmas', 'Art rock', 'Baroque', 'Bhangra', 'Big beat', 'Breakbeat', 'Chillout', 'Downtempo', 'Dub', 'EBM', 'Eclectic', 'Electro', 'Electroclash', 'Emo', 'Experimental', 'Garage', 'Global', 'IDM', 'Illbient', 'Industro-Goth', 'Jam Band', 'Krautrock', 'Leftfield', 'Lounge', 'Math rock', 'New romantic', 'Nu-breakz', 'Post-punk', 'Post-rock', 'Psytrance', 'Shoegaze', 'Space rock', 'Trop rock', 'World music', 'Neoclassical', 'Audiobook', 'Audio theatre', 'Neue Deutsche Welle', 'Podcast', 'Indie rock', 'G-Funk', 'Dubstep', 'Garage rock', 'Psybient', ]; export const parseId3V1Tag = (slice: FileSlice, tags: MetadataTags) => { const startPos = slice.filePos; tags.raw ??= {}; tags.raw['TAG'] ??= readBytes(slice, ID3_V1_TAG_SIZE - 3); // Dump the whole tag into the raw metadata slice.filePos = startPos; const title = readId3V1String(slice, 30); if (title) tags.title ??= title; const artist = readId3V1String(slice, 30); if (artist) tags.artist ??= artist; const album = readId3V1String(slice, 30); if (album) tags.album ??= album; const yearText = readId3V1String(slice, 4); const year = Number.parseInt(yearText, 10); if (Number.isInteger(year) && year > 0) { tags.date ??= new Date(String(year)); // String so that it parses as UTC } const commentBytes = readBytes(slice, 30); let comment: string; // Check for the ID3v1.1 track number format: // The 29th byte (index 28) is a null terminator, and the 30th byte is the track number. if (commentBytes[28] === 0 && commentBytes[29] !== 0) { const trackNum = commentBytes[29]!; if (trackNum > 0) { tags.trackNumber ??= trackNum; } slice.skip(-30); comment = readId3V1String(slice, 28); slice.skip(2); } else { slice.skip(-30); comment = readId3V1String(slice, 30); } if (comment) tags.comment ??= comment; const genreIndex = readU8(slice); if (genreIndex < ID3_V1_GENRES.length) { tags.genre ??= ID3_V1_GENRES[genreIndex]; } }; export const readId3V1String = (slice: FileSlice, length: number) => { const bytes = readBytes(slice, length); const endIndex = coalesceIndex(bytes.indexOf(0), bytes.length); const relevantBytes = bytes.subarray(0, endIndex); // Decode as ISO-8859-1 let str = ''; for (let i = 0; i < relevantBytes.length; i++) { str += String.fromCharCode(relevantBytes[i]!); } return str.trimEnd(); // String also may be padded with spaces }; export const readId3V2Header = (slice: FileSlice): Id3V2Header | null => { const startPos = slice.filePos; const tag = readAscii(slice, 3); const majorVersion = readU8(slice); const revision = readU8(slice); const flags = readU8(slice); const sizeRaw = readU32Be(slice); if (tag !== 'ID3' || majorVersion === 0xff || revision === 0xff || (sizeRaw & 0x80808080) !== 0) { slice.filePos = startPos; return null; } let size = decodeSynchsafe(sizeRaw); if (flags & Id3V2HeaderFlags.Footer) { size += ID3_V2_HEADER_SIZE; } return { majorVersion, revision, flags, size }; }; export const parseId3V2Tag = (slice: FileSlice, header: Id3V2Header, tags: MetadataTags) => { // https://id3.org/id3v2.3.0 if (![2, 3, 4].includes(header.majorVersion)) { Logging._warn(`Unsupported ID3v2 major version: ${header.majorVersion}`); return; } const dataSize = (header.flags & Id3V2HeaderFlags.Footer) ? header.size - ID3_V2_HEADER_SIZE : header.size; const bytes = readBytes(slice, dataSize); const reader = new Id3V2Reader(header, bytes); if ((header.flags & Id3V2HeaderFlags.Unsynchronisation) && header.majorVersion === 3) { reader.ununsynchronizeAll(); } if (header.flags & Id3V2HeaderFlags.ExtendedHeader) { const extendedHeaderSize = reader.readU32(); if (header.majorVersion === 3) { reader.pos += extendedHeaderSize; // The extended header size excludes itself } else { reader.pos += extendedHeaderSize - 4; // The extended header size includes itself } } while (reader.pos <= reader.bytes.length - reader.frameHeaderSize()) { const frame = reader.readId3V2Frame(); if (!frame) { break; } const frameStartPos = reader.pos; const frameEndPos = reader.pos + frame.size; let frameEncrypted = false; let frameCompressed = false; let frameUnsynchronized = false; if (header.majorVersion === 3) { frameEncrypted = !!(frame.flags & (1 << 6)); frameCompressed = !!(frame.flags & (1 << 7)); } else if (header.majorVersion === 4) { frameEncrypted = !!(frame.flags & (1 << 2)); frameCompressed = !!(frame.flags & (1 << 3)); frameUnsynchronized = !!(frame.flags & (1 << 1)) || !!(header.flags & Id3V2HeaderFlags.Unsynchronisation); } if (frameEncrypted) { Logging._warn(`Skipping encrypted ID3v2 frame ${frame.id}`); reader.pos = frameEndPos; continue; } if (frameCompressed) { Logging._warn(`Skipping compressed ID3v2 frame ${frame.id}`); // Maybe someday? Idk reader.pos = frameEndPos; continue; } if (frameUnsynchronized) { reader.ununsynchronizeRegion(reader.pos, frameEndPos); } tags.raw ??= {}; if (frame.id === 'TXXX') { const txxx = tags.raw['TXXX'] ??= {}; const encoding = reader.readId3V2TextEncoding(); const description = reader.readId3V2Text(encoding, frameEndPos); const value = reader.readId3V2Text(encoding, frameEndPos); (txxx as Record)[description] ??= value; } else if (frame.id[0] === 'T') { // It's a text frame, let's decode as text tags.raw[frame.id] ??= reader.readId3V2EncodingAndText(frameEndPos); } else { // For the others, let's just get the bytes tags.raw[frame.id] ??= reader.readBytes(frame.size); } reader.pos = frameStartPos; switch (frame.id) { case 'TIT2': case 'TT2': { tags.title ??= reader.readId3V2EncodingAndText(frameEndPos); }; break; case 'TIT3': case 'TT3': { tags.description ??= reader.readId3V2EncodingAndText(frameEndPos); }; break; case 'TPE1': case 'TP1': { tags.artist ??= reader.readId3V2EncodingAndText(frameEndPos); }; break; case 'TALB': case 'TAL': { tags.album ??= reader.readId3V2EncodingAndText(frameEndPos); }; break; case 'TPE2': case 'TP2': { tags.albumArtist ??= reader.readId3V2EncodingAndText(frameEndPos); }; break; case 'TRCK': case 'TRK': { const trackText = reader.readId3V2EncodingAndText(frameEndPos); const parts = trackText.split('/'); const trackNum = Number.parseInt(parts[0]!, 10); const tracksTotal = parts[1] && Number.parseInt(parts[1], 10); if (Number.isInteger(trackNum) && trackNum > 0) { tags.trackNumber ??= trackNum; } if (tracksTotal && Number.isInteger(tracksTotal) && tracksTotal > 0) { tags.tracksTotal ??= tracksTotal; } }; break; case 'TPOS': case 'TPA': { const discText = reader.readId3V2EncodingAndText(frameEndPos); const parts = discText.split('/'); const discNum = Number.parseInt(parts[0]!, 10); const discsTotal = parts[1] && Number.parseInt(parts[1], 10); if (Number.isInteger(discNum) && discNum > 0) { tags.discNumber ??= discNum; } if (discsTotal && Number.isInteger(discsTotal) && discsTotal > 0) { tags.discsTotal ??= discsTotal; } }; break; case 'TCON': case 'TCO': { const genreText = reader.readId3V2EncodingAndText(frameEndPos); let match = /^\((\d+)\)/.exec(genreText); if (match) { const genreNumber = Number.parseInt(match[1]!); if (ID3_V1_GENRES[genreNumber] !== undefined) { tags.genre ??= ID3_V1_GENRES[genreNumber]; break; } } match = /^\d+$/.exec(genreText); if (match) { const genreNumber = Number.parseInt(match[0]); if (ID3_V1_GENRES[genreNumber] !== undefined) { tags.genre ??= ID3_V1_GENRES[genreNumber]; break; } } tags.genre ??= genreText; }; break; case 'TDRC': case 'TDAT': { const dateText = reader.readId3V2EncodingAndText(frameEndPos); const date = new Date(dateText); if (!Number.isNaN(date.getTime())) { tags.date ??= date; } }; break; case 'TYER': case 'TYE': { const yearText = reader.readId3V2EncodingAndText(frameEndPos); const year = Number.parseInt(yearText, 10); if (Number.isInteger(year)) { tags.date ??= new Date(String(year)); // String so that it parses as UTC } }; break; case 'USLT': case 'ULT': { const encoding = reader.readU8(); reader.pos += 3; // Skip language reader.readId3V2Text(encoding, frameEndPos); // Short content description tags.lyrics ??= reader.readId3V2Text(encoding, frameEndPos); }; break; case 'COMM': case 'COM': { const encoding = reader.readU8(); reader.pos += 3; // Skip language reader.readId3V2Text(encoding, frameEndPos); // Short content description tags.comment ??= reader.readId3V2Text(encoding, frameEndPos); }; break; case 'APIC': case 'PIC': { const encoding = reader.readId3V2TextEncoding(); let mimeType: string; if (header.majorVersion === 2) { const imageFormat = reader.readAscii(3); mimeType = imageFormat === 'PNG' ? 'image/png' : imageFormat === 'JPG' ? 'image/jpeg' : 'image/*'; } else { mimeType = reader.readId3V2Text(encoding, frameEndPos); } const pictureType = reader.readU8(); const description = reader.readId3V2Text(encoding, frameEndPos).trimEnd(); // Trim ending spaces const imageDataSize = frameEndPos - reader.pos; if (imageDataSize >= 0) { const imageData = reader.readBytes(imageDataSize); if (!tags.images) tags.images = []; tags.images.push({ data: imageData, mimeType, kind: pictureType === 3 ? 'coverFront' : pictureType === 4 ? 'coverBack' : 'unknown', description, }); } }; break; default: { reader.pos += frame.size; }; break; } reader.pos = frameEndPos; } }; // https://id3.org/id3v2.3.0 export class Id3V2Reader { pos = 0; view: DataView; constructor(public header: Id3V2Header, public bytes: Uint8Array) { this.view = new DataView(bytes.buffer, bytes.byteOffset, bytes.byteLength); } frameHeaderSize() { return this.header.majorVersion === 2 ? 6 : 10; } ununsynchronizeAll() { const newBytes: number[] = []; for (let i = 0; i < this.bytes.length; i++) { const value1 = this.bytes[i]!; newBytes.push(value1); if (value1 === 0xff && i !== this.bytes.length - 1) { const value2 = this.bytes[i]!; if (value2 === 0x00) { i++; } } } this.bytes = new Uint8Array(newBytes); this.view = new DataView(this.bytes.buffer); } ununsynchronizeRegion(start: number, end: number) { const newBytes: number[] = []; for (let i = start; i < end; i++) { const value1 = this.bytes[i]!; newBytes.push(value1); if (value1 === 0xff && i !== end - 1) { const value2 = this.bytes[i + 1]!; if (value2 === 0x00) { i++; } } } const before = this.bytes.subarray(0, start); const after = this.bytes.subarray(end); this.bytes = new Uint8Array(before.length + newBytes.length + after.length); this.bytes.set(before, 0); this.bytes.set(newBytes, before.length); this.bytes.set(after, before.length + newBytes.length); this.view = new DataView(this.bytes.buffer); } readBytes(length: number) { const slice = this.bytes.subarray(this.pos, this.pos + length); this.pos += length; return slice; } readU8() { const value = this.view.getUint8(this.pos); this.pos += 1; return value; } readU16() { const value = this.view.getUint16(this.pos, false); this.pos += 2; return value; } readU24() { const high = this.view.getUint16(this.pos, false); const low = this.view.getUint8(this.pos + 2); this.pos += 3; return high * 0x100 + low; } readU32() { const value = this.view.getUint32(this.pos, false); this.pos += 4; return value; } readAscii(length: number) { let str = ''; for (let i = 0; i < length; i++) { str += String.fromCharCode(this.view.getUint8(this.pos + i)); } this.pos += length; return str; } readId3V2Frame() { if (this.header.majorVersion === 2) { const id = this.readAscii(3); if (id === '\x00\x00\x00') { return null; } const size = this.readU24(); return { id, size, flags: 0 }; } else { const id = this.readAscii(4); if (id === '\x00\x00\x00\x00') { // We've landed in the padding section return null; } const sizeRaw = this.readU32(); let size = this.header.majorVersion === 4 ? decodeSynchsafe(sizeRaw) : sizeRaw; const flags = this.readU16(); const headerEndPos = this.pos; // Some files may have incorrectly synchsafed/unsynchsafed sizes. To validate which interpretation is valid, // we validate a size by skipping ahead and seeing if we land at a valid frame header (or at the end of the // tag. const isSizeValid = (size: number) => { const nextPos = this.pos + size; if (nextPos > this.bytes.length) { return false; } if (nextPos <= this.bytes.length - this.frameHeaderSize()) { this.pos += size; const nextId = this.readAscii(4); if (nextId !== '\x00\x00\x00\x00' && !/[0-9A-Z]{4}/.test(nextId)) { return false; } } return true; }; if (!isSizeValid(size)) { // Flip the synchsafing, and try if this one makes more sense const otherSize = this.header.majorVersion === 4 ? sizeRaw : decodeSynchsafe(sizeRaw); if (isSizeValid(otherSize)) { size = otherSize; } } this.pos = headerEndPos; return { id, size, flags }; } } readId3V2TextEncoding(): Id3V2TextEncoding { const number = this.readU8(); if (number > 3) { throw new Error(`Unsupported text encoding: ${number}`); } return number; } readId3V2Text(encoding: Id3V2TextEncoding, until: number): string { const startPos = this.pos; const data = this.readBytes(until - this.pos); switch (encoding) { case Id3V2TextEncoding.ISO_8859_1: { let str = ''; for (let i = 0; i < data.length; i++) { const value = data[i]!; if (value === 0) { this.pos = startPos + i + 1; break; } str += String.fromCharCode(value); } return str; } case Id3V2TextEncoding.UTF_16_WITH_BOM: { if (data[0] === 0xff && data[1] === 0xfe) { const decoder = new TextDecoder('utf-16le'); const endIndex = coalesceIndex( data.findIndex((x, i) => x === 0 && data[i + 1] === 0 && i % 2 === 0), data.length, ); this.pos = startPos + Math.min(endIndex + 2, data.length); return decoder.decode(data.subarray(2, endIndex)); } else if (data[0] === 0xfe && data[1] === 0xff) { const decoder = new TextDecoder('utf-16be'); const endIndex = coalesceIndex( data.findIndex((x, i) => x === 0 && data[i + 1] === 0 && i % 2 === 0), data.length, ); this.pos = startPos + Math.min(endIndex + 2, data.length); return decoder.decode(data.subarray(2, endIndex)); } else { // Treat it like UTF-8, some files do this const endIndex = coalesceIndex(data.findIndex(x => x === 0), data.length); this.pos = startPos + Math.min(endIndex + 1, data.length); return textDecoder.decode(data.subarray(0, endIndex)); } } case Id3V2TextEncoding.UTF_16_BE_NO_BOM: { const decoder = new TextDecoder('utf-16be'); const endIndex = coalesceIndex( data.findIndex((x, i) => x === 0 && data[i + 1] === 0 && i % 2 === 0), data.length, ); this.pos = startPos + Math.min(endIndex + 2, data.length); return decoder.decode(data.subarray(0, endIndex)); } case Id3V2TextEncoding.UTF_8: { const endIndex = coalesceIndex(data.findIndex(x => x === 0), data.length); this.pos = startPos + Math.min(endIndex + 1, data.length); return textDecoder.decode(data.subarray(0, endIndex)); } } } readId3V2EncodingAndText(until: number) { if (this.pos >= until) { return ''; } const encoding = this.readId3V2TextEncoding(); return this.readId3V2Text(encoding, until); } } export class Id3V2Writer { writer: Writer; helper = new Uint8Array(8); helperView = toDataView(this.helper); constructor(writer: Writer) { this.writer = writer; } writeId3V2Tag(metadata: MetadataTags): number { const tagStartPos = this.writer.getPos(); // Write ID3v2.4 header this.writeAscii('ID3'); this.writeU8(0x04); // Version 2.4 this.writeU8(0x00); // Revision 0 this.writeU8(0x00); // Flags this.writeSynchsafeU32(0); // Size placeholder const framesStartPos = this.writer.getPos(); const writtenTags = new Set(); // Write all metadata frames for (const { key, value } of keyValueIterator(metadata)) { switch (key) { case 'title': { this.writeId3V2TextFrame('TIT2', value); writtenTags.add('TIT2'); }; break; case 'description': { this.writeId3V2TextFrame('TIT3', value); writtenTags.add('TIT3'); }; break; case 'artist': { this.writeId3V2TextFrame('TPE1', value); writtenTags.add('TPE1'); }; break; case 'album': { this.writeId3V2TextFrame('TALB', value); writtenTags.add('TALB'); }; break; case 'albumArtist': { this.writeId3V2TextFrame('TPE2', value); writtenTags.add('TPE2'); }; break; case 'trackNumber': { const string = metadata.tracksTotal !== undefined ? `${value}/${metadata.tracksTotal}` : value.toString(); this.writeId3V2TextFrame('TRCK', string); writtenTags.add('TRCK'); }; break; case 'discNumber': { const string = metadata.discsTotal !== undefined ? `${value}/${metadata.discsTotal}` : value.toString(); this.writeId3V2TextFrame('TPOS', string); writtenTags.add('TPOS'); }; break; case 'genre': { this.writeId3V2TextFrame('TCON', value); writtenTags.add('TCON'); }; break; case 'date': { this.writeId3V2TextFrame('TDRC', value.toISOString().slice(0, 10)); writtenTags.add('TDRC'); }; break; case 'lyrics': { this.writeId3V2LyricsFrame(value); writtenTags.add('USLT'); }; break; case 'comment': { this.writeId3V2CommentFrame(value); writtenTags.add('COMM'); }; break; case 'images': { const pictureTypeMap = { coverFront: 0x03, coverBack: 0x04, unknown: 0x00 }; for (const image of value) { const pictureType = pictureTypeMap[image.kind] ?? 0x00; const description = image.description ?? ''; this.writeId3V2ApicFrame(image.mimeType, pictureType, description, image.data); } }; break; case 'tracksTotal': case 'discsTotal': { // Handled with trackNumber and discNumber respectively }; break; case 'raw': { // Handled later }; break; default: { assertNever(key); } } } if (metadata.raw) { for (const key in metadata.raw) { const value = metadata.raw[key]; if (value == null || key.length !== 4 || writtenTags.has(key)) { continue; } let bytes: Uint8Array; if (typeof value === 'string') { const useIso88591 = isIso88591Compatible(value); if (useIso88591) { bytes = new Uint8Array(value.length + 2); bytes[0] = Id3V2TextEncoding.ISO_8859_1; for (let i = 0; i < value.length; i++) { bytes[i + 1] = value.charCodeAt(i); } // Last byte is the null terminator } else { const encoded = textEncoder.encode(value); bytes = new Uint8Array(encoded.byteLength + 2); bytes[0] = Id3V2TextEncoding.UTF_8; bytes.set(encoded, 1); // Last byte is the null terminator } } else if (value instanceof Uint8Array) { bytes = value; } else if (key === 'TXXX' && isRecordStringString(value)) { for (const description in value) { const frameValue = value[description]!; const useIso88591 = isIso88591Compatible(description) && isIso88591Compatible(frameValue); const encodedDescription = useIso88591 ? null : textEncoder.encode(description); const encodedValue = useIso88591 ? null : textEncoder.encode(frameValue); const descriptionDataLength = useIso88591 ? description.length : encodedDescription!.byteLength; const valueDataLength = useIso88591 ? frameValue.length : encodedValue!.byteLength; const frameSize = 1 + descriptionDataLength + 1 + valueDataLength + 1; this.writeAscii('TXXX'); this.writeSynchsafeU32(frameSize); this.writeU16(0x0000); this.writeU8(useIso88591 ? Id3V2TextEncoding.ISO_8859_1 : Id3V2TextEncoding.UTF_8); if (useIso88591) { this.writeIsoString(description); this.writeIsoString(frameValue); } else { this.writer.write(encodedDescription!); this.writeU8(0x00); this.writer.write(encodedValue!); this.writeU8(0x00); } } continue; } else { continue; } this.writeAscii(key); this.writeSynchsafeU32(bytes.byteLength); this.writeU16(0x0000); this.writer.write(bytes); } } const framesEndPos = this.writer.getPos(); const framesSize = framesEndPos - framesStartPos; // Update the size field in the header (synchsafe) this.writer.seek(tagStartPos + 6); // Skip 'ID3' + version + revision + flags this.writeSynchsafeU32(framesSize); this.writer.seek(framesEndPos); return framesSize + 10; // +10 for the header size } writeU8(value: number) { this.helper[0] = value; this.writer.write(this.helper.subarray(0, 1)); } writeU16(value: number) { this.helperView.setUint16(0, value, false); this.writer.write(this.helper.subarray(0, 2)); } writeU32(value: number) { this.helperView.setUint32(0, value, false); this.writer.write(this.helper.subarray(0, 4)); } writeAscii(text: string) { for (let i = 0; i < text.length; i++) { this.helper[i] = text.charCodeAt(i); } this.writer.write(this.helper.subarray(0, text.length)); } writeSynchsafeU32(value: number) { this.writeU32(encodeSynchsafe(value)); } writeIsoString(text: string) { const bytes = new Uint8Array(text.length + 1); for (let i = 0; i < text.length; i++) { bytes[i] = text.charCodeAt(i); } // Last byte is the null terminator this.writer.write(bytes); } writeUtf8String(text: string) { const utf8Data = textEncoder.encode(text); this.writer.write(utf8Data); this.writeU8(0x00); } writeId3V2TextFrame(frameId: string, text: string) { const useIso88591 = isIso88591Compatible(text); const textDataLength = useIso88591 ? text.length : textEncoder.encode(text).byteLength; const frameSize = 1 + textDataLength + 1; this.writeAscii(frameId); this.writeSynchsafeU32(frameSize); this.writeU16(0x0000); this.writeU8(useIso88591 ? Id3V2TextEncoding.ISO_8859_1 : Id3V2TextEncoding.UTF_8); if (useIso88591) { this.writeIsoString(text); } else { this.writeUtf8String(text); } } writeId3V2LyricsFrame(lyrics: string) { const useIso88591 = isIso88591Compatible(lyrics); const shortDescription = ''; const frameSize = 1 + 3 + shortDescription.length + 1 + lyrics.length + 1; this.writeAscii('USLT'); this.writeSynchsafeU32(frameSize); this.writeU16(0x0000); this.writeU8(useIso88591 ? Id3V2TextEncoding.ISO_8859_1 : Id3V2TextEncoding.UTF_8); this.writeAscii('und'); if (useIso88591) { this.writeIsoString(shortDescription); this.writeIsoString(lyrics); } else { this.writeUtf8String(shortDescription); this.writeUtf8String(lyrics); } } writeId3V2CommentFrame(comment: string) { const useIso88591 = isIso88591Compatible(comment); const textDataLength = useIso88591 ? comment.length : textEncoder.encode(comment).byteLength; const shortDescription = ''; const frameSize = 1 + 3 + shortDescription.length + 1 + textDataLength + 1; this.writeAscii('COMM'); this.writeSynchsafeU32(frameSize); this.writeU16(0x0000); this.writeU8(useIso88591 ? Id3V2TextEncoding.ISO_8859_1 : Id3V2TextEncoding.UTF_8); this.writeU8(0x75); // 'u' this.writeU8(0x6E); // 'n' this.writeU8(0x64); // 'd' if (useIso88591) { this.writeIsoString(shortDescription); this.writeIsoString(comment); } else { this.writeUtf8String(shortDescription); this.writeUtf8String(comment); } } writeId3V2ApicFrame(mimeType: string, pictureType: number, description: string, imageData: Uint8Array) { const useIso88591 = isIso88591Compatible(mimeType) && isIso88591Compatible(description); const descriptionDataLength = useIso88591 ? description.length : textEncoder.encode(description).byteLength; const frameSize = 1 + mimeType.length + 1 + 1 + descriptionDataLength + 1 + imageData.byteLength; this.writeAscii('APIC'); this.writeSynchsafeU32(frameSize); this.writeU16(0x0000); this.writeU8(useIso88591 ? Id3V2TextEncoding.ISO_8859_1 : Id3V2TextEncoding.UTF_8); if (useIso88591) { this.writeIsoString(mimeType); } else { this.writeUtf8String(mimeType); } this.writeU8(pictureType); if (useIso88591) { this.writeIsoString(description); } else { this.writeUtf8String(description); } this.writer.write(imageData); } } ===== src/codec-data.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AVC_LEVEL_TABLE, VideoCodec, VP9_LEVEL_TABLE } from './codec'; import { assert, assertNever, base64ToBytes, bytesToBase64, keyValueIterator, getUint24, last, readExpGolomb, readSignedExpGolomb, Rational, textDecoder, textEncoder, toDataView, toUint8Array, getChromiumVersion, isChromium, setUint24, } from './misc'; import { Logging } from './logging'; import { PacketType } from './packet'; import { MetadataTags } from './metadata'; import { AC3_SAMPLE_RATES, EAC3_REDUCED_SAMPLE_RATES } from '../shared/ac3-misc'; import { Bitstream } from '../shared/bitstream'; // References for AVC/HEVC code: // ISO 14496-15 // Rec. ITU-T H.264 // Rec. ITU-T H.265 // https://stackoverflow.com/questions/24884827 export enum AvcNalUnitType { NON_IDR_SLICE = 1, SLICE_DPA = 2, SLICE_DPB = 3, SLICE_DPC = 4, IDR = 5, SEI = 6, SPS = 7, PPS = 8, AUD = 9, SPS_EXT = 13, } export enum HevcNalUnitType { RASL_N = 8, RASL_R = 9, BLA_W_LP = 16, RSV_IRAP_VCL23 = 23, VPS_NUT = 32, SPS_NUT = 33, PPS_NUT = 34, AUD_NUT = 35, PREFIX_SEI_NUT = 39, SUFFIX_SEI_NUT = 40, } export type NalUnitLocation = { offset: number; length: number; }; export const iterateNalUnitsInAnnexB = function* (packetData: Uint8Array): Generator { let i = 0; let nalStart = -1; while (i < packetData.length - 2) { const zeroIndex = packetData.indexOf(0, i); if (zeroIndex === -1 || zeroIndex >= packetData.length - 2) { break; } i = zeroIndex; let startCodeLength = 0; // Check for 4-byte start code (0x00000001) if ( i + 3 < packetData.length && packetData[i + 1] === 0 && packetData[i + 2] === 0 && packetData[i + 3] === 1 ) { startCodeLength = 4; } else if (packetData[i + 1] === 0 && packetData[i + 2] === 1) { // Check for 3-byte start code (0x000001) startCodeLength = 3; } if (startCodeLength === 0) { i++; continue; } // If we had a previous NAL unit, yield it if (nalStart !== -1 && i > nalStart) { yield { offset: nalStart, length: i - nalStart, }; } nalStart = i + startCodeLength; i = nalStart; } // Yield the last NAL unit if there is one if (nalStart !== -1 && nalStart < packetData.length) { yield { offset: nalStart, length: packetData.length - nalStart, }; } }; export const iterateNalUnitsInLengthPrefixed = function* ( packetData: Uint8Array, lengthSize: 1 | 2 | 3 | 4, ): Generator { let offset = 0; const dataView = new DataView(packetData.buffer, packetData.byteOffset, packetData.byteLength); while (offset + lengthSize <= packetData.length) { let nalUnitLength: number; if (lengthSize === 1) { nalUnitLength = dataView.getUint8(offset); } else if (lengthSize === 2) { nalUnitLength = dataView.getUint16(offset, false); } else if (lengthSize === 3) { nalUnitLength = getUint24(dataView, offset, false); } else { assert(lengthSize === 4); nalUnitLength = dataView.getUint32(offset, false); } offset += lengthSize; yield { offset, length: nalUnitLength, }; offset += nalUnitLength; } }; export const iterateAvcNalUnits = (packetData: Uint8Array, decoderConfig: VideoDecoderConfig) => { if (decoderConfig.description) { const bytes = toUint8Array(decoderConfig.description); const lengthSizeMinusOne = bytes[4]! & 0b11; const lengthSize = (lengthSizeMinusOne + 1) as 1 | 2 | 3 | 4; return iterateNalUnitsInLengthPrefixed(packetData, lengthSize); } else { return iterateNalUnitsInAnnexB(packetData); } }; export const extractNalUnitTypeForAvc = (byte: number) => { return byte & 0x1F; }; const removeEmulationPreventionBytes = (data: Uint8Array) => { const result: number[] = []; const len = data.length; for (let i = 0; i < len; i++) { // Look for the 0x000003 pattern if (i + 2 < len && data[i] === 0x00 && data[i + 1] === 0x00 && data[i + 2] === 0x03) { result.push(0x00, 0x00); // Push the first two bytes i += 2; // Skip the 0x03 byte } else { result.push(data[i]!); } } return new Uint8Array(result); }; const ANNEX_B_START_CODE = new Uint8Array([0, 0, 0, 1]); export const concatNalUnitsInAnnexB = (nalUnits: Uint8Array[]) => { const totalLength = nalUnits.reduce((a, b) => a + ANNEX_B_START_CODE.byteLength + b.byteLength, 0); const result = new Uint8Array(totalLength); let offset = 0; for (const nalUnit of nalUnits) { result.set(ANNEX_B_START_CODE, offset); offset += ANNEX_B_START_CODE.byteLength; result.set(nalUnit, offset); offset += nalUnit.byteLength; } return result; }; export const concatNalUnitsInLengthPrefixed = (nalUnits: Uint8Array[], lengthSize: 1 | 2 | 3 | 4) => { const totalLength = nalUnits.reduce((a, b) => a + lengthSize + b.byteLength, 0); const result = new Uint8Array(totalLength); let offset = 0; for (const nalUnit of nalUnits) { const dataView = new DataView(result.buffer, result.byteOffset, result.byteLength); switch (lengthSize) { case 1: dataView.setUint8(offset, nalUnit.byteLength); break; case 2: dataView.setUint16(offset, nalUnit.byteLength, false); break; case 3: setUint24(dataView, offset, nalUnit.byteLength, false); break; case 4: dataView.setUint32(offset, nalUnit.byteLength, false); break; } offset += lengthSize; result.set(nalUnit, offset); offset += nalUnit.byteLength; } return result; }; // Data specified in ISO 14496-15 export type AvcDecoderConfigurationRecord = { configurationVersion: number; avcProfileIndication: number; profileCompatibility: number; avcLevelIndication: number; lengthSizeMinusOne: number; sequenceParameterSets: Uint8Array[]; pictureParameterSets: Uint8Array[]; // Fields only for specific profiles: chromaFormat: number | null; bitDepthLumaMinus8: number | null; bitDepthChromaMinus8: number | null; sequenceParameterSetExt: Uint8Array[] | null; }; export const concatAvcNalUnits = (nalUnits: Uint8Array[], decoderConfig: VideoDecoderConfig) => { if (decoderConfig.description) { // Stream is length-prefixed. Let's extract the size of the length prefix from the decoder config const bytes = toUint8Array(decoderConfig.description); const lengthSizeMinusOne = bytes[4]! & 0b11; const lengthSize = (lengthSizeMinusOne + 1) as 1 | 2 | 3 | 4; return concatNalUnitsInLengthPrefixed(nalUnits, lengthSize); } else { // Stream is in Annex B format return concatNalUnitsInAnnexB(nalUnits); } }; /** Builds an AvcDecoderConfigurationRecord from an AVC packet in Annex B format. */ export const extractAvcDecoderConfigurationRecord = (packetData: Uint8Array): AvcDecoderConfigurationRecord | null => { try { const spsUnits: Uint8Array[] = []; const ppsUnits: Uint8Array[] = []; const spsExtUnits: Uint8Array[] = []; for (const loc of iterateNalUnitsInAnnexB(packetData)) { const nalUnit = packetData.subarray(loc.offset, loc.offset + loc.length); const type = extractNalUnitTypeForAvc(nalUnit[0]!); if (type === AvcNalUnitType.SPS) { spsUnits.push(nalUnit); } else if (type === AvcNalUnitType.PPS) { ppsUnits.push(nalUnit); } else if (type === AvcNalUnitType.SPS_EXT) { spsExtUnits.push(nalUnit); } } if (spsUnits.length === 0) { return null; } if (ppsUnits.length === 0) { return null; } // Let's get the first SPS for profile and level information const spsData = spsUnits[0]!; const spsInfo = parseAvcSps(spsData); assert(spsInfo !== null); const hasExtendedData = spsInfo.profileIdc === 100 || spsInfo.profileIdc === 110 || spsInfo.profileIdc === 122 || spsInfo.profileIdc === 144; return { configurationVersion: 1, avcProfileIndication: spsInfo.profileIdc, profileCompatibility: spsInfo.constraintFlags, avcLevelIndication: spsInfo.levelIdc, lengthSizeMinusOne: 3, // Typically 4 bytes for length field sequenceParameterSets: spsUnits, pictureParameterSets: ppsUnits, chromaFormat: hasExtendedData ? spsInfo.chromaFormatIdc : null, bitDepthLumaMinus8: hasExtendedData ? spsInfo.bitDepthLumaMinus8 : null, bitDepthChromaMinus8: hasExtendedData ? spsInfo.bitDepthChromaMinus8 : null, sequenceParameterSetExt: hasExtendedData ? spsExtUnits : null, }; } catch (error) { Logging._error('Error building AVC Decoder Configuration Record:', error); return null; } }; /** Serializes an AvcDecoderConfigurationRecord into the format specified in Section 5.3.3.1 of ISO 14496-15. */ export const serializeAvcDecoderConfigurationRecord = (record: AvcDecoderConfigurationRecord) => { const bytes: number[] = []; // Write header bytes.push(record.configurationVersion); bytes.push(record.avcProfileIndication); bytes.push(record.profileCompatibility); bytes.push(record.avcLevelIndication); bytes.push(0xFC | (record.lengthSizeMinusOne & 0x03)); // Reserved bits (6) + lengthSizeMinusOne (2) // Reserved bits (3) + numOfSequenceParameterSets (5) bytes.push(0xE0 | (record.sequenceParameterSets.length & 0x1F)); // Write SPS for (const sps of record.sequenceParameterSets) { const length = sps.byteLength; bytes.push(length >> 8); // High byte bytes.push(length & 0xFF); // Low byte for (let i = 0; i < length; i++) { bytes.push(sps[i]!); } } bytes.push(record.pictureParameterSets.length); // Write PPS for (const pps of record.pictureParameterSets) { const length = pps.byteLength; bytes.push(length >> 8); // High byte bytes.push(length & 0xFF); // Low byte for (let i = 0; i < length; i++) { bytes.push(pps[i]!); } } if ( record.avcProfileIndication === 100 || record.avcProfileIndication === 110 || record.avcProfileIndication === 122 || record.avcProfileIndication === 144 ) { assert(record.chromaFormat !== null); assert(record.bitDepthLumaMinus8 !== null); assert(record.bitDepthChromaMinus8 !== null); assert(record.sequenceParameterSetExt !== null); bytes.push(0xFC | (record.chromaFormat & 0x03)); // Reserved bits + chroma_format bytes.push(0xF8 | (record.bitDepthLumaMinus8 & 0x07)); // Reserved bits + bit_depth_luma_minus8 bytes.push(0xF8 | (record.bitDepthChromaMinus8 & 0x07)); // Reserved bits + bit_depth_chroma_minus8 bytes.push(record.sequenceParameterSetExt.length); // Write SPS Ext for (const spsExt of record.sequenceParameterSetExt) { const length = spsExt.byteLength; bytes.push(length >> 8); // High byte bytes.push(length & 0xFF); // Low byte for (let i = 0; i < length; i++) { bytes.push(spsExt[i]!); } } } return new Uint8Array(bytes); }; /** Deserializes an AvcDecoderConfigurationRecord from the format specified in Section 5.3.3.1 of ISO 14496-15. */ export const deserializeAvcDecoderConfigurationRecord = (data: Uint8Array): AvcDecoderConfigurationRecord | null => { try { const view = toDataView(data); let offset = 0; // Read header const configurationVersion = view.getUint8(offset++); const avcProfileIndication = view.getUint8(offset++); const profileCompatibility = view.getUint8(offset++); const avcLevelIndication = view.getUint8(offset++); const lengthSizeMinusOne = view.getUint8(offset++) & 0x03; const numOfSequenceParameterSets = view.getUint8(offset++) & 0x1F; // Read SPS const sequenceParameterSets: Uint8Array[] = []; for (let i = 0; i < numOfSequenceParameterSets; i++) { const length = view.getUint16(offset, false); offset += 2; sequenceParameterSets.push(data.subarray(offset, offset + length)); offset += length; } const numOfPictureParameterSets = view.getUint8(offset++); // Read PPS const pictureParameterSets: Uint8Array[] = []; for (let i = 0; i < numOfPictureParameterSets; i++) { const length = view.getUint16(offset, false); offset += 2; pictureParameterSets.push(data.subarray(offset, offset + length)); offset += length; } const record: AvcDecoderConfigurationRecord = { configurationVersion, avcProfileIndication, profileCompatibility, avcLevelIndication, lengthSizeMinusOne, sequenceParameterSets, pictureParameterSets, chromaFormat: null, bitDepthLumaMinus8: null, bitDepthChromaMinus8: null, sequenceParameterSetExt: null, }; // Check if there are extended profile fields if ( ( avcProfileIndication === 100 || avcProfileIndication === 110 || avcProfileIndication === 122 || avcProfileIndication === 144 ) && offset + 4 <= data.length ) { const chromaFormat = view.getUint8(offset++) & 0x03; const bitDepthLumaMinus8 = view.getUint8(offset++) & 0x07; const bitDepthChromaMinus8 = view.getUint8(offset++) & 0x07; const numOfSequenceParameterSetExt = view.getUint8(offset++); record.chromaFormat = chromaFormat; record.bitDepthLumaMinus8 = bitDepthLumaMinus8; record.bitDepthChromaMinus8 = bitDepthChromaMinus8; // Read SPS Ext const sequenceParameterSetExt: Uint8Array[] = []; for (let i = 0; i < numOfSequenceParameterSetExt; i++) { const length = view.getUint16(offset, false); offset += 2; sequenceParameterSetExt.push(data.subarray(offset, offset + length)); offset += length; } record.sequenceParameterSetExt = sequenceParameterSetExt; } return record; } catch (error) { Logging._error('Error deserializing AVC Decoder Configuration Record:', error); return null; } }; export type AvcSpsInfo = { profileIdc: number; constraintFlags: number; levelIdc: number; frameMbsOnlyFlag: number; chromaFormatIdc: number; bitDepthLumaMinus8: number; bitDepthChromaMinus8: number; codedWidth: number; codedHeight: number; displayWidth: number; displayHeight: number; pixelAspectRatio: Rational; colourPrimaries: number; transferCharacteristics: number; matrixCoefficients: number; fullRangeFlag: number; numReorderFrames: number; maxDecFrameBuffering: number; }; const AVC_HEVC_ASPECT_RATIO_IDC_TABLE: Partial> = { 1: { num: 1, den: 1 }, 2: { num: 12, den: 11 }, 3: { num: 10, den: 11 }, 4: { num: 16, den: 11 }, 5: { num: 40, den: 33 }, 6: { num: 24, den: 11 }, 7: { num: 20, den: 11 }, 8: { num: 32, den: 11 }, 9: { num: 80, den: 33 }, 10: { num: 18, den: 11 }, 11: { num: 15, den: 11 }, 12: { num: 64, den: 33 }, 13: { num: 160, den: 99 }, 14: { num: 4, den: 3 }, 15: { num: 3, den: 2 }, 16: { num: 2, den: 1 }, }; /** Parses an AVC SPS (Sequence Parameter Set) to extract basic information. */ export const parseAvcSps = (sps: Uint8Array): AvcSpsInfo | null => { try { const bitstream = new Bitstream(removeEmulationPreventionBytes(sps)); bitstream.skipBits(1); // forbidden_zero_bit bitstream.skipBits(2); // nal_ref_idc const nalUnitType = bitstream.readBits(5); if (nalUnitType !== 7) { // SPS NAL unit type is 7 return null; } const profileIdc = bitstream.readAlignedByte(); const constraintFlags = bitstream.readAlignedByte(); const levelIdc = bitstream.readAlignedByte(); readExpGolomb(bitstream); // seq_parameter_set_id // "When chroma_format_idc is not present, it shall be inferred to be equal to 1 (4:2:0 chroma format)." let chromaFormatIdc = 1; // "When bit_depth_luma_minus8 is not present, it shall be inferred to be equal to 0."" let bitDepthLumaMinus8 = 0; // "When bit_depth_chroma_minus8 is not present, it shall be inferred to be equal to 0." let bitDepthChromaMinus8 = 0; // "When separate_colour_plane_flag is not present, it shall be inferred to be equal to 0." let separateColourPlaneFlag = 0; // Handle high profile chroma_format_idc if ( profileIdc === 100 || profileIdc === 110 || profileIdc === 122 || profileIdc === 244 || profileIdc === 44 || profileIdc === 83 || profileIdc === 86 || profileIdc === 118 || profileIdc === 128 ) { chromaFormatIdc = readExpGolomb(bitstream); if (chromaFormatIdc === 3) { separateColourPlaneFlag = bitstream.readBits(1); } bitDepthLumaMinus8 = readExpGolomb(bitstream); bitDepthChromaMinus8 = readExpGolomb(bitstream); bitstream.skipBits(1); // qpprime_y_zero_transform_bypass_flag const seqScalingMatrixPresentFlag = bitstream.readBits(1); if (seqScalingMatrixPresentFlag) { for (let i = 0; i < (chromaFormatIdc !== 3 ? 8 : 12); i++) { const seqScalingListPresentFlag = bitstream.readBits(1); if (seqScalingListPresentFlag) { const sizeOfScalingList = i < 6 ? 16 : 64; let lastScale = 8; let nextScale = 8; for (let j = 0; j < sizeOfScalingList; j++) { if (nextScale !== 0) { const deltaScale = readSignedExpGolomb(bitstream); nextScale = (lastScale + deltaScale + 256) % 256; } lastScale = nextScale === 0 ? lastScale : nextScale; } } } } } readExpGolomb(bitstream); // log2_max_frame_num_minus4 const picOrderCntType = readExpGolomb(bitstream); if (picOrderCntType === 0) { readExpGolomb(bitstream); // log2_max_pic_order_cnt_lsb_minus4 } else if (picOrderCntType === 1) { bitstream.skipBits(1); // delta_pic_order_always_zero_flag readSignedExpGolomb(bitstream); // offset_for_non_ref_pic readSignedExpGolomb(bitstream); // offset_for_top_to_bottom_field const numRefFramesInPicOrderCntCycle = readExpGolomb(bitstream); for (let i = 0; i < numRefFramesInPicOrderCntCycle; i++) { readSignedExpGolomb(bitstream); // offset_for_ref_frame[i] } } readExpGolomb(bitstream); // max_num_ref_frames bitstream.skipBits(1); // gaps_in_frame_num_value_allowed_flag const picWidthInMbsMinus1 = readExpGolomb(bitstream); const picHeightInMapUnitsMinus1 = readExpGolomb(bitstream); const codedWidth = 16 * (picWidthInMbsMinus1 + 1); const codedHeight = 16 * (picHeightInMapUnitsMinus1 + 1); let displayWidth = codedWidth; let displayHeight = codedHeight; const frameMbsOnlyFlag = bitstream.readBits(1); if (!frameMbsOnlyFlag) { bitstream.skipBits(1); // mb_adaptive_frame_field_flag } bitstream.skipBits(1); // direct_8x8_inference_flag const frameCroppingFlag = bitstream.readBits(1); if (frameCroppingFlag) { const frameCropLeftOffset = readExpGolomb(bitstream); const frameCropRightOffset = readExpGolomb(bitstream); const frameCropTopOffset = readExpGolomb(bitstream); const frameCropBottomOffset = readExpGolomb(bitstream); let cropUnitX: number; let cropUnitY: number; const chromaArrayType = separateColourPlaneFlag === 0 ? chromaFormatIdc : 0; if (chromaArrayType === 0) { // "If ChromaArrayType is equal to 0, CropUnitX and CropUnitY are derived as:" cropUnitX = 1; cropUnitY = 2 - frameMbsOnlyFlag; } else { // "Otherwise (ChromaArrayType is equal to 1, 2, or 3), CropUnitX and CropUnitY are derived as:" const subWidthC = chromaFormatIdc === 3 ? 1 : 2; const subHeightC = chromaFormatIdc === 1 ? 2 : 1; cropUnitX = subWidthC; cropUnitY = subHeightC * (2 - frameMbsOnlyFlag); } displayWidth -= (cropUnitX * (frameCropLeftOffset + frameCropRightOffset)); displayHeight -= (cropUnitY * (frameCropTopOffset + frameCropBottomOffset)); } // 2 = unspecified let colourPrimaries = 2; let transferCharacteristics = 2; let matrixCoefficients = 2; let fullRangeFlag = 0; let pixelAspectRatio: Rational = { num: 1, den: 1 }; let numReorderFrames: number | null = null; let maxDecFrameBuffering: number | null = null; const vuiParametersPresentFlag = bitstream.readBits(1); if (vuiParametersPresentFlag) { const aspectRatioInfoPresentFlag = bitstream.readBits(1); if (aspectRatioInfoPresentFlag) { const aspectRatioIdc = bitstream.readBits(8); if (aspectRatioIdc === 255) { // Extended_SAR pixelAspectRatio = { num: bitstream.readBits(16), den: bitstream.readBits(16), }; } else { const aspectRatio = AVC_HEVC_ASPECT_RATIO_IDC_TABLE[aspectRatioIdc]; if (aspectRatio) { pixelAspectRatio = aspectRatio; } } } const overscanInfoPresentFlag = bitstream.readBits(1); if (overscanInfoPresentFlag) { bitstream.skipBits(1); // overscan_appropriate_flag } const videoSignalTypePresentFlag = bitstream.readBits(1); if (videoSignalTypePresentFlag) { bitstream.skipBits(3); // video_format fullRangeFlag = bitstream.readBits(1); const colourDescriptionPresentFlag = bitstream.readBits(1); if (colourDescriptionPresentFlag) { colourPrimaries = bitstream.readBits(8); transferCharacteristics = bitstream.readBits(8); matrixCoefficients = bitstream.readBits(8); } } const chromaLocInfoPresentFlag = bitstream.readBits(1); if (chromaLocInfoPresentFlag) { readExpGolomb(bitstream); // chroma_sample_loc_type_top_field readExpGolomb(bitstream); // chroma_sample_loc_type_bottom_field } const timingInfoPresentFlag = bitstream.readBits(1); if (timingInfoPresentFlag) { bitstream.skipBits(32); // num_units_in_tick bitstream.skipBits(32); // time_scale bitstream.skipBits(1); // fixed_frame_rate_flag } const nalHrdParametersPresentFlag = bitstream.readBits(1); if (nalHrdParametersPresentFlag) { skipAvcHrdParameters(bitstream); } const vclHrdParametersPresentFlag = bitstream.readBits(1); if (vclHrdParametersPresentFlag) { skipAvcHrdParameters(bitstream); } if (nalHrdParametersPresentFlag || vclHrdParametersPresentFlag) { bitstream.skipBits(1); // low_delay_hrd_flag } bitstream.skipBits(1); // pic_struct_present_flag const bitstreamRestrictionFlag = bitstream.readBits(1); if (bitstreamRestrictionFlag) { bitstream.skipBits(1); // motion_vectors_over_pic_boundaries_flag readExpGolomb(bitstream); // max_bytes_per_pic_denom readExpGolomb(bitstream); // max_bits_per_mb_denom readExpGolomb(bitstream); // log2_max_mv_length_horizontal readExpGolomb(bitstream); // log2_max_mv_length_vertical numReorderFrames = readExpGolomb(bitstream); maxDecFrameBuffering = readExpGolomb(bitstream); } } if (numReorderFrames === null) { assert(maxDecFrameBuffering === null); const constraintSet3Flag = constraintFlags & 0b00010000; if ( (profileIdc === 44 || profileIdc === 86 || profileIdc === 100 || profileIdc === 110 || profileIdc === 122 || profileIdc === 244 ) && constraintSet3Flag ) { // "If profile_idc is equal to 44, 86, 100, 110, 122, or 244 and constraint_set3_flag is equal to 1, the // value of num_reorder_frames shall be inferred to be equal to 0." numReorderFrames = 0; maxDecFrameBuffering = 0; } else { const picWidthInMbs = picWidthInMbsMinus1 + 1; const picHeightInMapUnits = picHeightInMapUnitsMinus1 + 1; const frameHeightInMbs = (2 - frameMbsOnlyFlag) * picHeightInMapUnits; const levelInfo = AVC_LEVEL_TABLE.find( x => x.level >= levelIdc, ) ?? last(AVC_LEVEL_TABLE)!; // "MaxDpbFrames is equal to // Min( MaxDpbMbs / ( picWidthInMbs * frameHeightInMbs ), 16 ) and MaxDpbMbs is given in Table A-1." const maxDpbFrames = Math.min( Math.floor(levelInfo.maxDpbMbs / (picWidthInMbs * frameHeightInMbs)), 16, ); // "Otherwise, [...] the value of num_reorder_frames shall be inferred to be equal to MaxDpbFrames." numReorderFrames = maxDpbFrames; maxDecFrameBuffering = maxDpbFrames; } } assert(maxDecFrameBuffering !== null); return { profileIdc, constraintFlags, levelIdc, frameMbsOnlyFlag, chromaFormatIdc, bitDepthLumaMinus8, bitDepthChromaMinus8, codedWidth, codedHeight, displayWidth, displayHeight, pixelAspectRatio, colourPrimaries, matrixCoefficients, transferCharacteristics, fullRangeFlag, numReorderFrames, maxDecFrameBuffering, }; } catch (error) { Logging._error('Error parsing AVC SPS:', error); return null; } }; const skipAvcHrdParameters = (bitstream: Bitstream) => { const cpb_cnt_minus1 = readExpGolomb(bitstream); bitstream.skipBits(4); // bit_rate_scale bitstream.skipBits(4); // cpb_size_scale for (let i = 0; i <= cpb_cnt_minus1; i++) { readExpGolomb(bitstream); // bit_rate_value_minus1[i] readExpGolomb(bitstream); // cpb_size_value_minus1[i] bitstream.skipBits(1); // cbr_flag[i] } bitstream.skipBits(5); // initial_cpb_removal_delay_length_minus1 bitstream.skipBits(5); // cpb_removal_delay_length_minus1 bitstream.skipBits(5); // dpb_output_delay_length_minus1 bitstream.skipBits(5); // time_offset_length }; // Data specified in ISO 14496-15 export type HevcDecoderConfigurationRecord = { configurationVersion: number; generalProfileSpace: number; generalTierFlag: number; generalProfileIdc: number; generalProfileCompatibilityFlags: number; generalConstraintIndicatorFlags: Uint8Array; // 6 bytes long generalLevelIdc: number; minSpatialSegmentationIdc: number; parallelismType: number; chromaFormatIdc: number; bitDepthLumaMinus8: number; bitDepthChromaMinus8: number; avgFrameRate: number; constantFrameRate: number; numTemporalLayers: number; temporalIdNested: number; lengthSizeMinusOne: number; arrays: { arrayCompleteness: number; nalUnitType: number; nalUnits: Uint8Array[]; }[]; }; export type HevcSpsInfo = { displayWidth: number; displayHeight: number; pixelAspectRatio: Rational; colourPrimaries: number; transferCharacteristics: number; matrixCoefficients: number; fullRangeFlag: number; maxDecFrameBuffering: number; spsMaxSubLayersMinus1: number; spsTemporalIdNestingFlag: number; generalProfileSpace: number; generalTierFlag: number; generalProfileIdc: number; generalProfileCompatibilityFlags: number; generalConstraintIndicatorFlags: Uint8Array; generalLevelIdc: number; chromaFormatIdc: number; bitDepthLumaMinus8: number; bitDepthChromaMinus8: number; minSpatialSegmentationIdc: number; }; export const concatHevcNalUnits = (nalUnits: Uint8Array[], decoderConfig: VideoDecoderConfig) => { if (decoderConfig.description) { // Stream is length-prefixed. Let's extract the size of the length prefix from the decoder config const bytes = toUint8Array(decoderConfig.description); const lengthSizeMinusOne = bytes[21]! & 0b11; const lengthSize = (lengthSizeMinusOne + 1) as 1 | 2 | 3 | 4; return concatNalUnitsInLengthPrefixed(nalUnits, lengthSize); } else { // Stream is in Annex B format return concatNalUnitsInAnnexB(nalUnits); } }; export const iterateHevcNalUnits = (packetData: Uint8Array, decoderConfig: VideoDecoderConfig) => { if (decoderConfig.description) { const bytes = toUint8Array(decoderConfig.description); const lengthSizeMinusOne = bytes[21]! & 0b11; const lengthSize = (lengthSizeMinusOne + 1) as 1 | 2 | 3 | 4; return iterateNalUnitsInLengthPrefixed(packetData, lengthSize); } else { return iterateNalUnitsInAnnexB(packetData); } }; export const extractNalUnitTypeForHevc = (byte: number) => { return (byte >> 1) & 0x3F; }; /** Parses an HEVC SPS (Sequence Parameter Set) to extract video information. */ export const parseHevcSps = (sps: Uint8Array): HevcSpsInfo | null => { try { const bitstream = new Bitstream(removeEmulationPreventionBytes(sps)); bitstream.skipBits(16); // NAL header bitstream.readBits(4); // sps_video_parameter_set_id const spsMaxSubLayersMinus1 = bitstream.readBits(3); const spsTemporalIdNestingFlag = bitstream.readBits(1); const { general_profile_space, general_tier_flag, general_profile_idc, general_profile_compatibility_flags, general_constraint_indicator_flags, general_level_idc, } = parseProfileTierLevel(bitstream, spsMaxSubLayersMinus1); readExpGolomb(bitstream); // sps_seq_parameter_set_id const chromaFormatIdc = readExpGolomb(bitstream); let separateColourPlaneFlag = 0; if (chromaFormatIdc === 3) { separateColourPlaneFlag = bitstream.readBits(1); } const picWidthInLumaSamples = readExpGolomb(bitstream); const picHeightInLumaSamples = readExpGolomb(bitstream); let displayWidth = picWidthInLumaSamples; let displayHeight = picHeightInLumaSamples; if (bitstream.readBits(1)) { // conformance_window_flag const confWinLeftOffset = readExpGolomb(bitstream); const confWinRightOffset = readExpGolomb(bitstream); const confWinTopOffset = readExpGolomb(bitstream); const confWinBottomOffset = readExpGolomb(bitstream); // SubWidthC and SubHeightC depend on chroma_format_idc and separate_colour_plane_flag let subWidthC = 1; let subHeightC = 1; const chromaArrayType = separateColourPlaneFlag === 0 ? chromaFormatIdc : 0; if (chromaArrayType === 1) { subWidthC = 2; subHeightC = 2; } else if (chromaArrayType === 2) { subWidthC = 2; subHeightC = 1; } displayWidth -= (confWinLeftOffset + confWinRightOffset) * subWidthC; displayHeight -= (confWinTopOffset + confWinBottomOffset) * subHeightC; } const bitDepthLumaMinus8 = readExpGolomb(bitstream); const bitDepthChromaMinus8 = readExpGolomb(bitstream); readExpGolomb(bitstream); // log2_max_pic_order_cnt_lsb_minus4 const spsSubLayerOrderingInfoPresentFlag = bitstream.readBits(1); const startI = spsSubLayerOrderingInfoPresentFlag ? 0 : spsMaxSubLayersMinus1; let spsMaxNumReorderPics = 0; for (let i = startI; i <= spsMaxSubLayersMinus1; i++) { readExpGolomb(bitstream); // sps_max_dec_pic_buffering_minus1[i] spsMaxNumReorderPics = readExpGolomb(bitstream); // sps_max_num_reorder_pics[i] readExpGolomb(bitstream); // sps_max_latency_increase_plus1[i] } readExpGolomb(bitstream); // log2_min_luma_coding_block_size_minus3 readExpGolomb(bitstream); // log2_diff_max_min_luma_coding_block_size readExpGolomb(bitstream); // log2_min_luma_transform_block_size_minus2 readExpGolomb(bitstream); // log2_diff_max_min_luma_transform_block_size readExpGolomb(bitstream); // max_transform_hierarchy_depth_inter readExpGolomb(bitstream); // max_transform_hierarchy_depth_intra if (bitstream.readBits(1)) { // scaling_list_enabled_flag if (bitstream.readBits(1)) { skipScalingListData(bitstream); } } bitstream.skipBits(1); // amp_enabled_flag bitstream.skipBits(1); // sample_adaptive_offset_enabled_flag if (bitstream.readBits(1)) { // pcm_enabled_flag bitstream.skipBits(4); // pcm_sample_bit_depth_luma_minus1 bitstream.skipBits(4); // pcm_sample_bit_depth_chroma_minus1 readExpGolomb(bitstream); // log2_min_pcm_luma_coding_block_size_minus3 readExpGolomb(bitstream); // log2_diff_max_min_pcm_luma_coding_block_size bitstream.skipBits(1); // pcm_loop_filter_disabled_flag } const numShortTermRefPicSets = readExpGolomb(bitstream); skipAllStRefPicSets(bitstream, numShortTermRefPicSets); if (bitstream.readBits(1)) { // long_term_ref_pics_present_flag const numLongTermRefPicsSps = readExpGolomb(bitstream); for (let i = 0; i < numLongTermRefPicsSps; i++) { readExpGolomb(bitstream); // lt_ref_pic_poc_lsb_sps[i] bitstream.skipBits(1); // used_by_curr_pic_lt_sps_flag[i] } } bitstream.skipBits(1); // sps_temporal_mvp_enabled_flag bitstream.skipBits(1); // strong_intra_smoothing_enabled_flag let colourPrimaries = 2; let transferCharacteristics = 2; let matrixCoefficients = 2; let fullRangeFlag = 0; let minSpatialSegmentationIdc = 0; let pixelAspectRatio: Rational = { num: 1, den: 1 }; if (bitstream.readBits(1)) { // vui_parameters_present_flag const vui = parseHevcVui(bitstream, spsMaxSubLayersMinus1); pixelAspectRatio = vui.pixelAspectRatio; colourPrimaries = vui.colourPrimaries; transferCharacteristics = vui.transferCharacteristics; matrixCoefficients = vui.matrixCoefficients; fullRangeFlag = vui.fullRangeFlag; minSpatialSegmentationIdc = vui.minSpatialSegmentationIdc; } return { displayWidth, displayHeight, pixelAspectRatio, colourPrimaries, transferCharacteristics, matrixCoefficients, fullRangeFlag, maxDecFrameBuffering: spsMaxNumReorderPics + 1, spsMaxSubLayersMinus1, spsTemporalIdNestingFlag, generalProfileSpace: general_profile_space, generalTierFlag: general_tier_flag, generalProfileIdc: general_profile_idc, generalProfileCompatibilityFlags: general_profile_compatibility_flags, generalConstraintIndicatorFlags: general_constraint_indicator_flags, generalLevelIdc: general_level_idc, chromaFormatIdc, bitDepthLumaMinus8, bitDepthChromaMinus8, minSpatialSegmentationIdc, }; } catch (error) { Logging._error('Error parsing HEVC SPS:', error); return null; } }; /** Builds a HevcDecoderConfigurationRecord from an HEVC packet in Annex B format. */ export const extractHevcDecoderConfigurationRecord = (packetData: Uint8Array) => { try { const vpsUnits: Uint8Array[] = []; const spsUnits: Uint8Array[] = []; const ppsUnits: Uint8Array[] = []; const seiUnits: Uint8Array[] = []; for (const loc of iterateNalUnitsInAnnexB(packetData)) { const nalUnit = packetData.subarray(loc.offset, loc.offset + loc.length); const type = extractNalUnitTypeForHevc(nalUnit[0]!); if (type === HevcNalUnitType.VPS_NUT) { vpsUnits.push(nalUnit); } else if (type === HevcNalUnitType.SPS_NUT) { spsUnits.push(nalUnit); } else if (type === HevcNalUnitType.PPS_NUT) { ppsUnits.push(nalUnit); } else if (type === HevcNalUnitType.PREFIX_SEI_NUT || type === HevcNalUnitType.SUFFIX_SEI_NUT) { seiUnits.push(nalUnit); } } if (spsUnits.length === 0 || ppsUnits.length === 0) return null; const spsInfo = parseHevcSps(spsUnits[0]!); if (!spsInfo) return null; // Parse PPS for parallelismType let parallelismType = 0; if (ppsUnits.length > 0) { const pps = ppsUnits[0]!; const ppsBitstream = new Bitstream(removeEmulationPreventionBytes(pps)); ppsBitstream.skipBits(16); // NAL header readExpGolomb(ppsBitstream); // pps_pic_parameter_set_id readExpGolomb(ppsBitstream); // pps_seq_parameter_set_id ppsBitstream.skipBits(1); // dependent_slice_segments_enabled_flag ppsBitstream.skipBits(1); // output_flag_present_flag ppsBitstream.skipBits(3); // num_extra_slice_header_bits ppsBitstream.skipBits(1); // sign_data_hiding_enabled_flag ppsBitstream.skipBits(1); // cabac_init_present_flag readExpGolomb(ppsBitstream); // num_ref_idx_l0_default_active_minus1 readExpGolomb(ppsBitstream); // num_ref_idx_l1_default_active_minus1 readSignedExpGolomb(ppsBitstream); // init_qp_minus26 ppsBitstream.skipBits(1); // constrained_intra_pred_flag ppsBitstream.skipBits(1); // transform_skip_enabled_flag if (ppsBitstream.readBits(1)) { // cu_qp_delta_enabled_flag readExpGolomb(ppsBitstream); // diff_cu_qp_delta_depth } readSignedExpGolomb(ppsBitstream); // pps_cb_qp_offset readSignedExpGolomb(ppsBitstream); // pps_cr_qp_offset ppsBitstream.skipBits(1); // pps_slice_chroma_qp_offsets_present_flag ppsBitstream.skipBits(1); // weighted_pred_flag ppsBitstream.skipBits(1); // weighted_bipred_flag ppsBitstream.skipBits(1); // transquant_bypass_enabled_flag const tiles_enabled_flag = ppsBitstream.readBits(1); const entropy_coding_sync_enabled_flag = ppsBitstream.readBits(1); if (!tiles_enabled_flag && !entropy_coding_sync_enabled_flag) parallelismType = 0; else if (tiles_enabled_flag && !entropy_coding_sync_enabled_flag) parallelismType = 2; else if (!tiles_enabled_flag && entropy_coding_sync_enabled_flag) parallelismType = 3; else parallelismType = 0; } const arrays = [ ...(vpsUnits.length ? [ { arrayCompleteness: 1, nalUnitType: HevcNalUnitType.VPS_NUT, nalUnits: vpsUnits, }, ] : []), ...(spsUnits.length ? [ { arrayCompleteness: 1, nalUnitType: HevcNalUnitType.SPS_NUT, nalUnits: spsUnits, }, ] : []), ...(ppsUnits.length ? [ { arrayCompleteness: 1, nalUnitType: HevcNalUnitType.PPS_NUT, nalUnits: ppsUnits, }, ] : []), ...(seiUnits.length ? [ { arrayCompleteness: 1, nalUnitType: extractNalUnitTypeForHevc(seiUnits[0]![0]!), nalUnits: seiUnits, }, ] : []), ]; const record: HevcDecoderConfigurationRecord = { configurationVersion: 1, generalProfileSpace: spsInfo.generalProfileSpace, generalTierFlag: spsInfo.generalTierFlag, generalProfileIdc: spsInfo.generalProfileIdc, generalProfileCompatibilityFlags: spsInfo.generalProfileCompatibilityFlags, generalConstraintIndicatorFlags: spsInfo.generalConstraintIndicatorFlags, generalLevelIdc: spsInfo.generalLevelIdc, minSpatialSegmentationIdc: spsInfo.minSpatialSegmentationIdc, parallelismType, chromaFormatIdc: spsInfo.chromaFormatIdc, bitDepthLumaMinus8: spsInfo.bitDepthLumaMinus8, bitDepthChromaMinus8: spsInfo.bitDepthChromaMinus8, avgFrameRate: 0, constantFrameRate: 0, numTemporalLayers: spsInfo.spsMaxSubLayersMinus1 + 1, temporalIdNested: spsInfo.spsTemporalIdNestingFlag, lengthSizeMinusOne: 3, arrays, }; return record; } catch (error) { Logging._error('Error building HEVC Decoder Configuration Record:', error); return null; } }; const parseProfileTierLevel = ( bitstream: Bitstream, maxNumSubLayersMinus1: number, ) => { const general_profile_space = bitstream.readBits(2); const general_tier_flag = bitstream.readBits(1); const general_profile_idc = bitstream.readBits(5); let general_profile_compatibility_flags = 0; for (let i = 0; i < 32; i++) { general_profile_compatibility_flags = (general_profile_compatibility_flags << 1) | bitstream.readBits(1); } const general_constraint_indicator_flags = new Uint8Array(6); for (let i = 0; i < 6; i++) { general_constraint_indicator_flags[i] = bitstream.readBits(8); } const general_level_idc = bitstream.readBits(8); const sub_layer_profile_present_flag: number[] = []; const sub_layer_level_present_flag: number[] = []; for (let i = 0; i < maxNumSubLayersMinus1; i++) { sub_layer_profile_present_flag.push(bitstream.readBits(1)); sub_layer_level_present_flag.push(bitstream.readBits(1)); } if (maxNumSubLayersMinus1 > 0) { for (let i = maxNumSubLayersMinus1; i < 8; i++) { bitstream.skipBits(2); // reserved_zero_2bits } } for (let i = 0; i < maxNumSubLayersMinus1; i++) { if (sub_layer_profile_present_flag[i]) bitstream.skipBits(88); if (sub_layer_level_present_flag[i]) bitstream.skipBits(8); } return { general_profile_space, general_tier_flag, general_profile_idc, general_profile_compatibility_flags, general_constraint_indicator_flags, general_level_idc, }; }; const skipScalingListData = (bitstream: Bitstream) => { for (let sizeId = 0; sizeId < 4; sizeId++) { for (let matrixId = 0; matrixId < (sizeId === 3 ? 2 : 6); matrixId++) { const scaling_list_pred_mode_flag = bitstream.readBits(1); if (!scaling_list_pred_mode_flag) { readExpGolomb(bitstream); // scaling_list_pred_matrix_id_delta } else { const coefNum = Math.min(64, 1 << (4 + (sizeId << 1))); if (sizeId > 1) { readSignedExpGolomb(bitstream); // scaling_list_dc_coef_minus8 } for (let i = 0; i < coefNum; i++) { readSignedExpGolomb(bitstream); // scaling_list_delta_coef } } } } }; const skipAllStRefPicSets = (bitstream: Bitstream, num_short_term_ref_pic_sets: number) => { const NumDeltaPocs: number[] = []; for (let stRpsIdx = 0; stRpsIdx < num_short_term_ref_pic_sets; stRpsIdx++) { NumDeltaPocs[stRpsIdx] = skipStRefPicSet(bitstream, stRpsIdx, num_short_term_ref_pic_sets, NumDeltaPocs); } }; const skipStRefPicSet = ( bitstream: Bitstream, stRpsIdx: number, num_short_term_ref_pic_sets: number, NumDeltaPocs: number[], ) => { let NumDeltaPocsThis = 0; let inter_ref_pic_set_prediction_flag = 0; let RefRpsIdx = 0; if (stRpsIdx !== 0) { inter_ref_pic_set_prediction_flag = bitstream.readBits(1); } if (inter_ref_pic_set_prediction_flag) { if (stRpsIdx === num_short_term_ref_pic_sets) { const delta_idx_minus1 = readExpGolomb(bitstream); RefRpsIdx = stRpsIdx - (delta_idx_minus1 + 1); } else { RefRpsIdx = stRpsIdx - 1; } bitstream.readBits(1); // delta_rps_sign readExpGolomb(bitstream); // abs_delta_rps_minus1 // The number of iterations is NumDeltaPocs[RefRpsIdx] + 1 const numDelta = NumDeltaPocs[RefRpsIdx] ?? 0; for (let j = 0; j <= numDelta; j++) { const used_by_curr_pic_flag = bitstream.readBits(1); if (!used_by_curr_pic_flag) { bitstream.readBits(1); // use_delta_flag } } NumDeltaPocsThis = NumDeltaPocs[RefRpsIdx]!; } else { const num_negative_pics = readExpGolomb(bitstream); const num_positive_pics = readExpGolomb(bitstream); for (let i = 0; i < num_negative_pics; i++) { readExpGolomb(bitstream); // delta_poc_s0_minus1[i] bitstream.readBits(1); // used_by_curr_pic_s0_flag[i] } for (let i = 0; i < num_positive_pics; i++) { readExpGolomb(bitstream); // delta_poc_s1_minus1[i] bitstream.readBits(1); // used_by_curr_pic_s1_flag[i] } NumDeltaPocsThis = num_negative_pics + num_positive_pics; } return NumDeltaPocsThis; }; const parseHevcVui = (bitstream: Bitstream, sps_max_sub_layers_minus1: number) => { // Defaults: 2 = unspecified let colourPrimaries = 2; let transferCharacteristics = 2; let matrixCoefficients = 2; let fullRangeFlag = 0; let minSpatialSegmentationIdc = 0; let pixelAspectRatio: Rational = { num: 1, den: 1 }; if (bitstream.readBits(1)) { // aspect_ratio_info_present_flag const aspect_ratio_idc = bitstream.readBits(8); if (aspect_ratio_idc === 255) { pixelAspectRatio = { num: bitstream.readBits(16), den: bitstream.readBits(16), }; } else { const aspectRatio = AVC_HEVC_ASPECT_RATIO_IDC_TABLE[aspect_ratio_idc]; if (aspectRatio) { pixelAspectRatio = aspectRatio; } } } if (bitstream.readBits(1)) { // overscan_info_present_flag bitstream.readBits(1); // overscan_appropriate_flag } if (bitstream.readBits(1)) { // video_signal_type_present_flag bitstream.readBits(3); // video_format fullRangeFlag = bitstream.readBits(1); if (bitstream.readBits(1)) { // colour_description_present_flag colourPrimaries = bitstream.readBits(8); transferCharacteristics = bitstream.readBits(8); matrixCoefficients = bitstream.readBits(8); } } if (bitstream.readBits(1)) { // chroma_loc_info_present_flag readExpGolomb(bitstream); // chroma_sample_loc_type_top_field readExpGolomb(bitstream); // chroma_sample_loc_type_bottom_field } bitstream.readBits(1); // neutral_chroma_indication_flag bitstream.readBits(1); // field_seq_flag bitstream.readBits(1); // frame_field_info_present_flag if (bitstream.readBits(1)) { // default_display_window_flag readExpGolomb(bitstream); // def_disp_win_left_offset readExpGolomb(bitstream); // def_disp_win_right_offset readExpGolomb(bitstream); // def_disp_win_top_offset readExpGolomb(bitstream); // def_disp_win_bottom_offset } if (bitstream.readBits(1)) { // vui_timing_info_present_flag bitstream.readBits(32); // vui_num_units_in_tick bitstream.readBits(32); // vui_time_scale if (bitstream.readBits(1)) { // vui_poc_proportional_to_timing_flag readExpGolomb(bitstream); // vui_num_ticks_poc_diff_one_minus1 } if (bitstream.readBits(1)) { skipHevcHrdParameters(bitstream, true, sps_max_sub_layers_minus1); } } if (bitstream.readBits(1)) { // bitstream_restriction_flag bitstream.readBits(1); // tiles_fixed_structure_flag bitstream.readBits(1); // motion_vectors_over_pic_boundaries_flag bitstream.readBits(1); // restricted_ref_pic_lists_flag minSpatialSegmentationIdc = readExpGolomb(bitstream); readExpGolomb(bitstream); // max_bytes_per_pic_denom readExpGolomb(bitstream); // max_bits_per_min_cu_denom readExpGolomb(bitstream); // log2_max_mv_length_horizontal readExpGolomb(bitstream); // log2_max_mv_length_vertical } return { pixelAspectRatio, colourPrimaries, transferCharacteristics, matrixCoefficients, fullRangeFlag, minSpatialSegmentationIdc, }; }; const skipHevcHrdParameters = ( bitstream: Bitstream, commonInfPresentFlag: boolean, maxNumSubLayersMinus1: number, ) => { let nal_hrd_parameters_present_flag = false; let vcl_hrd_parameters_present_flag = false; let sub_pic_hrd_params_present_flag = false; if (commonInfPresentFlag) { nal_hrd_parameters_present_flag = bitstream.readBits(1) === 1; vcl_hrd_parameters_present_flag = bitstream.readBits(1) === 1; if (nal_hrd_parameters_present_flag || vcl_hrd_parameters_present_flag) { sub_pic_hrd_params_present_flag = bitstream.readBits(1) === 1; if (sub_pic_hrd_params_present_flag) { bitstream.readBits(8); // tick_divisor_minus2 bitstream.readBits(5); // du_cpb_removal_delay_increment_length_minus1 bitstream.readBits(1); // sub_pic_cpb_params_in_pic_timing_sei_flag bitstream.readBits(5); // dpb_output_delay_du_length_minus1 } bitstream.readBits(4); // bit_rate_scale bitstream.readBits(4); // cpb_size_scale if (sub_pic_hrd_params_present_flag) { bitstream.readBits(4); // cpb_size_du_scale } bitstream.readBits(5); // initial_cpb_removal_delay_length_minus1 bitstream.readBits(5); // au_cpb_removal_delay_length_minus1 bitstream.readBits(5); // dpb_output_delay_length_minus1 } } for (let i = 0; i <= maxNumSubLayersMinus1; i++) { const fixed_pic_rate_general_flag = bitstream.readBits(1) === 1; let fixed_pic_rate_within_cvs_flag = true; // Default assumption if general is true if (!fixed_pic_rate_general_flag) { fixed_pic_rate_within_cvs_flag = bitstream.readBits(1) === 1; } let low_delay_hrd_flag = false; // Default assumption if (fixed_pic_rate_within_cvs_flag) { readExpGolomb(bitstream); // elemental_duration_in_tc_minus1[i] } else { low_delay_hrd_flag = bitstream.readBits(1) === 1; } let CpbCnt = 1; // Default if low_delay is true if (!low_delay_hrd_flag) { const cpb_cnt_minus1 = readExpGolomb(bitstream); // cpb_cnt_minus1[i] CpbCnt = cpb_cnt_minus1 + 1; } if (nal_hrd_parameters_present_flag) { skipSubLayerHrdParameters(bitstream, CpbCnt, sub_pic_hrd_params_present_flag); } if (vcl_hrd_parameters_present_flag) { skipSubLayerHrdParameters(bitstream, CpbCnt, sub_pic_hrd_params_present_flag); } } }; const skipSubLayerHrdParameters = ( bitstream: Bitstream, CpbCnt: number, sub_pic_hrd_params_present_flag: boolean, ) => { for (let i = 0; i < CpbCnt; i++) { readExpGolomb(bitstream); // bit_rate_value_minus1[i] readExpGolomb(bitstream); // cpb_size_value_minus1[i] if (sub_pic_hrd_params_present_flag) { readExpGolomb(bitstream); // cpb_size_du_value_minus1[i] readExpGolomb(bitstream); // bit_rate_du_value_minus1[i] } bitstream.readBits(1); // cbr_flag[i] } }; /** Serializes an HevcDecoderConfigurationRecord into the format specified in Section 8.3.3.1 of ISO 14496-15. */ export const serializeHevcDecoderConfigurationRecord = (record: HevcDecoderConfigurationRecord) => { const bytes: number[] = []; bytes.push(record.configurationVersion); bytes.push( ((record.generalProfileSpace & 0x3) << 6) | ((record.generalTierFlag & 0x1) << 5) | (record.generalProfileIdc & 0x1F), ); bytes.push((record.generalProfileCompatibilityFlags >>> 24) & 0xFF); bytes.push((record.generalProfileCompatibilityFlags >>> 16) & 0xFF); bytes.push((record.generalProfileCompatibilityFlags >>> 8) & 0xFF); bytes.push(record.generalProfileCompatibilityFlags & 0xFF); bytes.push(...record.generalConstraintIndicatorFlags); bytes.push(record.generalLevelIdc & 0xFF); bytes.push(0xF0 | ((record.minSpatialSegmentationIdc >> 8) & 0x0F)); // Reserved + high nibble bytes.push(record.minSpatialSegmentationIdc & 0xFF); // Low byte bytes.push(0xFC | (record.parallelismType & 0x03)); bytes.push(0xFC | (record.chromaFormatIdc & 0x03)); bytes.push(0xF8 | (record.bitDepthLumaMinus8 & 0x07)); bytes.push(0xF8 | (record.bitDepthChromaMinus8 & 0x07)); bytes.push((record.avgFrameRate >> 8) & 0xFF); // High byte bytes.push(record.avgFrameRate & 0xFF); // Low byte bytes.push( ((record.constantFrameRate & 0x03) << 6) | ((record.numTemporalLayers & 0x07) << 3) | ((record.temporalIdNested & 0x01) << 2) | (record.lengthSizeMinusOne & 0x03), ); bytes.push(record.arrays.length & 0xFF); for (const arr of record.arrays) { bytes.push( ((arr.arrayCompleteness & 0x01) << 7) | (0 << 6) | (arr.nalUnitType & 0x3F), ); bytes.push((arr.nalUnits.length >> 8) & 0xFF); // High byte bytes.push(arr.nalUnits.length & 0xFF); // Low byte for (const nal of arr.nalUnits) { bytes.push((nal.length >> 8) & 0xFF); // High byte bytes.push(nal.length & 0xFF); // Low byte for (let i = 0; i < nal.length; i++) { bytes.push(nal[i]!); } } } return new Uint8Array(bytes); }; /** Deserializes an HevcDecoderConfigurationRecord from the format specified in Section 8.3.3.1 of ISO 14496-15. */ export const deserializeHevcDecoderConfigurationRecord = (data: Uint8Array): HevcDecoderConfigurationRecord | null => { try { const view = toDataView(data); let offset = 0; const configurationVersion = view.getUint8(offset++); const byte1 = view.getUint8(offset++); const generalProfileSpace = (byte1 >> 6) & 0x3; const generalTierFlag = (byte1 >> 5) & 0x1; const generalProfileIdc = byte1 & 0x1F; const generalProfileCompatibilityFlags = view.getUint32(offset, false); offset += 4; const generalConstraintIndicatorFlags = data.subarray(offset, offset + 6); offset += 6; const generalLevelIdc = view.getUint8(offset++); const minSpatialSegmentationIdc = ((view.getUint8(offset++) & 0x0F) << 8) | view.getUint8(offset++); const parallelismType = view.getUint8(offset++) & 0x03; const chromaFormatIdc = view.getUint8(offset++) & 0x03; const bitDepthLumaMinus8 = view.getUint8(offset++) & 0x07; const bitDepthChromaMinus8 = view.getUint8(offset++) & 0x07; const avgFrameRate = view.getUint16(offset, false); offset += 2; const byte21 = view.getUint8(offset++); const constantFrameRate = (byte21 >> 6) & 0x03; const numTemporalLayers = (byte21 >> 3) & 0x07; const temporalIdNested = (byte21 >> 2) & 0x01; const lengthSizeMinusOne = byte21 & 0x03; const numOfArrays = view.getUint8(offset++); const arrays: HevcDecoderConfigurationRecord['arrays'] = []; for (let i = 0; i < numOfArrays; i++) { const arrByte = view.getUint8(offset++); const arrayCompleteness = (arrByte >> 7) & 0x01; const nalUnitType = arrByte & 0x3F; const numNalus = view.getUint16(offset, false); offset += 2; const nalUnits: Uint8Array[] = []; for (let j = 0; j < numNalus; j++) { const nalUnitLength = view.getUint16(offset, false); offset += 2; nalUnits.push(data.subarray(offset, offset + nalUnitLength)); offset += nalUnitLength; } arrays.push({ arrayCompleteness, nalUnitType, nalUnits, }); } return { configurationVersion, generalProfileSpace, generalTierFlag, generalProfileIdc, generalProfileCompatibilityFlags, generalConstraintIndicatorFlags, generalLevelIdc, minSpatialSegmentationIdc, parallelismType, chromaFormatIdc, bitDepthLumaMinus8, bitDepthChromaMinus8, avgFrameRate, constantFrameRate, numTemporalLayers, temporalIdNested, lengthSizeMinusOne, arrays, }; } catch (error) { Logging._error('Error deserializing HEVC Decoder Configuration Record:', error); return null; } }; enum HevcNaluOrderState { audAllowed, beforeFirstVcl, afterFirstVcl, eoBitstreamAllowed, noMoreDataAllowed, } // This function sanitzes the contents of an HEVC packet such that // https://source.chromium.org/chromium/chromium/src/+/main:media/formats/mp4/hevc.cc's validation logic does not trip // up on its contents. The validation is often too strict and rejects packets that Chromium could decode just fine. // Chromium code retrieved on 2026-04-29. // See https://issues.chromium.org/issues/507611247. export const sanitizeHevcPacketForChromium = ( packetData: Uint8Array, decoderConfig: VideoDecoderConfig, ): Uint8Array | null => { const removedNalUnits = new Set(); let orderState: HevcNaluOrderState = HevcNaluOrderState.audAllowed; for (const loc of iterateHevcNalUnits(packetData, decoderConfig)) { if (orderState === HevcNaluOrderState.noMoreDataAllowed) { removedNalUnits.add(loc.offset); continue; } const type = extractNalUnitTypeForHevc(packetData[loc.offset]!); if (orderState === HevcNaluOrderState.eoBitstreamAllowed && type !== 37 /* EOB_NUT */) { removedNalUnits.add(loc.offset); continue; } let remove = false; if (type === 35) { // AUD_NUT if (orderState > HevcNaluOrderState.audAllowed) { remove = true; } else { orderState = HevcNaluOrderState.beforeFirstVcl; } } else if (type <= 31) { // VCL (0-31) if (orderState > HevcNaluOrderState.afterFirstVcl) { remove = true; } else { orderState = HevcNaluOrderState.afterFirstVcl; } } else if (type === 36) { // EOS_NUT if (orderState !== HevcNaluOrderState.afterFirstVcl) { remove = true; } else { orderState = HevcNaluOrderState.eoBitstreamAllowed; } } else if (type === 37) { // EOB_NUT if (orderState < HevcNaluOrderState.afterFirstVcl) { remove = true; } else { orderState = HevcNaluOrderState.noMoreDataAllowed; } } else if ( type === 32 || type === 33 || type === 34 || type === 39 || (type >= 41 && type <= 44) || (type >= 48 && type <= 55) ) { // VPS, SPS, PPS, PREFIX_SEI, RSV_NVCL41..44, UNSPEC48..55 if (orderState > HevcNaluOrderState.beforeFirstVcl) { remove = true; } else { orderState = HevcNaluOrderState.beforeFirstVcl; } } else if ( type === 38 || type === 40 || (type >= 45 && type <= 47) || (type >= 56 && type <= 63) ) { // FD, SUFFIX_SEI, RSV_NVCL45..47, UNSPEC56..63 if (orderState < HevcNaluOrderState.afterFirstVcl) { remove = true; } } if (remove) { removedNalUnits.add(loc.offset); } } // If nothing violated the rules, return null to signal that if (removedNalUnits.size === 0) { return null; } const filteredNalUnits: Uint8Array[] = []; for (const loc of iterateHevcNalUnits(packetData, decoderConfig)) { if (!removedNalUnits.has(loc.offset)) { filteredNalUnits.push(packetData.subarray(loc.offset, loc.offset + loc.length)); } } return concatHevcNalUnits(filteredNalUnits, decoderConfig); }; export type Vp9CodecInfo = { profile: number; level: number; bitDepth: number; chromaSubsampling: number; videoFullRangeFlag: number; colourPrimaries: number; transferCharacteristics: number; matrixCoefficients: number; }; export const extractVp9CodecInfoFromPacket = ( packet: Uint8Array, ): Vp9CodecInfo | null => { // eslint-disable-next-line @stylistic/max-len // https://storage.googleapis.com/downloads.webmproject.org/docs/vp9/vp9-bitstream-specification-v0.7-20170222-draft.pdf // http://downloads.webmproject.org/docs/vp9/vp9-bitstream_superframe-and-uncompressed-header_v1.0.pdf const bitstream = new Bitstream(packet); // Frame marker (0b10) const frameMarker = bitstream.readBits(2); if (frameMarker !== 2) { return null; } // Profile const profileLowBit = bitstream.readBits(1); const profileHighBit = bitstream.readBits(1); const profile = (profileHighBit << 1) + profileLowBit; // Skip reserved bit for profile 3 if (profile === 3) { bitstream.skipBits(1); } // show_existing_frame const showExistingFrame = bitstream.readBits(1); if (showExistingFrame === 1) { return null; } // frame_type (0 = key frame) const frameType = bitstream.readBits(1); if (frameType !== 0) { return null; } // Skip show_frame and error_resilient_mode bitstream.skipBits(2); // Sync code (0x498342) const syncCode = bitstream.readBits(24); if (syncCode !== 0x498342) { return null; } // Color config let bitDepth = 8; if (profile >= 2) { const tenOrTwelveBit = bitstream.readBits(1); bitDepth = tenOrTwelveBit ? 12 : 10; } // Color space const colorSpace = bitstream.readBits(3); let chromaSubsampling = 0; let videoFullRangeFlag = 0; if (colorSpace !== 7) { // 7 is CS_RGB const colorRange = bitstream.readBits(1); videoFullRangeFlag = colorRange; if (profile === 1 || profile === 3) { const subsamplingX = bitstream.readBits(1); const subsamplingY = bitstream.readBits(1); // 0 = 4:2:0 vertical // 1 = 4:2:0 colocated // 2 = 4:2:2 // 3 = 4:4:4 chromaSubsampling = !subsamplingX && !subsamplingY ? 3 // 0,0 = 4:4:4 : subsamplingX && !subsamplingY ? 2 // 1,0 = 4:2:2 : 1; // 1,1 = 4:2:0 colocated (default) // Skip reserved bit bitstream.skipBits(1); } else { // For profile 0 and 2, always 4:2:0 chromaSubsampling = 1; // Using colocated as default } } else { // RGB is always 4:4:4 chromaSubsampling = 3; videoFullRangeFlag = 1; } // Parse frame size const widthMinusOne = bitstream.readBits(16); const heightMinusOne = bitstream.readBits(16); const width = widthMinusOne + 1; const height = heightMinusOne + 1; // Calculate level based on dimensions const pictureSize = width * height; let level = last(VP9_LEVEL_TABLE)!.level; // Default to highest level for (const entry of VP9_LEVEL_TABLE) { if (pictureSize <= entry.maxPictureSize) { level = entry.level; break; } } // Map color_space to standard values const matrixCoefficients = colorSpace === 7 ? 0 : colorSpace === 2 ? 1 : colorSpace === 1 ? 6 : 2; const colourPrimaries = colorSpace === 2 ? 1 : colorSpace === 1 ? 6 : 2; const transferCharacteristics = colorSpace === 2 ? 1 : colorSpace === 1 ? 6 : 2; return { profile, level, bitDepth, chromaSubsampling, videoFullRangeFlag, colourPrimaries, transferCharacteristics, matrixCoefficients, }; }; export type Av1CodecInfo = { profile: number; level: number; tier: number; bitDepth: number; monochrome: number; chromaSubsamplingX: number; chromaSubsamplingY: number; chromaSamplePosition: number; }; /** Iterates over all OBUs in an AV1 packet bistream. */ export const iterateAv1PacketObus = function* (packet: Uint8Array) { // https://aomediacodec.github.io/av1-spec/av1-spec.pdf const bitstream = new Bitstream(packet); const readLeb128 = (): number | null => { let value = 0; for (let i = 0; i < 8; i++) { const byte = bitstream.readAlignedByte(); value |= ((byte & 0x7f) << (i * 7)); if (!(byte & 0x80)) { break; } // Spec requirement if (i === 7 && (byte & 0x80)) { return null; } } // Spec requirement if (value >= 2 ** 32 - 1) { return null; } return value; }; while (bitstream.getBitsLeft() >= 8) { // Parse OBU header bitstream.skipBits(1); const obuType = bitstream.readBits(4); const obuExtension = bitstream.readBits(1); const obuHasSizeField = bitstream.readBits(1); bitstream.skipBits(1); // Skip extension header if present if (obuExtension) { bitstream.skipBits(8); } // Read OBU size if present let obuSize: number; if (obuHasSizeField) { const obuSizeValue = readLeb128(); if (obuSizeValue === null) return; // It was invalid obuSize = obuSizeValue; } else { // Calculate remaining bits and convert to bytes, rounding down obuSize = Math.floor(bitstream.getBitsLeft() / 8); } assert(bitstream.pos % 8 === 0); yield { type: obuType, data: packet.subarray(bitstream.pos / 8, bitstream.pos / 8 + obuSize), }; // Move to next OBU bitstream.skipBits(obuSize * 8); } }; /** * When AV1 codec information is not provided by the container, we can still try to extract the information by digging * into the AV1 bitstream. */ export const extractAv1CodecInfoFromPacket = ( packet: Uint8Array, ): Av1CodecInfo | null => { // https://aomediacodec.github.io/av1-spec/av1-spec.pdf for (const { type, data } of iterateAv1PacketObus(packet)) { if (type !== 1) { continue; // 1 == OBU_SEQUENCE_HEADER } const bitstream = new Bitstream(data); // Read sequence header fields const seqProfile = bitstream.readBits(3); // eslint-disable-next-line @typescript-eslint/no-unused-vars const stillPicture = bitstream.readBits(1); const reducedStillPictureHeader = bitstream.readBits(1); let seqLevel = 0; let seqTier = 0; let bufferDelayLengthMinus1 = 0; if (reducedStillPictureHeader) { seqLevel = bitstream.readBits(5); } else { // Parse timing_info_present_flag const timingInfoPresentFlag = bitstream.readBits(1); if (timingInfoPresentFlag) { // Skip timing info (num_units_in_display_tick, time_scale, equal_picture_interval) bitstream.skipBits(32); // num_units_in_display_tick bitstream.skipBits(32); // time_scale const equalPictureInterval = bitstream.readBits(1); if (equalPictureInterval) { // Skip num_ticks_per_picture_minus_1 (uvlc) // Since this is variable length, we'd need to implement uvlc reading // For now, we'll return null as this is rare return null; } } // Parse decoder_model_info_present_flag const decoderModelInfoPresentFlag = bitstream.readBits(1); if (decoderModelInfoPresentFlag) { // Store buffer_delay_length_minus_1 instead of just skipping bufferDelayLengthMinus1 = bitstream.readBits(5); bitstream.skipBits(32); // num_units_in_decoding_tick bitstream.skipBits(5); // buffer_removal_time_length_minus_1 bitstream.skipBits(5); // frame_presentation_time_length_minus_1 } // Parse operating_points_cnt_minus_1 const operatingPointsCntMinus1 = bitstream.readBits(5); // For each operating point for (let i = 0; i <= operatingPointsCntMinus1; i++) { // operating_point_idc[i] bitstream.skipBits(12); // seq_level_idx[i] const seqLevelIdx = bitstream.readBits(5); if (i === 0) { seqLevel = seqLevelIdx; } if (seqLevelIdx > 7) { // seq_tier[i] const seqTierTemp = bitstream.readBits(1); if (i === 0) { seqTier = seqTierTemp; } } if (decoderModelInfoPresentFlag) { // decoder_model_present_for_this_op[i] const decoderModelPresentForThisOp = bitstream.readBits(1); if (decoderModelPresentForThisOp) { const n = bufferDelayLengthMinus1 + 1; bitstream.skipBits(n); // decoder_buffer_delay[op] bitstream.skipBits(n); // encoder_buffer_delay[op] bitstream.skipBits(1); // low_delay_mode_flag[op] } } // initial_display_delay_present_flag const initialDisplayDelayPresentFlag = bitstream.readBits(1); if (initialDisplayDelayPresentFlag) { // initial_display_delay_minus_1[i] bitstream.skipBits(4); } } } // Frame size const frameWidthBitsMinus1 = bitstream.readBits(4); const frameHeightBitsMinus1 = bitstream.readBits(4); const n1 = frameWidthBitsMinus1 + 1; bitstream.skipBits(n1); // max_frame_width_minus_1 const n2 = frameHeightBitsMinus1 + 1; bitstream.skipBits(n2); // max_frame_height_minus_1 // Frame IDs let frameIdNumbersPresentFlag = 0; if (reducedStillPictureHeader) { frameIdNumbersPresentFlag = 0; } else { frameIdNumbersPresentFlag = bitstream.readBits(1); } if (frameIdNumbersPresentFlag) { bitstream.skipBits(4); // delta_frame_id_length_minus_2 bitstream.skipBits(3); // additional_frame_id_length_minus_1 } bitstream.skipBits(1); // use_128x128_superblock bitstream.skipBits(1); // enable_filter_intra bitstream.skipBits(1); // enable_intra_edge_filter if (!reducedStillPictureHeader) { bitstream.skipBits(1); // enable_interintra_compound bitstream.skipBits(1); // enable_masked_compound bitstream.skipBits(1); // enable_warped_motion bitstream.skipBits(1); // enable_dual_filter const enableOrderHint = bitstream.readBits(1); if (enableOrderHint) { bitstream.skipBits(1); // enable_jnt_comp bitstream.skipBits(1); // enable_ref_frame_mvs } const seqChooseScreenContentTools = bitstream.readBits(1); let seqForceScreenContentTools = 0; if (seqChooseScreenContentTools) { seqForceScreenContentTools = 2; // SELECT_SCREEN_CONTENT_TOOLS } else { seqForceScreenContentTools = bitstream.readBits(1); } if (seqForceScreenContentTools > 0) { const seqChooseIntegerMv = bitstream.readBits(1); if (!seqChooseIntegerMv) { bitstream.skipBits(1); // seq_force_integer_mv } } if (enableOrderHint) { bitstream.skipBits(3); // order_hint_bits_minus_1 } } bitstream.skipBits(1); // enable_superres bitstream.skipBits(1); // enable_cdef bitstream.skipBits(1); // enable_restoration // color_config() const highBitdepth = bitstream.readBits(1); let bitDepth = 8; if (seqProfile === 2 && highBitdepth) { const twelveBit = bitstream.readBits(1); bitDepth = twelveBit ? 12 : 10; } else if (seqProfile <= 2) { bitDepth = highBitdepth ? 10 : 8; } let monochrome = 0; if (seqProfile !== 1) { monochrome = bitstream.readBits(1); } let chromaSubsamplingX = 1; let chromaSubsamplingY = 1; let chromaSamplePosition = 0; if (!monochrome) { if (seqProfile === 0) { chromaSubsamplingX = 1; chromaSubsamplingY = 1; } else if (seqProfile === 1) { chromaSubsamplingX = 0; chromaSubsamplingY = 0; } else { if (bitDepth === 12) { chromaSubsamplingX = bitstream.readBits(1); if (chromaSubsamplingX) { chromaSubsamplingY = bitstream.readBits(1); } } } if (chromaSubsamplingX && chromaSubsamplingY) { chromaSamplePosition = bitstream.readBits(2); } } return { profile: seqProfile, level: seqLevel, tier: seqTier, bitDepth, monochrome, chromaSubsamplingX, chromaSubsamplingY, chromaSamplePosition, }; } return null; }; export const parseOpusIdentificationHeader = (bytes: Uint8Array) => { const view = toDataView(bytes); const outputChannelCount = view.getUint8(9); const preSkip = view.getUint16(10, true); const inputSampleRate = view.getUint32(12, true); const outputGain = view.getInt16(16, true); const channelMappingFamily = view.getUint8(18); let channelMappingTable: Uint8Array | null = null; if (channelMappingFamily) { channelMappingTable = bytes.subarray(19, 19 + 2 + outputChannelCount); } return { outputChannelCount, preSkip, inputSampleRate, outputGain, channelMappingFamily, channelMappingTable, }; }; // From https://datatracker.ietf.org/doc/html/rfc6716, in 48 kHz samples const OPUS_FRAME_DURATION_TABLE = [ 480, 960, 1920, 2880, 480, 960, 1920, 2880, 480, 960, 1920, 2880, 480, 960, 480, 960, 120, 240, 480, 960, 120, 240, 480, 960, 120, 240, 480, 960, 120, 240, 480, 960, ]; export const parseOpusTocByte = (packet: Uint8Array) => { const config = packet[0]! >> 3; return { durationInSamples: OPUS_FRAME_DURATION_TABLE[config]!, }; }; // Based on vorbis_parser.c from FFmpeg. export const parseModesFromVorbisSetupPacket = (setupHeader: Uint8Array) => { // Verify that this is a Setup header. if (setupHeader.length < 7) { throw new Error('Setup header is too short.'); } if (setupHeader[0] !== 5) { throw new Error('Wrong packet type in Setup header.'); } const signature = String.fromCharCode(...setupHeader.slice(1, 7)); if (signature !== 'vorbis') { throw new Error('Invalid packet signature in Setup header.'); } // Reverse the entire buffer. const bufSize = setupHeader.length; const revBuffer = new Uint8Array(bufSize); for (let i = 0; i < bufSize; i++) { revBuffer[i] = setupHeader[bufSize - 1 - i]!; } // Initialize a Bitstream on the reversed buffer. const bitstream = new Bitstream(revBuffer); // --- Find the framing bit. // In FFmpeg code, we scan until get_bits1() returns 1. let gotFramingBit = 0; while (bitstream.getBitsLeft() > 97) { if (bitstream.readBits(1) === 1) { gotFramingBit = bitstream.pos; break; } } if (gotFramingBit === 0) { throw new Error('Invalid Setup header: framing bit not found.'); } // --- Search backwards for a valid mode header. // We try to “guess” the number of modes by reading a fixed pattern. let modeCount = 0; let gotModeHeader = false; let lastModeCount = 0; while (bitstream.getBitsLeft() >= 97) { const tempPos = bitstream.pos; const a = bitstream.readBits(8); const b = bitstream.readBits(16); const c = bitstream.readBits(16); // If a > 63 or b or c nonzero, assume we’ve gone too far. if (a > 63 || b !== 0 || c !== 0) { bitstream.pos = tempPos; break; } bitstream.skipBits(1); modeCount++; if (modeCount > 64) { break; } const bsClone = bitstream.clone(); const candidate = bsClone.readBits(6) + 1; if (candidate === modeCount) { gotModeHeader = true; lastModeCount = modeCount; } } if (!gotModeHeader) { throw new Error('Invalid Setup header: mode header not found.'); } if (lastModeCount > 63) { throw new Error(`Unsupported mode count: ${lastModeCount}.`); } const finalModeCount = lastModeCount; // --- Reinitialize the bitstream. bitstream.pos = 0; // Skip the bits up to the found framing bit. bitstream.skipBits(gotFramingBit); // --- Now read, for each mode (in reverse order), 40 bits then one bit. // That one bit is the mode blockflag. const modeBlockflags = Array(finalModeCount).fill(0) as number[]; for (let i = finalModeCount - 1; i >= 0; i--) { bitstream.skipBits(40); modeBlockflags[i] = bitstream.readBits(1); } return { modeBlockflags }; }; /** Determines a packet's type (key or delta) by digging into the packet bitstream. */ export const determineVideoPacketType = ( codec: VideoCodec, decoderConfig: VideoDecoderConfig, packetData: Uint8Array, ): PacketType | null => { switch (codec) { case 'avc': { for (const loc of iterateAvcNalUnits(packetData, decoderConfig)) { const nalTypeByte = packetData[loc.offset]!; const type = extractNalUnitTypeForAvc(nalTypeByte); if (type >= AvcNalUnitType.NON_IDR_SLICE && type <= AvcNalUnitType.SLICE_DPC) { return 'delta'; } if (type === AvcNalUnitType.IDR) { return 'key'; } // In addition to IDR, Recovery Point SEI also counts as a valid H.264 keyframe by current consensus. // See https://github.com/w3c/webcodecs/issues/650 for the relevant discussion. WebKit and Firefox have // always supported them, but Chromium hasn't, therefore the (admittedly dirty) version check. if (type === AvcNalUnitType.SEI && (!isChromium() || getChromiumVersion()! >= 144)) { const nalUnit = packetData.subarray(loc.offset, loc.offset + loc.length); const bytes = removeEmulationPreventionBytes(nalUnit); let pos = 1; // Skip NALU header // sei_rbsp() do { // sei_message() let payloadType = 0; while (true) { const nextByte = bytes[pos++]; if (nextByte === undefined) break; payloadType += nextByte; if (nextByte < 255) { break; } } let payloadSize = 0; while (true) { const nextByte = bytes[pos++]; if (nextByte === undefined) break; payloadSize += nextByte; if (nextByte < 255) { break; } } // sei_payload() const PAYLOAD_TYPE_RECOVERY_POINT = 6; if (payloadType === PAYLOAD_TYPE_RECOVERY_POINT) { const bitstream = new Bitstream(bytes); bitstream.pos = 8 * pos; const recoveryFrameCount = readExpGolomb(bitstream); const exactMatchFlag = bitstream.readBits(1); if (recoveryFrameCount === 0 && exactMatchFlag === 1) { // https://github.com/w3c/webcodecs/pull/910 // "recovery_frame_cnt == 0 and exact_match_flag=1 in the SEI recovery payload" return 'key'; } } pos += payloadSize; } while (pos < bytes.length - 1); } } return 'delta'; }; case 'hevc': { for (const loc of iterateHevcNalUnits(packetData, decoderConfig)) { const type = extractNalUnitTypeForHevc(packetData[loc.offset]!); if (type < HevcNalUnitType.BLA_W_LP) { return 'delta'; } if (type <= HevcNalUnitType.RSV_IRAP_VCL23) { return 'key'; } } return 'delta'; }; case 'vp8': { // VP8, once again, by far the easiest to deal with. const frameType = packetData[0]! & 0b1; return frameType === 0 ? 'key' : 'delta'; }; case 'vp9': { const bitstream = new Bitstream(packetData); if (bitstream.readBits(2) !== 2) { return null; }; const profileLowBit = bitstream.readBits(1); const profileHighBit = bitstream.readBits(1); const profile = (profileHighBit << 1) + profileLowBit; // Skip reserved bit for profile 3 if (profile === 3) { bitstream.skipBits(1); } const showExistingFrame = bitstream.readBits(1); if (showExistingFrame) { return null; } const frameType = bitstream.readBits(1); return frameType === 0 ? 'key' : 'delta'; }; case 'av1': { let reducedStillPictureHeader = false; for (const { type, data } of iterateAv1PacketObus(packetData)) { if (type === 1) { // OBU_SEQUENCE_HEADER const bitstream = new Bitstream(data); bitstream.skipBits(4); reducedStillPictureHeader = !!bitstream.readBits(1); } else if ( type === 3 // OBU_FRAME_HEADER || type === 6 // OBU_FRAME || type === 7 // OBU_REDUNDANT_FRAME_HEADER ) { if (reducedStillPictureHeader) { return 'key'; } const bitstream = new Bitstream(data); const showExistingFrame = bitstream.readBits(1); if (showExistingFrame) { return null; } const frameType = bitstream.readBits(2); return frameType === 0 ? 'key' : 'delta'; } } return null; }; default: { assertNever(codec); assert(false); }; } }; export enum FlacBlockType { STREAMINFO = 0, VORBIS_COMMENT = 4, PICTURE = 6, } export const readVorbisComments = (bytes: Uint8Array, metadataTags: MetadataTags) => { // https://datatracker.ietf.org/doc/html/rfc7845#section-5.2 const commentView = toDataView(bytes); let commentPos = 0; const vendorStringLength = commentView.getUint32(commentPos, true); commentPos += 4; const vendorString = textDecoder.decode( bytes.subarray(commentPos, commentPos + vendorStringLength), ); commentPos += vendorStringLength; if (vendorStringLength > 0) { // Expose the vendor string in the raw metadata metadataTags.raw ??= {}; metadataTags.raw['vendor'] ??= vendorString; } const listLength = commentView.getUint32(commentPos, true); commentPos += 4; // Loop over all metadata tags for (let i = 0; i < listLength; i++) { const stringLength = commentView.getUint32(commentPos, true); commentPos += 4; const string = textDecoder.decode( bytes.subarray(commentPos, commentPos + stringLength), ); commentPos += stringLength; const separatorIndex = string.indexOf('='); if (separatorIndex === -1) { continue; } const key = string.slice(0, separatorIndex).toUpperCase(); const value = string.slice(separatorIndex + 1); metadataTags.raw ??= {}; metadataTags.raw[key] ??= value; switch (key) { case 'TITLE': { metadataTags.title ??= value; }; break; case 'DESCRIPTION': { metadataTags.description ??= value; }; break; case 'ARTIST': { metadataTags.artist ??= value; }; break; case 'ALBUM': { metadataTags.album ??= value; }; break; case 'ALBUMARTIST': { metadataTags.albumArtist ??= value; }; break; case 'COMMENT': { metadataTags.comment ??= value; }; break; case 'LYRICS': { metadataTags.lyrics ??= value; }; break; case 'TRACKNUMBER': { const parts = value.split('/'); const trackNum = Number.parseInt(parts[0]!, 10); const tracksTotal = parts[1] && Number.parseInt(parts[1], 10); if (Number.isInteger(trackNum) && trackNum > 0) { metadataTags.trackNumber ??= trackNum; } if (tracksTotal && Number.isInteger(tracksTotal) && tracksTotal > 0) { metadataTags.tracksTotal ??= tracksTotal; } }; break; case 'TRACKTOTAL': { const tracksTotal = Number.parseInt(value, 10); if (Number.isInteger(tracksTotal) && tracksTotal > 0) { metadataTags.tracksTotal ??= tracksTotal; } }; break; case 'DISCNUMBER': { const parts = value.split('/'); const discNum = Number.parseInt(parts[0]!, 10); const discsTotal = parts[1] && Number.parseInt(parts[1], 10); if (Number.isInteger(discNum) && discNum > 0) { metadataTags.discNumber ??= discNum; } if (discsTotal && Number.isInteger(discsTotal) && discsTotal > 0) { metadataTags.discsTotal ??= discsTotal; } }; break; case 'DISCTOTAL': { const discsTotal = Number.parseInt(value, 10); if (Number.isInteger(discsTotal) && discsTotal > 0) { metadataTags.discsTotal ??= discsTotal; } }; break; case 'DATE': { const date = new Date(value); if (!Number.isNaN(date.getTime())) { metadataTags.date ??= date; } }; break; case 'GENRE': { metadataTags.genre ??= value; }; break; case 'METADATA_BLOCK_PICTURE': { // https://datatracker.ietf.org/doc/rfc9639/ Section 8.8 const decoded = base64ToBytes(value); const view = toDataView(decoded); const pictureType = view.getUint32(0, false); const mediaTypeLength = view.getUint32(4, false); const mediaType = String.fromCharCode(...decoded.subarray(8, 8 + mediaTypeLength)); // ASCII const descriptionLength = view.getUint32(8 + mediaTypeLength, false); const description = textDecoder.decode(decoded.subarray( 12 + mediaTypeLength, 12 + mediaTypeLength + descriptionLength, )); const dataLength = view.getUint32(mediaTypeLength + descriptionLength + 28); const data = decoded.subarray( mediaTypeLength + descriptionLength + 32, mediaTypeLength + descriptionLength + 32 + dataLength, ); metadataTags.images ??= []; metadataTags.images.push({ data, mimeType: mediaType, kind: pictureType === 3 ? 'coverFront' : pictureType === 4 ? 'coverBack' : 'unknown', name: undefined, description: description || undefined, }); }; break; } } }; export const createVorbisComments = (headerBytes: Uint8Array, tags: MetadataTags, writeImages: boolean) => { // https://datatracker.ietf.org/doc/html/rfc7845#section-5.2 const commentHeaderParts: Uint8Array[] = [ headerBytes, ]; const vendorString = 'Mediabunny'; const encodedVendorString = textEncoder.encode(vendorString); let currentBuffer = new Uint8Array(4 + encodedVendorString.length); let currentView = new DataView(currentBuffer.buffer); currentView.setUint32(0, encodedVendorString.length, true); currentBuffer.set(encodedVendorString, 4); commentHeaderParts.push(currentBuffer); const writtenTags = new Set(); const addCommentTag = (key: string, value: string) => { const joined = `${key}=${value}`; const encoded = textEncoder.encode(joined); currentBuffer = new Uint8Array(4 + encoded.length); currentView = new DataView(currentBuffer.buffer); currentView.setUint32(0, encoded.length, true); currentBuffer.set(encoded, 4); commentHeaderParts.push(currentBuffer); writtenTags.add(key); }; for (const { key, value } of keyValueIterator(tags)) { switch (key) { case 'title': { addCommentTag('TITLE', value); }; break; case 'description': { addCommentTag('DESCRIPTION', value); }; break; case 'artist': { addCommentTag('ARTIST', value); }; break; case 'album': { addCommentTag('ALBUM', value); }; break; case 'albumArtist': { addCommentTag('ALBUMARTIST', value); }; break; case 'genre': { addCommentTag('GENRE', value); }; break; case 'date': { const rawVersion = tags.raw?.['DATE'] ?? tags.raw?.['date']; if (rawVersion && typeof rawVersion === 'string') { addCommentTag('DATE', rawVersion); } else { addCommentTag('DATE', value.toISOString().slice(0, 10)); } }; break; case 'comment': { addCommentTag('COMMENT', value); }; break; case 'lyrics': { addCommentTag('LYRICS', value); }; break; case 'trackNumber': { addCommentTag('TRACKNUMBER', value.toString()); }; break; case 'tracksTotal': { addCommentTag('TRACKTOTAL', value.toString()); }; break; case 'discNumber': { addCommentTag('DISCNUMBER', value.toString()); }; break; case 'discsTotal': { addCommentTag('DISCTOTAL', value.toString()); }; break; case 'images': { // For example, in .flac, we put the pictures in a different section, // not in the Vorbis comment header. if (!writeImages) { break; } for (const image of value) { // https://datatracker.ietf.org/doc/rfc9639/ Section 8.8 const pictureType = image.kind === 'coverFront' ? 3 : image.kind === 'coverBack' ? 4 : 0; const encodedMediaType = new Uint8Array(image.mimeType.length); for (let i = 0; i < image.mimeType.length; i++) { encodedMediaType[i] = image.mimeType.charCodeAt(i); } const encodedDescription = textEncoder.encode(image.description ?? ''); const buffer = new Uint8Array( 4 // Picture type + 4 // MIME type length + encodedMediaType.length // MIME type + 4 // Description length + encodedDescription.length // Description + 16 // Width, height, color depth, number of colors + 4 // Picture data length + image.data.length, // Picture data ); const view = toDataView(buffer); view.setUint32(0, pictureType, false); view.setUint32(4, encodedMediaType.length, false); buffer.set(encodedMediaType, 8); view.setUint32(8 + encodedMediaType.length, encodedDescription.length, false); buffer.set(encodedDescription, 12 + encodedMediaType.length); // Skip a bunch of fields (width, height, color depth, number of colors) view.setUint32( 28 + encodedMediaType.length + encodedDescription.length, image.data.length, false, ); buffer.set( image.data, 32 + encodedMediaType.length + encodedDescription.length, ); const encoded = bytesToBase64(buffer); addCommentTag('METADATA_BLOCK_PICTURE', encoded); } }; break; case 'raw': { // Handled later }; break; default: assertNever(key); } } if (tags.raw) { for (const key in tags.raw) { const value = tags.raw[key] ?? tags.raw[key.toLowerCase()]; if (key === 'vendor' || value == null || writtenTags.has(key)) { continue; } if (typeof value === 'string') { addCommentTag(key, value); } } } const listLengthBuffer = new Uint8Array(4); toDataView(listLengthBuffer).setUint32(0, writtenTags.size, true); commentHeaderParts.splice(2, 0, listLengthBuffer); // Insert after the header and vendor section // Merge all comment header parts into a single buffer const commentHeaderLength = commentHeaderParts.reduce((a, b) => a + b.length, 0); const commentHeader = new Uint8Array(commentHeaderLength); let pos = 0; for (const part of commentHeaderParts) { commentHeader.set(part, pos); pos += part.length; } return commentHeader; }; // ============================================================================ // AC-3 / E-AC-3 Parsing // Reference: ETSI TS 102 366 V1.4.1 // ============================================================================ /** * Channel counts indexed by acmod (Table 4.3). * Does NOT include LFE - add lfeon to get total channel count. */ export const AC3_ACMOD_CHANNEL_COUNTS = [2, 1, 2, 3, 3, 4, 4, 5]; export interface Ac3FrameInfo { /** Sample rate code */ fscod: number; /** Bitstream ID */ bsid: number; /** Bitstream mode */ bsmod: number; /** Audio coding mode */ acmod: number; /** LFE channel on */ lfeon: number; /** Bit rate code (0-18, maps to bitrate via Table F.4.1) */ bitRateCode: number; } /** * Parse an AC-3 syncframe to extract BSI (Bit Stream Information) fields. * Section 4.3 */ export const parseAc3SyncFrame = (data: Uint8Array): Ac3FrameInfo | null => { if (data.length < 7) { return null; } // Check sync word (0x0B77) if (data[0] !== 0x0B || data[1] !== 0x77) { return null; } const bitstream = new Bitstream(data); bitstream.skipBits(16); // sync word bitstream.skipBits(16); // crc1 const fscod = bitstream.readBits(2); if (fscod === 3) { return null; // Reserved, invalid } const frmsizecod = bitstream.readBits(6); const bsid = bitstream.readBits(5); // Verify this is AC-3 if (bsid > 8) { return null; } const bsmod = bitstream.readBits(3); const acmod = bitstream.readBits(3); // Skip cmixlev (center downmix level) if three front channels are in use (L, C, R). if ((acmod & 0x1) !== 0 && acmod !== 0x1) { bitstream.skipBits(2); } // Skip surmixlev (surround downmix level) if surround channels are in use. if ((acmod & 0x4) !== 0) { bitstream.skipBits(2); } // Skip dsurmod if stereo (acmod === 2) if (acmod === 0x2) { bitstream.skipBits(2); } const lfeon = bitstream.readBits(1); const bitRateCode = Math.floor(frmsizecod / 2); return { fscod, bsid, bsmod, acmod, lfeon, bitRateCode }; }; /** * AC-3 frame sizes in bytes, indexed by [3 * frmsizecod + fscod]. * fscod: 0=48kHz, 1=44.1kHz, 2=32kHz * Values are 16-bit words * 2 (to convert to bytes). * Table 4.13 */ export const AC3_FRAME_SIZES = [ // frmsizecod, [48kHz, 44.1kHz, 32kHz] in bytes 64 * 2, 69 * 2, 96 * 2, 64 * 2, 70 * 2, 96 * 2, 80 * 2, 87 * 2, 120 * 2, 80 * 2, 88 * 2, 120 * 2, 96 * 2, 104 * 2, 144 * 2, 96 * 2, 105 * 2, 144 * 2, 112 * 2, 121 * 2, 168 * 2, 112 * 2, 122 * 2, 168 * 2, 128 * 2, 139 * 2, 192 * 2, 128 * 2, 140 * 2, 192 * 2, 160 * 2, 174 * 2, 240 * 2, 160 * 2, 175 * 2, 240 * 2, 192 * 2, 208 * 2, 288 * 2, 192 * 2, 209 * 2, 288 * 2, 224 * 2, 243 * 2, 336 * 2, 224 * 2, 244 * 2, 336 * 2, 256 * 2, 278 * 2, 384 * 2, 256 * 2, 279 * 2, 384 * 2, 320 * 2, 348 * 2, 480 * 2, 320 * 2, 349 * 2, 480 * 2, 384 * 2, 417 * 2, 576 * 2, 384 * 2, 418 * 2, 576 * 2, 448 * 2, 487 * 2, 672 * 2, 448 * 2, 488 * 2, 672 * 2, 512 * 2, 557 * 2, 768 * 2, 512 * 2, 558 * 2, 768 * 2, 640 * 2, 696 * 2, 960 * 2, 640 * 2, 697 * 2, 960 * 2, 768 * 2, 835 * 2, 1152 * 2, 768 * 2, 836 * 2, 1152 * 2, 896 * 2, 975 * 2, 1344 * 2, 896 * 2, 976 * 2, 1344 * 2, 1024 * 2, 1114 * 2, 1536 * 2, 1024 * 2, 1115 * 2, 1536 * 2, 1152 * 2, 1253 * 2, 1728 * 2, 1152 * 2, 1254 * 2, 1728 * 2, 1280 * 2, 1393 * 2, 1920 * 2, 1280 * 2, 1394 * 2, 1920 * 2, ]; /** Number of samples per AC-3 syncframe (always 1536) */ export const AC3_SAMPLES_PER_FRAME = 1536; /** * AC-3 registration_descriptor for MPEG-TS. * Section A.2.3 */ export const AC3_REGISTRATION_DESCRIPTOR = new Uint8Array([0x05, 0x04, 0x41, 0x43, 0x2d, 0x33]); /** E-AC-3 registration_descriptor for MPEG-TS/ */ export const EAC3_REGISTRATION_DESCRIPTOR = new Uint8Array([0x05, 0x04, 0x45, 0x41, 0x43, 0x33]); /** Number of audio blocks per syncframe, indexed by numblkscod */ export const EAC3_NUMBLKS_TABLE = [1, 2, 3, 6]; /** * E-AC-3 independent substream info. * Each independent substream represents a separate audio program. */ export interface Eac3SubstreamInfo { /** Sample rate code */ fscod: number; /** Sample rate code 2 (ATSC A/52:2018) */ fscod2: number | null; /** Bitstream ID */ bsid: number; /** Bitstream mode */ bsmod: number; /** Audio coding mode */ acmod: number; /** LFE channel on */ lfeon: number; /** Number of dependent substreams */ numDepSub: number; /** Channel locations for dependent substreams */ chanLoc: number; } /** * E-AC-3 decoder configuration (dec3 box contents). */ export interface Eac3FrameInfo { /** Data rate in kbps */ dataRate: number; /** Independent substreams */ substreams: Eac3SubstreamInfo[]; } /** * Parse an E-AC-3 syncframe to extract BSI fields. * Section E.1.2 */ export const parseEac3SyncFrame = (data: Uint8Array): Eac3FrameInfo | null => { if (data.length < 6) { return null; } // Check sync word (0x0B77) if (data[0] !== 0x0B || data[1] !== 0x77) { return null; } const bitstream = new Bitstream(data); bitstream.skipBits(16); // sync word const strmtyp = bitstream.readBits(2); bitstream.skipBits(3); // substreamid // Only parse independent substreams (strmtyp 0 or 2) if (strmtyp !== 0 && strmtyp !== 2) { return null; } const frmsiz = bitstream.readBits(11); const fscod = bitstream.readBits(2); let fscod2 = 0; let numblkscod: number; if (fscod === 3) { // fscod2 enables reduced sample rates (24/22.05/16 kHz) per ATSC A/52:2018 fscod2 = bitstream.readBits(2); numblkscod = 3; // Implicitly 6 blocks when fscod=3 } else { numblkscod = bitstream.readBits(2); } const acmod = bitstream.readBits(3); const lfeon = bitstream.readBits(1); const bsid = bitstream.readBits(5); // Verify this is E-AC-3 if (bsid < 11 || bsid > 16) { return null; } // Calculate data rate: ((frmsiz + 1) * fs) / (numblks * 16) const numblks = EAC3_NUMBLKS_TABLE[numblkscod]!; let fs: number; if (fscod < 3) { fs = AC3_SAMPLE_RATES[fscod]! / 1000; } else { fs = EAC3_REDUCED_SAMPLE_RATES[fscod2]! / 1000; } const dataRate = Math.round(((frmsiz + 1) * fs) / (numblks * 16)); // These fields require parsing beyond the first frame. // Defaults are correct for almost all content. const bsmod = 0; const numDepSub = 0; const chanLoc = 0; const substream: Eac3SubstreamInfo = { fscod, fscod2, bsid, bsmod, acmod, lfeon, numDepSub, chanLoc, }; return { dataRate, substreams: [substream], }; }; /** * Parse a dec3 box to extract E-AC-3 parameters. * Section F.6 */ export const parseEac3Config = (data: Uint8Array): Eac3FrameInfo | null => { if (data.length < 2) { return null; } const bitstream = new Bitstream(data); const dataRate = bitstream.readBits(13); const numIndSub = bitstream.readBits(3); const substreams: Eac3SubstreamInfo[] = []; for (let i = 0; i <= numIndSub; i++) { // Check we have enough data for this substream // Each substream needs at least 24 bits (3 bytes) without dependent subs if (Math.ceil(bitstream.pos / 8) + 3 > data.length) { break; } const fscod = bitstream.readBits(2); const bsid = bitstream.readBits(5); bitstream.skipBits(1); // reserved bitstream.skipBits(1); // asvc const bsmod = bitstream.readBits(3); const acmod = bitstream.readBits(3); const lfeon = bitstream.readBits(1); bitstream.skipBits(3); // reserved const numDepSub = bitstream.readBits(4); let chanLoc = 0; if (numDepSub > 0) { chanLoc = bitstream.readBits(9); } else { bitstream.skipBits(1); // reserved } substreams.push({ fscod, fscod2: null, bsid, bsmod, acmod, lfeon, numDepSub, chanLoc, }); } if (substreams.length === 0) { return null; } return { dataRate, substreams }; }; /** * Get sample rate from E-AC-3 config. * See ATSC A/52:2018 for handling fscod2. */ export const getEac3SampleRate = (config: Eac3FrameInfo): number | null => { const sub = config.substreams[0]; assert(sub); if (sub.fscod < 3) { return AC3_SAMPLE_RATES[sub.fscod]!; } else if (sub.fscod2 !== null && sub.fscod2 < 3) { return EAC3_REDUCED_SAMPLE_RATES[sub.fscod2]!; } return null; }; /** * Get channel count from E-AC-3 config (first independent substream only). */ export const getEac3ChannelCount = (config: Eac3FrameInfo): number => { const sub = config.substreams[0]; assert(sub); let channels = AC3_ACMOD_CHANNEL_COUNTS[sub.acmod]! + sub.lfeon; // Add channels from dependent substreams if (sub.numDepSub > 0) { const CHAN_LOC_COUNTS = [2, 2, 1, 1, 2, 2, 2, 1, 1]; for (let bit = 0; bit < 9; bit++) { if (sub.chanLoc & (1 << (8 - bit))) { channels += CHAN_LOC_COUNTS[bit]!; } } } return channels; }; ===== src/output-format.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AdtsMuxer } from './adts/adts-muxer'; import { AUDIO_CODECS, AudioCodec, MediaCodec, NON_PCM_AUDIO_CODECS, PCM_AUDIO_CODECS, SUBTITLE_CODECS, SubtitleCodec, VIDEO_CODECS, VideoCodec, } from './codec'; import { FlacMuxer } from './flac/flac-muxer'; import { IsobmffMuxer } from './isobmff/isobmff-muxer'; import { MatroskaMuxer } from './matroska/matroska-muxer'; import { MediaSource } from './media-source'; import { Mp3Muxer } from './mp3/mp3-muxer'; import { Muxer } from './muxer'; import { OggMuxer } from './ogg/ogg-muxer'; import { Output, OutputTrack, TrackType } from './output'; import { MpegTsMuxer } from './mpeg-ts/mpeg-ts-muxer'; import { WaveMuxer } from './wave/wave-muxer'; import { HlsMuxer } from './hls/hls-muxer'; import { HLS_MIME_TYPE } from './hls/hls-misc'; import { MaybePromise, FilePath, toArray } from './misc'; import { Target } from './target'; /** * Specifies an inclusive range of integers. * @group Miscellaneous * @public */ export type InclusiveIntegerRange = { /** The integer cannot be less than this. */ min: number; /** The integer cannot be greater than this. */ max: number; }; /** * Specifies the number of tracks (for each track type and in total) that an output format supports. * @group Output formats * @public */ export type TrackCountLimits = { [K in TrackType]: InclusiveIntegerRange; } & { /** Specifies the overall allowed range of track counts for the output format. */ total: InclusiveIntegerRange; }; /** * Base class representing an output media file format. * @group Output formats * @public */ export abstract class OutputFormat { /** @internal */ abstract _createMuxer(output: Output): Muxer; /** @internal */ abstract get _name(): string; /** The file extension used by this output format, beginning with a dot. */ abstract get fileExtension(): string; /** The base MIME type of the output format. */ abstract get mimeType(): string; /** Returns a list of media codecs that this output format can contain. */ abstract getSupportedCodecs(): MediaCodec[]; /** Returns the number of tracks that this output format supports. */ abstract getSupportedTrackCounts(): TrackCountLimits; /** Whether this output format supports video rotation metadata. */ abstract get supportsVideoRotationMetadata(): boolean; /** * Whether this output format's tracks store timestamped media data. When `true`, the timestamps of added packets * will be respected, allowing things like gaps in media data or non-zero start times. When `false`, the format's * media data implicitly starts at zero and follows an implicit sequential timing from there, using the intrinsic * durations of the media data. */ abstract get supportsTimestampedMediaData(): boolean; /** Returns a list of video codecs that this output format can contain. */ getSupportedVideoCodecs() { return this.getSupportedCodecs() .filter(codec => (VIDEO_CODECS as readonly string[]).includes(codec)) as VideoCodec[]; } /** Returns a list of audio codecs that this output format can contain. */ getSupportedAudioCodecs() { return this.getSupportedCodecs() .filter(codec => (AUDIO_CODECS as readonly string[]).includes(codec)) as AudioCodec[]; } /** Returns a list of subtitle codecs that this output format can contain. */ getSupportedSubtitleCodecs() { return this.getSupportedCodecs() .filter(codec => (SUBTITLE_CODECS as readonly string[]).includes(codec)) as SubtitleCodec[]; } /** @internal */ // eslint-disable-next-line @typescript-eslint/no-unused-vars _codecUnsupportedHint(codec: MediaCodec) { return ''; } } /** * ISOBMFF-specific output options. * @group Output formats * @public */ export type IsobmffOutputFormatOptions = { /** * Controls the placement of metadata in the file. Placing metadata at the start of the file is known as "Fast * Start", which results in better playback at the cost of more required processing or memory. * * Use `false` to disable Fast Start, placing the metadata at the end of the file. Fastest and uses the least * memory. * * Use `'in-memory'` to produce a file with Fast Start by keeping all media chunks in memory until the file is * finalized. This produces a high-quality and compact output at the cost of a more expensive finalization step and * higher memory requirements. Data will be written monotonically (in order) when this option is set. * * Use `'reserve'` to reserve space at the start of the file into which the metadata will be written later. This * produces a file with Fast Start but requires knowledge about the expected length of the file beforehand. When * using this option, you must set the {@link BaseTrackMetadata.maximumPacketCount} field in the track metadata * for all tracks. * * Use `'fragmented'` to place metadata at the start of the file by creating a fragmented file (fMP4). In a * fragmented file, chunks of media and their metadata are written to the file in "fragments", eliminating the need * to put all metadata in one place. Fragmented files are useful for streaming contexts, as each fragment can be * played individually without requiring knowledge of the other fragments. Furthermore, they remain lightweight to * create even for very large files, as they don't require all media to be kept in memory. However, fragmented files * are not as widely and wholly supported as regular MP4/MOV files. Data will be written monotonically (in order) * when this option is set. * * When this field is not defined, either `false` or `'in-memory'` will be used, automatically determined based on * the type of output target used. */ fastStart?: false | 'in-memory' | 'reserve' | 'fragmented'; /** * When using `fastStart: 'fragmented'`, this field controls the minimum duration of each fragment, in seconds. * New fragments will only be created when the current fragment is longer than this value. Defaults to 1 second. */ minimumFragmentDuration?: number; /** * The metadata format to use for writing metadata tags. * * - `'auto'` (default): Behaves like `'mdir'` for MP4 and like `'udta'` for QuickTime, matching FFmpeg's default * behavior. * - `'mdir'`: Write tags into `moov/udta/meta` using the 'mdir' handler format. * - `'mdta'`: Write tags into `moov/udta/meta` using the 'mdta' handler format, equivalent to FFmpeg's * `use_metadata_tags` flag. This allows for custom keys of arbitrary length. * - `'udta'`: Write tags directly into `moov/udta`. */ metadataFormat?: 'auto' | 'mdir' | 'mdta' | 'udta'; /** * Will be called once the ftyp (File Type) box of the output file has been written. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. */ onFtyp?: (data: Uint8Array, position: number) => unknown; /** * Will be called once the moov (Movie) box of the output file has been written. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. */ onMoov?: (data: Uint8Array, position: number) => unknown; /** * Will be called for each finalized mdat (Media Data) box of the output file. Usage of this callback is not * recommended when not using `fastStart: 'fragmented'`, as there will be one monolithic mdat box which might * require large amounts of memory. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. */ onMdat?: (data: Uint8Array, position: number) => unknown; /** * Will be called for each finalized moof (Movie Fragment) box of the output file. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. * @param timestamp - The start timestamp of the fragment in seconds. */ onMoof?: (data: Uint8Array, position: number, timestamp: number) => unknown; }; /** * Format representing files compatible with the ISO base media file format (ISOBMFF), like MP4 or MOV files. * @group Output formats * @public */ export abstract class IsobmffOutputFormat extends OutputFormat { /** @internal */ _options: IsobmffOutputFormatOptions; /** Internal constructor. */ constructor(options: IsobmffOutputFormatOptions = {}) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if ( options.fastStart !== undefined && ![false, 'in-memory', 'reserve', 'fragmented'].includes(options.fastStart) ) { throw new TypeError( 'options.fastStart, when provided, must be false, \'in-memory\', \'reserve\', or \'fragmented\'.', ); } if ( options.minimumFragmentDuration !== undefined && (!Number.isFinite(options.minimumFragmentDuration) || options.minimumFragmentDuration < 0) ) { throw new TypeError('options.minimumFragmentDuration, when provided, must be a non-negative number.'); } if (options.onFtyp !== undefined && typeof options.onFtyp !== 'function') { throw new TypeError('options.onFtyp, when provided, must be a function.'); } if (options.onMoov !== undefined && typeof options.onMoov !== 'function') { throw new TypeError('options.onMoov, when provided, must be a function.'); } if (options.onMdat !== undefined && typeof options.onMdat !== 'function') { throw new TypeError('options.onMdat, when provided, must be a function.'); } if (options.onMoof !== undefined && typeof options.onMoof !== 'function') { throw new TypeError('options.onMoof, when provided, must be a function.'); } if ( options.metadataFormat !== undefined && !['mdir', 'mdta', 'udta', 'auto'].includes(options.metadataFormat) ) { throw new TypeError( 'options.metadataFormat, when provided, must be either \'auto\', \'mdir\', \'mdta\', or \'udta\'.', ); } super(); this._options = options; } getSupportedTrackCounts(): TrackCountLimits { const max = 2 ** 32 - 1; // Have fun reaching this one return { video: { min: 0, max }, audio: { min: 0, max }, subtitle: { min: 0, max }, total: { min: 1, max }, }; } get supportsVideoRotationMetadata() { return true; } get supportsTimestampedMediaData() { return true; } /** @internal */ _createMuxer(output: Output) { return new IsobmffMuxer(output, this); } } /** * MPEG-4 Part 14 (MP4) file format. Supports most codecs. * @group Output formats * @public */ export class Mp4OutputFormat extends IsobmffOutputFormat { /** Creates a new {@link Mp4OutputFormat} configured with the specified `options`. */ constructor(options?: IsobmffOutputFormatOptions) { super(options); } /** @internal */ get _name() { return 'MP4'; } get fileExtension() { return '.mp4'; } get mimeType() { return 'video/mp4'; } getSupportedCodecs(): MediaCodec[] { return [ ...VIDEO_CODECS, ...NON_PCM_AUDIO_CODECS, // These are supported via ISO/IEC 23003-5: 'pcm-s16', 'pcm-s16be', 'pcm-s24', 'pcm-s24be', 'pcm-s32', 'pcm-s32be', 'pcm-f32', 'pcm-f32be', 'pcm-f64', 'pcm-f64be', ...SUBTITLE_CODECS, ]; } /** @internal */ override _codecUnsupportedHint(codec: MediaCodec) { if (new MovOutputFormat().getSupportedCodecs().includes(codec)) { return ' Switching to MOV will grant support for this codec.'; } return ''; } } /** * CMAF-specific output options. * @group Output formats * @public */ export type CmafOutputFormatOptions = Omit & { /** * Controls the minimum duration of each fragment, in seconds. New fragments will only be created when the current * fragment is longer than this value. Defaults to `Infinity`, meaning the file will contain only one fragment. */ minimumFragmentDuration?: number; }; /** * Creates a single Common Media Application Format (CMAF) segment. An init segment will be written to the * {@link Target} specified in {@link OutputOptions.initTarget}. Supports most codecs. * @group Output formats * @public */ export class CmafOutputFormat extends IsobmffOutputFormat { /** Creates a new {@link CmafOutputFormat} configured with the specified `options`. */ constructor(options?: CmafOutputFormatOptions) { super(options); } /** @internal */ get _name() { return 'CMAF'; } get fileExtension() { return '.m4s'; } get mimeType() { return 'video/mp4'; } getSupportedCodecs(): MediaCodec[] { return [ ...VIDEO_CODECS, ...NON_PCM_AUDIO_CODECS, // These are supported via ISO/IEC 23003-5: 'pcm-s16', 'pcm-s16be', 'pcm-s24', 'pcm-s24be', 'pcm-s32', 'pcm-s32be', 'pcm-f32', 'pcm-f32be', 'pcm-f64', 'pcm-f64be', ...SUBTITLE_CODECS, ]; } } /** * QuickTime File Format (QTFF), often called MOV. Supports all video and audio codecs, but not subtitle codecs. * @group Output formats * @public */ export class MovOutputFormat extends IsobmffOutputFormat { /** Creates a new {@link MovOutputFormat} configured with the specified `options`. */ constructor(options?: IsobmffOutputFormatOptions) { super(options); } /** @internal */ get _name() { return 'MOV'; } get fileExtension() { return '.mov'; } get mimeType() { return 'video/quicktime'; } getSupportedCodecs(): MediaCodec[] { return [ ...VIDEO_CODECS, ...AUDIO_CODECS, ]; } /** @internal */ override _codecUnsupportedHint(codec: MediaCodec) { if (new Mp4OutputFormat().getSupportedCodecs().includes(codec)) { return ' Switching to MP4 will grant support for this codec.'; } return ''; } } /** * Matroska-specific output options. * @group Output formats * @public */ export type MkvOutputFormatOptions = { /** * Configures the output to only append new data at the end, useful for live-streaming the file as it's being * created. When enabled, some features such as storing duration and seeking will be disabled or impacted, so don't * use this option when you want to write out a clean file for later use. */ appendOnly?: boolean; /** * This field controls the minimum duration of each Matroska cluster, in seconds. New clusters will only be created * when the current cluster is longer than this value. Defaults to 1 second. */ minimumClusterDuration?: number; /** * Will be called once the EBML header of the output file has been written. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. */ onEbmlHeader?: (data: Uint8Array, position: number) => void; /** * Will be called once the header part of the Matroska Segment element has been written. The header data includes * the Segment element and everything inside it, up to (but excluding) the first Matroska Cluster. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. */ onSegmentHeader?: (data: Uint8Array, position: number) => unknown; /** * Will be called for each finalized Matroska Cluster of the output file. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. * @param timestamp - The start timestamp of the cluster in seconds. */ onCluster?: (data: Uint8Array, position: number, timestamp: number) => unknown; }; /** * Matroska file format. * * Supports writing transparent video. For a video track to be marked as transparent, the first packet added must * contain alpha side data. * * @group Output formats * @public */ export class MkvOutputFormat extends OutputFormat { /** @internal */ _options: MkvOutputFormatOptions; /** Creates a new {@link MkvOutputFormat} configured with the specified `options`. */ constructor(options: MkvOutputFormatOptions = {}) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.appendOnly !== undefined && typeof options.appendOnly !== 'boolean') { throw new TypeError('options.appendOnly, when provided, must be a boolean.'); } if ( options.minimumClusterDuration !== undefined && (!Number.isFinite(options.minimumClusterDuration) || options.minimumClusterDuration < 0) ) { throw new TypeError('options.minimumClusterDuration, when provided, must be a non-negative number.'); } if (options.onEbmlHeader !== undefined && typeof options.onEbmlHeader !== 'function') { throw new TypeError('options.onEbmlHeader, when provided, must be a function.'); } if (options.onSegmentHeader !== undefined && typeof options.onSegmentHeader !== 'function') { throw new TypeError('options.onHeader, when provided, must be a function.'); } if (options.onCluster !== undefined && typeof options.onCluster !== 'function') { throw new TypeError('options.onCluster, when provided, must be a function.'); } super(); this._options = options; } /** @internal */ _createMuxer(output: Output) { return new MatroskaMuxer(output, this); } /** @internal */ get _name() { return 'Matroska'; } getSupportedTrackCounts(): TrackCountLimits { const max = 127; return { video: { min: 0, max }, audio: { min: 0, max }, subtitle: { min: 0, max }, total: { min: 1, max }, }; } get fileExtension() { return '.mkv'; } get mimeType() { return 'video/x-matroska'; } getSupportedCodecs(): MediaCodec[] { return [ ...VIDEO_CODECS, ...NON_PCM_AUDIO_CODECS, ...PCM_AUDIO_CODECS.filter(codec => !['pcm-s8', 'pcm-f32be', 'pcm-f64be', 'ulaw', 'alaw'].includes(codec)), ...SUBTITLE_CODECS, ]; } get supportsVideoRotationMetadata() { // While it technically does support it with ProjectionPoseRoll, many players appear to ignore this value return false; } get supportsTimestampedMediaData() { return true; } } /** * WebM-specific output options. * @group Output formats * @public */ export type WebMOutputFormatOptions = MkvOutputFormatOptions; /** * WebM file format, based on Matroska. * * Supports writing transparent video. For a video track to be marked as transparent, the first packet added must * contain alpha side data. * * @group Output formats * @public */ export class WebMOutputFormat extends MkvOutputFormat { /** Creates a new {@link WebMOutputFormat} configured with the specified `options`. */ constructor(options?: MkvOutputFormatOptions) { super(options); } override getSupportedCodecs(): MediaCodec[] { return [ ...VIDEO_CODECS.filter(codec => ['vp8', 'vp9', 'av1'].includes(codec)), ...AUDIO_CODECS.filter(codec => ['opus', 'vorbis'].includes(codec)), ...SUBTITLE_CODECS, ]; } /** @internal */ override get _name() { return 'WebM'; } override get fileExtension() { return '.webm'; } override get mimeType() { return 'video/webm'; } /** @internal */ override _codecUnsupportedHint(codec: MediaCodec) { if (new MkvOutputFormat().getSupportedCodecs().includes(codec)) { return ' Switching to MKV will grant support for this codec.'; } return ''; } } /** * MP3-specific output options. * @group Output formats * @public */ export type Mp3OutputFormatOptions = { /** * Controls whether the Xing header, which contains additional metadata as well as an index, is written to the start * of the MP3 file. When disabled, the writing process becomes append-only. Defaults to `true`. */ xingHeader?: boolean; /** * Will be called once the Xing metadata frame is finalized. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. */ onXingFrame?: (data: Uint8Array, position: number) => unknown; }; /** * MP3 file format. * @group Output formats * @public */ export class Mp3OutputFormat extends OutputFormat { /** @internal */ _options: Mp3OutputFormatOptions; /** Creates a new {@link Mp3OutputFormat} configured with the specified `options`. */ constructor(options: Mp3OutputFormatOptions = {}) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.xingHeader !== undefined && typeof options.xingHeader !== 'boolean') { throw new TypeError('options.xingHeader, when provided, must be a boolean.'); } if (options.onXingFrame !== undefined && typeof options.onXingFrame !== 'function') { throw new TypeError('options.onXingFrame, when provided, must be a function.'); } super(); this._options = options; } /** @internal */ _createMuxer(output: Output) { return new Mp3Muxer(output, this); } /** @internal */ get _name() { return 'MP3'; } getSupportedTrackCounts(): TrackCountLimits { return { video: { min: 0, max: 0 }, audio: { min: 1, max: 1 }, subtitle: { min: 0, max: 0 }, total: { min: 1, max: 1 }, }; } get fileExtension() { return '.mp3'; } get mimeType() { return 'audio/mpeg'; } getSupportedCodecs(): MediaCodec[] { return ['mp3']; } get supportsVideoRotationMetadata() { return false; } get supportsTimestampedMediaData() { return false; } } /** * WAVE-specific output options. * @group Output formats * @public */ export type WavOutputFormatOptions = { /** * When enabled, an RF64 file will be written, allowing for file sizes to exceed 4 GiB, which is otherwise not * possible for regular WAVE files. */ large?: boolean; /** * The metadata format to use for writing metadata tags. * * - `'info'` (default): Writes metadata into a RIFF INFO LIST chunk, the default way to contain metadata tags * within WAVE. Only allows for a limited subset of tags to be written. * - `'id3'`: Writes metadata into an ID3 chunk. Non-default, but used by many taggers in practice. Allows for a * much larger and richer set of tags to be written. */ metadataFormat?: 'info' | 'id3'; /** * Will be called once the file header is written. The header consists of the RIFF header, the format chunk, * metadata chunks, and the start of the data chunk (with a placeholder size of 0). */ onHeader?: (data: Uint8Array, position: number) => unknown; }; /** * WAVE file format, based on RIFF. * @group Output formats * @public */ export class WavOutputFormat extends OutputFormat { /** @internal */ _options: WavOutputFormatOptions; /** Creates a new {@link WavOutputFormat} configured with the specified `options`. */ constructor(options: WavOutputFormatOptions = {}) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.large !== undefined && typeof options.large !== 'boolean') { throw new TypeError('options.large, when provided, must be a boolean.'); } if (options.metadataFormat !== undefined && !['info', 'id3'].includes(options.metadataFormat)) { throw new TypeError('options.metadataFormat, when provided, must be either \'info\' or \'id3\'.'); } if (options.onHeader !== undefined && typeof options.onHeader !== 'function') { throw new TypeError('options.onHeader, when provided, must be a function.'); } super(); this._options = options; } /** @internal */ _createMuxer(output: Output) { return new WaveMuxer(output, this); } /** @internal */ get _name() { return 'WAVE'; } getSupportedTrackCounts(): TrackCountLimits { return { video: { min: 0, max: 0 }, audio: { min: 1, max: 1 }, subtitle: { min: 0, max: 0 }, total: { min: 1, max: 1 }, }; } get fileExtension() { return '.wav'; } get mimeType() { return 'audio/wav'; } getSupportedCodecs(): MediaCodec[] { return [ ...PCM_AUDIO_CODECS.filter(codec => ['pcm-s16', 'pcm-s24', 'pcm-s32', 'pcm-f32', 'pcm-u8', 'ulaw', 'alaw'].includes(codec), ), ]; } get supportsVideoRotationMetadata() { return false; } get supportsTimestampedMediaData() { return false; } } /** * Ogg-specific output options. * @group Output formats * @public */ export type OggOutputFormatOptions = { /** * The maximum duration of each Ogg page, in seconds. This is useful for streaming contexts where more frequent page * output is desired. By default, pages are only flushed when they exceed a certain size. */ maximumPageDuration?: number; /** * Will be called for each Ogg page that is written. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. * @param source - The {@link MediaSource} backing the page's logical bitstream (track). */ onPage?: (data: Uint8Array, position: number, source: MediaSource) => unknown; }; /** * Ogg file format. * @group Output formats * @public */ export class OggOutputFormat extends OutputFormat { /** @internal */ _options: OggOutputFormatOptions; /** Creates a new {@link OggOutputFormat} configured with the specified `options`. */ constructor(options: OggOutputFormatOptions = {}) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if ( options.maximumPageDuration !== undefined && (!Number.isFinite(options.maximumPageDuration) || options.maximumPageDuration <= 0) ) { throw new TypeError('options.maximumPageDuration, when provided, must be a positive number.'); } if (options.onPage !== undefined && typeof options.onPage !== 'function') { throw new TypeError('options.onPage, when provided, must be a function.'); } super(); this._options = options; } /** @internal */ _createMuxer(output: Output) { return new OggMuxer(output, this); } /** @internal */ get _name() { return 'Ogg'; } getSupportedTrackCounts(): TrackCountLimits { const max = 2 ** 32; // Have fun reaching this one return { video: { min: 0, max: 0 }, audio: { min: 0, max }, subtitle: { min: 0, max: 0 }, total: { min: 1, max }, }; } get fileExtension() { return '.ogg'; } get mimeType() { return 'application/ogg'; } getSupportedCodecs(): MediaCodec[] { return [ ...AUDIO_CODECS.filter(codec => ['vorbis', 'opus'].includes(codec)), ]; } get supportsVideoRotationMetadata() { return false; } get supportsTimestampedMediaData() { return false; } } /** * ADTS-specific output options. * @group Output formats * @public */ export type AdtsOutputFormatOptions = { /** * Will be called for each ADTS frame that is written. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. */ onFrame?: (data: Uint8Array, position: number) => unknown; }; /** * ADTS file format. * @group Output formats * @public */ export class AdtsOutputFormat extends OutputFormat { /** @internal */ _options: AdtsOutputFormatOptions; /** Creates a new {@link AdtsOutputFormat} configured with the specified `options`. */ constructor(options: AdtsOutputFormatOptions = {}) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.onFrame !== undefined && typeof options.onFrame !== 'function') { throw new TypeError('options.onFrame, when provided, must be a function.'); } super(); this._options = options; } /** @internal */ _createMuxer(output: Output) { return new AdtsMuxer(output, this); } /** @internal */ get _name() { return 'ADTS'; } getSupportedTrackCounts(): TrackCountLimits { return { video: { min: 0, max: 0 }, audio: { min: 1, max: 1 }, subtitle: { min: 0, max: 0 }, total: { min: 1, max: 1 }, }; } get fileExtension() { return '.aac'; } get mimeType() { return 'audio/aac'; } getSupportedCodecs(): MediaCodec[] { return ['aac']; } get supportsVideoRotationMetadata() { return false; } get supportsTimestampedMediaData() { return false; } } /** * FLAC-specific output options. * @group Output formats * @public */ export type FlacOutputFormatOptions = { /** * Configures the output to only append new data at the end, useful for live-streaming the file as it's being * created. When enabled, the STREAMINFO block will not be finalized with accurate min/max block sizes, frame sizes, * or total sample count, so don't use this option when you want to write out a clean file for later use. */ appendOnly?: boolean; /** * Will be called for each FLAC frame that is written. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. */ onFrame?: (data: Uint8Array, position: number) => unknown; }; /** * FLAC file format. * @group Output formats * @public */ export class FlacOutputFormat extends OutputFormat { /** @internal */ _options: FlacOutputFormatOptions; /** Creates a new {@link FlacOutputFormat} configured with the specified `options`. */ constructor(options: FlacOutputFormatOptions = {}) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.appendOnly !== undefined && typeof options.appendOnly !== 'boolean') { throw new TypeError('options.appendOnly, when provided, must be a boolean.'); } super(); this._options = options; } /** @internal */ _createMuxer(output: Output) { return new FlacMuxer(output, this); } /** @internal */ get _name() { return 'FLAC'; } getSupportedTrackCounts(): TrackCountLimits { return { video: { min: 0, max: 0 }, audio: { min: 1, max: 1 }, subtitle: { min: 0, max: 0 }, total: { min: 1, max: 1 }, }; } get fileExtension() { return '.flac'; } get mimeType() { return 'audio/flac'; } getSupportedCodecs(): MediaCodec[] { return ['flac']; } get supportsVideoRotationMetadata() { return false; } get supportsTimestampedMediaData() { return false; } } /** * MPEG-TS-specific output options. * @group Output formats * @public */ export type MpegTsOutputFormatOptions = { /** * Will be called for each 188-byte Transport Stream packet that is written. * * @param data - The raw bytes. * @param position - The byte offset of the data in the file. */ onPacket?: (data: Uint8Array, position: number) => unknown; }; /** * MPEG Transport Stream file format. * @group Output formats * @public */ export class MpegTsOutputFormat extends OutputFormat { /** @internal */ _options: MpegTsOutputFormatOptions; /** Creates a new {@link MpegTsOutputFormat} configured with the specified `options`. */ constructor(options: MpegTsOutputFormatOptions = {}) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (options.onPacket !== undefined && typeof options.onPacket !== 'function') { throw new TypeError('options.onPacket, when provided, must be a function.'); } super(); this._options = options; } /** @internal */ _createMuxer(output: Output) { return new MpegTsMuxer(output, this); } /** @internal */ get _name() { return 'MPEG-TS'; } getSupportedTrackCounts(): TrackCountLimits { const maxVideo = 16; // Stream IDs 0xE0-0xEF const maxAudio = 32; const maxTotal = maxVideo + maxAudio; return { video: { min: 0, max: maxVideo }, audio: { min: 0, max: maxAudio }, subtitle: { min: 0, max: 0 }, total: { min: 1, max: maxTotal }, }; } get fileExtension() { return '.ts'; } get mimeType() { return 'video/MP2T'; } getSupportedCodecs(): MediaCodec[] { return [ ...VIDEO_CODECS.filter(codec => ['avc', 'hevc'].includes(codec)), ...AUDIO_CODECS.filter(codec => ['aac', 'mp3', 'ac3', 'eac3'].includes(codec)), ]; } get supportsVideoRotationMetadata() { return false; } get supportsTimestampedMediaData() { return true; } } /** * Info about an HLS media playlist. * @group Output formats * @public */ export type HlsOutputPlaylistInfo = { /** The 1-based index of the media playlist in the master playlist. */ n: number; /** The output tracks contained in this playlist. */ tracks: OutputTrack[]; /** The format of the media segments in this playlist. */ segmentFormat: OutputFormat; }; /** * Info about an HLS media segment. * @group Output formats * @public */ export type HlsOutputSegmentInfo = { /** The 1-based index of the segment in the containing media playlist. */ n: number; /** If the segment is a single file, meaning it is a single segment file that covers the entire playlist. */ isSingleFile: boolean; /** The format of the media segment. */ format: OutputFormat; /** The media playlist to which this segment belongs. */ playlist: HlsOutputPlaylistInfo; }; /** * HLS-specific output options. * @group Output formats * @public */ export type HlsOutputFormatOptions = { /** * Specifies the file format of each media segment. Not all formats are supported by all players; prefer sticking * to the most commonly used ones: {@link MpegTsOutputFormat}, {@link CmafOutputFormat}, {@link AdtsOutputFormat}, * and {@link Mp3OutputFormat}. * * When an array of formats is specified, for each playlist, the first format that can contain all of the playlist's * tracks is chosen. This allows you to, for example, package audio into .aac files and video into .ts files. */ segmentFormat: OutputFormat | OutputFormat[]; /** * Specifies the target (max) duration in seconds for each media segment, defaulting to 2 seconds. * * Mediabunny will try not to emit media segments longer than the target duration, but it is forced to if key frames * are provided with a longer period than the target duration. Therefore, make sure to encode a key frame at least * every `targetDuration` seconds to guarantee segment length, controllable via * {@link VideoEncodingConfig.keyFrameInterval}. */ targetDuration?: number; /** * Whether to bundle all media segments for a playlist into a single file. Individual segments are then extracted * via range requests. */ singleFilePerPlaylist?: boolean; /** * If `true`, the muxer will be in "live mode", continuously emitting updated playlists as new segments are created. * The master playlist will be emitted as soon as all playlists have been emitted at least once, and will continue * to be emitted each time a segment is finalized to further refine the accuracy of the `BANDWIDTH` attribute. * * When `false` (the default), all playlists will only be emitted once, upon output finalization. */ live?: boolean; /** * When in live mode, this controls the maximum number of segments contained in each playlist. Defaults to * `Infinity`, meaning playlists continually grow in size. */ maxLiveSegmentCount?: number; /** * Returns the file path for a given media playlist. If the returned path is relative, it is relative to the root * path. * * Defaults to `'playlist-{n}.m3u8'`, where `n` is the 1-based index of the media playlist in the master playlist. */ getPlaylistPath?: (info: HlsOutputPlaylistInfo) => MaybePromise; /** * Returns the file path for a given media segment. If the returned path is relative, it is relative to the path * of the containing playlist. * * Defaults to `'segment-{n}-{k}{ext}'`, where `n` is the 1-based index of the containing media playlist in the * master playlist, `k` is the 1-based index of the segment in its playlist, and `ext` is the file extension of the * segment format (including the leading dot). * * If {@link HlsOutputFormatOptions.singleFilePerPlaylist} is true, it defaults to `'segments-{n}{ext}'` instead. */ getSegmentPath?: (info: HlsOutputSegmentInfo) => MaybePromise; /** * Returns the file path for a given media init segment. If the returned path is relative, it is relative to the * path of the containing playlist. * * Only necessary for segment formats that require an init file, such as {@link CmafOutputFormat}. * * Defaults to `'init-{n}{ext}'`, where `n` is the 1-based index of the containing media playlist in the master * playlist and `ext` is the file extension of the segment format (including the leading dot). */ getInitPath?: (info: HlsOutputPlaylistInfo) => MaybePromise; /** Called whenever the master playlist is written. */ onMaster?: (content: string) => unknown; /** Called whenever a media playlist is written. */ onPlaylist?: (content: string, info: HlsOutputPlaylistInfo) => unknown; /** * Called whenever a media segment has been fully written. In single-file mode, this function will only be called * once when the playlist is finalized. */ onSegment?: (target: Target, info: HlsOutputSegmentInfo) => unknown; /** * Called when a media playlist is initialized, before any segments have been written. In single-file mode, this * function is never called. */ onInit?: (target: Target, info: HlsOutputPlaylistInfo) => unknown; /** * Called when a media segment is removed from the start of a media playlist due to * {@link HlsOutputFormatOptions.maxLiveSegmentCount}. Will not be called when * {@link HlsOutputFormatOptions.singleFilePerPlaylist} is `true`. */ onSegmentPopped?: (path: string, info: HlsOutputSegmentInfo) => unknown; }; /** * HTTP Live Streaming (HLS) output format. HLS media is represented by a set of .m3u8 playlist files and media segment * files, meaning this format writes out multiple files, requiring the use of a _pathed Output_ * ({@link OutputOptions.target} must be a {@link PathedTarget}). * * This output format creates the following files: * - A master playlist .m3u8 file, containing the list of available playlists. A master playlist is always emitted, * written to the root path. * - One .m3u8 file for each playlist, each containing a list of media segments. * - Many media segments, containing the actual media data. * * To emit media playlists that use the `#EXT-X-PROGRAM-DATE-TIME` tag to map segment timestamps to real-world time, * set {@link BaseTrackMetadata.isRelativeToUnixEpoch} to `true` for all tracks. * * @group Output formats * @public */ export class HlsOutputFormat extends OutputFormat { /** @internal */ _options: HlsOutputFormatOptions; /** Creates a new {@link HlsOutputFormat} configured with the specified `options`. */ constructor(options: HlsOutputFormatOptions) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if ( !(options.segmentFormat instanceof OutputFormat) && ( !Array.isArray(options.segmentFormat) || options.segmentFormat.length === 0 || !options.segmentFormat.every(format => format instanceof OutputFormat) ) ) { throw new TypeError( 'options.segmentFormat must be an OutputFormat or a non-empty array of OutputFormat instances.', ); } if ( options.targetDuration !== undefined && (typeof options.targetDuration !== 'number' || options.targetDuration <= 0) ) { throw new TypeError('options.targetDuration, when provided, must be a positive number.'); } if (options.singleFilePerPlaylist !== undefined && typeof options.singleFilePerPlaylist !== 'boolean') { throw new TypeError('options.singleFilePerPlaylist, when provided, must be a boolean.'); } if (options.live !== undefined && typeof options.live !== 'boolean') { throw new TypeError('options.live, when provided, must be a boolean.'); } if ( options.maxLiveSegmentCount !== undefined && (typeof options.maxLiveSegmentCount !== 'number' || options.maxLiveSegmentCount < 1 || (Number.isFinite(options.maxLiveSegmentCount) && !Number.isInteger(options.maxLiveSegmentCount))) ) { throw new TypeError('options.maxLiveSegmentCount, when provided, must be a positive integer or Infinity.'); } if (options.getPlaylistPath !== undefined && typeof options.getPlaylistPath !== 'function') { throw new TypeError('options.getPlaylistPath, when provided, must be a function.'); } if (options.getSegmentPath !== undefined && typeof options.getSegmentPath !== 'function') { throw new TypeError('options.getSegmentPath, when provided, must be a function.'); } if (options.getInitPath !== undefined && typeof options.getInitPath !== 'function') { throw new TypeError('options.getInitPath, when provided, must be a function.'); } if (options.onMaster !== undefined && typeof options.onMaster !== 'function') { throw new TypeError('options.onMaster, when provided, must be a function.'); } if (options.onPlaylist !== undefined && typeof options.onPlaylist !== 'function') { throw new TypeError('options.onPlaylist, when provided, must be a function.'); } if (options.onSegment !== undefined && typeof options.onSegment !== 'function') { throw new TypeError('options.onSegment, when provided, must be a function.'); } if (options.onInit !== undefined && typeof options.onInit !== 'function') { throw new TypeError('options.onInit, when provided, must be a function.'); } if (options.onSegmentPopped !== undefined && typeof options.onSegmentPopped !== 'function') { throw new TypeError('options.onSegmentPopped, when provided, must be a function.'); } super(); this._options = options; } /** @internal */ _createMuxer(output: Output): Muxer { return new HlsMuxer(output, this); } /** @internal */ get _name() { return 'HTTP Live Streaming (HLS)'; } get fileExtension() { return '.m3u8'; } get mimeType() { return HLS_MIME_TYPE; } getSupportedCodecs(): MediaCodec[] { const uniqueCodecs = new Set(toArray(this._options.segmentFormat).flatMap(x => x.getSupportedCodecs())); return [...uniqueCodecs]; } getSupportedTrackCounts(): TrackCountLimits { let supportsVideo = false; let supportsAudio = false; let supportsSubtitle = false; for (const format of toArray(this._options.segmentFormat)) { const trackCounts = format.getSupportedTrackCounts(); supportsVideo ||= trackCounts.video.max > 0; supportsAudio ||= trackCounts.audio.max > 0; supportsSubtitle ||= trackCounts.subtitle.max > 0; } return { video: { min: 0, max: supportsVideo ? Infinity : 0 }, audio: { min: 0, max: supportsAudio ? Infinity : 0 }, subtitle: { min: 0, max: 0 }, // Currently disabled total: { min: 1, max: Infinity }, }; } get supportsVideoRotationMetadata(): boolean { return toArray(this._options.segmentFormat).some(format => format.supportsVideoRotationMetadata); } get supportsTimestampedMediaData(): boolean { return true; // I guess?? } /** @internal */ // eslint-disable-next-line @typescript-eslint/no-unused-vars override _codecUnsupportedHint(codec: MediaCodec): string { return ` Using different segment formats may grant support for this codec.`; } } ===== src/muxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AsyncMutex } from './misc'; import { Output, OutputAudioTrack, OutputSubtitleTrack, OutputTrack, OutputVideoTrack } from './output'; import { EncodedPacket } from './packet'; import { SubtitleCue, SubtitleMetadata } from './subtitles'; export abstract class Muxer { output: Output; mutex = new AsyncMutex(); constructor(output: Output) { this.output = output; } abstract start(): Promise; abstract getMimeType(): Promise; abstract addEncodedVideoPacket( track: OutputVideoTrack, packet: EncodedPacket, meta?: EncodedVideoChunkMetadata ): Promise; abstract addEncodedAudioPacket( track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata ): Promise; abstract addSubtitleCue(track: OutputSubtitleTrack, cue: SubtitleCue, meta?: SubtitleMetadata): Promise; abstract finalize(): Promise; // eslint-disable-next-line @typescript-eslint/no-unused-vars onTrackClose(track: OutputTrack) {} private trackTimestampInfo = new WeakMap(); protected validateTimestamp(track: OutputTrack, timestampInSeconds: number, isKeyPacket: boolean) { if (timestampInSeconds < 0) { throw new Error(`Timestamps must be non-negative (got ${timestampInSeconds}s).`); } let timestampInfo = this.trackTimestampInfo.get(track); if (!timestampInfo) { if (!isKeyPacket) { throw new Error('First packet must be a key packet.'); } timestampInfo = { maxTimestamp: timestampInSeconds, maxTimestampBeforeLastKeyPacket: null, }; this.trackTimestampInfo.set(track, timestampInfo); } else { if (isKeyPacket) { timestampInfo.maxTimestampBeforeLastKeyPacket = timestampInfo.maxTimestamp; } if ( timestampInfo.maxTimestampBeforeLastKeyPacket !== null && timestampInSeconds < timestampInfo.maxTimestampBeforeLastKeyPacket ) { throw new Error( `Timestamps cannot be smaller than the largest timestamp of the previous GOP (a GOP begins with a` + ` key packet and ends right before the next key packet). Got ${timestampInSeconds}s, but largest` + ` timestamp is ${timestampInfo.maxTimestampBeforeLastKeyPacket}s.`, ); } timestampInfo.maxTimestamp = Math.max(timestampInfo.maxTimestamp, timestampInSeconds); } } } ===== src/reader.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { InputDisposedError } from './input'; import { assert, clamp, getUint24, MaybePromise, textDecoder, toDataView } from './misc'; import { DEFAULT_MAX_READ_POSITION, DEFAULT_MIN_READ_POSITION, Source } from './source'; export class Reader { constructor(public source: Source) {} get fileSize(): number | null { const size = this.source._getFileSize(); if (size === undefined) { throw new Error('Reading file size too early; read required first.'); } return size; } get fileSizeNonStrict() { return this.source._getFileSize() ?? null; } requestSlice(start: number, length: number): MaybePromise { if (this.source._disposed) { throw new InputDisposedError(); } if (start < 0) { return null; } if (this.fileSizeNonStrict !== null && start + length > this.fileSizeNonStrict) { return null; } if (length === 0) { const buffer = new Uint8Array(0); return new FileSlice(buffer, toDataView(buffer), 0, start, start); } const end = start + length; const result = this.source._read(start, end, DEFAULT_MIN_READ_POSITION, DEFAULT_MAX_READ_POSITION); if (result instanceof Promise) { return result.then((x) => { if (!x) { return null; } return new FileSlice(x.bytes, x.view, x.offset, start, end); }); } else { if (!result) { return null; } return new FileSlice(result.bytes, result.view, result.offset, start, end); } } requestSliceRange(start: number, minLength: number, maxLength: number): MaybePromise { if (this.source._disposed) { throw new InputDisposedError(); } if (start < 0) { return null; } if (this.fileSizeNonStrict !== null) { return this.requestSlice( start, clamp(this.fileSizeNonStrict - start, minLength, maxLength), ); } else { const promisedAttempt = this.requestSlice(start, maxLength); const handleAttempt = (attempt: FileSlice | null) => { if (attempt) { return attempt; } // The slice couldn't fit, meaning we must know the file size now assert(this.fileSizeNonStrict !== null); return this.requestSlice( start, clamp(this.fileSizeNonStrict - start, minLength, maxLength), ); }; if (promisedAttempt instanceof Promise) { return promisedAttempt.then(handleAttempt); } else { return handleAttempt(promisedAttempt); } } } requestEntireFile(): MaybePromise { if (this.fileSizeNonStrict !== null) { return this.requestSlice(0, this.fileSizeNonStrict); } const CHUNK_SIZE = 1024; return (async () => { const chunks: Uint8Array[] = []; let currentSize = 0; while (true) { if (chunks.length === 1 && this.fileSizeNonStrict !== null) { // It only took one read to get to know the whole file size return this.requestSlice(0, this.fileSizeNonStrict); } let slice = this.requestSliceRange(currentSize, 0, CHUNK_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice || slice.length === 0) { break; } const chunk = readBytes(slice, slice.length); chunks.push(chunk); currentSize += slice.length; } const joined = new Uint8Array(currentSize); let offset = 0; for (const chunk of chunks) { joined.set(chunk, offset); offset += chunk.length; } return new FileSlice(joined, toDataView(joined), 0, 0, currentSize); })(); } } export class FileSlice { /** The current position in the backing buffer. Do not modify directly, prefer `.skip()` instead. */ bufferPos: number; constructor( /** The underlying bytes backing this slice. Avoid using this directly and prefer reader functions instead. */ public readonly bytes: Uint8Array, /** A view into the bytes backing this slice. Avoid using this directly and prefer reader functions instead. */ public readonly view: DataView, /** The offset in "file bytes" at which `bytes` begins in the file. */ private readonly offset: number, /** The offset in "file bytes" where this slice begins. */ public readonly start: number, /** The offset in "file bytes" where this slice ends (exclusive). */ public readonly end: number, ) { this.bufferPos = start - offset; } static tempFromBytes(bytes: Uint8Array) { return new FileSlice( bytes, toDataView(bytes), 0, 0, bytes.length, ); } get length() { return this.end - this.start; } get filePos() { return this.offset + this.bufferPos; } set filePos(value: number) { this.bufferPos = value - this.offset; } /** The number of bytes left from the current pos to the end of the slice. */ get remainingLength() { return Math.max(this.end - this.filePos, 0); } skip(byteCount: number) { this.bufferPos += byteCount; } /** Creates a new subslice of this slice whose byte range must be contained within this slice. */ slice(filePos: number, length = this.end - filePos) { if (filePos < this.start || filePos + length > this.end) { throw new RangeError('Slicing outside of original slice.'); } return new FileSlice( this.bytes, this.view, this.offset, filePos, filePos + length, ); } } const checkIsInRange = (slice: FileSlice, bytesToRead: number) => { if (slice.filePos < slice.start || slice.filePos + bytesToRead > slice.end) { throw new RangeError( `Tried reading [${slice.filePos}, ${slice.filePos + bytesToRead}), but slice is` + ` [${slice.start}, ${slice.end}). This is likely an internal error, please report it alongside the file` + ` that caused it.`, ); } }; export const readBytes = (slice: FileSlice, length: number) => { checkIsInRange(slice, length); const bytes = slice.bytes.subarray(slice.bufferPos, slice.bufferPos + length); slice.bufferPos += length; return bytes; }; export const readU8 = (slice: FileSlice) => { checkIsInRange(slice, 1); return slice.view.getUint8(slice.bufferPos++); }; export const readU16 = (slice: FileSlice, littleEndian: boolean) => { checkIsInRange(slice, 2); const value = slice.view.getUint16(slice.bufferPos, littleEndian); slice.bufferPos += 2; return value; }; export const readU16Be = (slice: FileSlice) => { checkIsInRange(slice, 2); const value = slice.view.getUint16(slice.bufferPos, false); slice.bufferPos += 2; return value; }; export const readU24Be = (slice: FileSlice) => { checkIsInRange(slice, 3); const value = getUint24(slice.view, slice.bufferPos, false); slice.bufferPos += 3; return value; }; export const readI16Be = (slice: FileSlice) => { checkIsInRange(slice, 2); const value = slice.view.getInt16(slice.bufferPos, false); slice.bufferPos += 2; return value; }; export const readU32 = (slice: FileSlice, littleEndian: boolean) => { checkIsInRange(slice, 4); const value = slice.view.getUint32(slice.bufferPos, littleEndian); slice.bufferPos += 4; return value; }; export const readU32Be = (slice: FileSlice) => { checkIsInRange(slice, 4); const value = slice.view.getUint32(slice.bufferPos, false); slice.bufferPos += 4; return value; }; export const readU32Le = (slice: FileSlice) => { checkIsInRange(slice, 4); const value = slice.view.getUint32(slice.bufferPos, true); slice.bufferPos += 4; return value; }; export const readI32Be = (slice: FileSlice) => { checkIsInRange(slice, 4); const value = slice.view.getInt32(slice.bufferPos, false); slice.bufferPos += 4; return value; }; export const readI32Le = (slice: FileSlice) => { checkIsInRange(slice, 4); const value = slice.view.getInt32(slice.bufferPos, true); slice.bufferPos += 4; return value; }; export const readU64 = (slice: FileSlice, littleEndian: boolean) => { let low: number; let high: number; if (littleEndian) { low = readU32(slice, true); high = readU32(slice, true); } else { high = readU32(slice, false); low = readU32(slice, false); } return high * 0x100000000 + low; }; export const readU64Be = (slice: FileSlice) => { const high = readU32Be(slice); const low = readU32Be(slice); return high * 0x100000000 + low; }; export const readI64Be = (slice: FileSlice) => { const high = readI32Be(slice); const low = readU32Be(slice); return high * 0x100000000 + low; }; export const readI64Le = (slice: FileSlice) => { const low = readU32Le(slice); const high = readI32Le(slice); return high * 0x100000000 + low; }; export const readF32Be = (slice: FileSlice) => { checkIsInRange(slice, 4); const value = slice.view.getFloat32(slice.bufferPos, false); slice.bufferPos += 4; return value; }; export const readF64Be = (slice: FileSlice) => { checkIsInRange(slice, 8); const value = slice.view.getFloat64(slice.bufferPos, false); slice.bufferPos += 8; return value; }; export const readAscii = (slice: FileSlice, length: number) => { checkIsInRange(slice, length); let str = ''; for (let i = 0; i < length; i++) { str += String.fromCharCode(slice.bytes[slice.bufferPos++]!); } return str; }; export const readAllLines = (slice: FileSlice, length: number, options?: { ignore?: (line: string) => boolean; }) => { const text = textDecoder.decode(readBytes(slice, length)); const lines = text.split('\n') .map(x => x.trim()) .filter(x => x.length > 0 && !options?.ignore?.(x)); return lines; }; ===== src/flac/flac-demuxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { FlacBlockType, readVorbisComments } from '../codec-data'; import { Demuxer } from '../demuxer'; import { Input } from '../input'; import { InputAudioTrackBacking } from '../input-track'; import { PacketRetrievalOptions } from '../media-sink'; import { assert, AsyncMutex, binarySearchLessOrEqual, textDecoder, UNDETERMINED_LANGUAGE, } from '../misc'; import { EncodedPacket, PLACEHOLDER_DATA } from '../packet'; import { FileSlice, readBytes, Reader, readU24Be, readU32Be, readU8, } from '../reader'; import { DEFAULT_TRACK_DISPOSITION, MetadataTags } from '../metadata'; import { calculateCrc8, readBlockSize, getBlockSizeOrUncommon, readCodedNumber, readSampleRate, getSampleRateOrUncommon, } from './flac-misc'; import { Bitstream } from '../../shared/bitstream'; import { ID3_V2_HEADER_SIZE, parseId3V2Tag, readId3V2Header } from '../id3'; type FlacAudioInfo = { numberOfChannels: number; sampleRate: number; totalSamples: number; minimumBlockSize: number; maximumBlockSize: number; minimumFrameSize: number; maximumFrameSize: number; description: Uint8Array; }; type Sample = { blockOffset: number; blockSize: number; byteOffset: number; byteSize: number; }; type NextFlacFrameResult = { num: number; blockSize: number; sampleRate: number; size: number; isLastFrame: boolean; }; export class FlacDemuxer extends Demuxer { reader: Reader; loadedSamples: Sample[] = []; // All samples from the start of the file to lastLoadedPos metadataPromise: Promise | null = null; trackBacking: FlacAudioTrackBacking | null = null; metadataTags: MetadataTags = {}; audioInfo: FlacAudioInfo | null = null; lastLoadedPos: number | null = null; blockingBit: number | null = null; readingMutex = new AsyncMutex(); lastSampleLoaded = false; constructor(input: Input) { super(input); this.reader = input._reader; } override async getMetadataTags(): Promise { await this.readMetadata(); return this.metadataTags; } async getTrackBackings() { await this.readMetadata(); assert(this.trackBacking); return [this.trackBacking]; } async getMimeType() { return 'audio/flac'; } async readMetadata() { return (this.metadataPromise ??= (async () => { // Read all ID3v2 tags at the start of the file let currentPos = 0; while (true) { let headerSlice = this.reader.requestSlice(currentPos, ID3_V2_HEADER_SIZE); if (headerSlice instanceof Promise) headerSlice = await headerSlice; if (!headerSlice) { this.lastSampleLoaded = true; return; } const id3V2Header = readId3V2Header(headerSlice); if (!id3V2Header) { break; } let contentSlice = this.reader.requestSlice(headerSlice.filePos, id3V2Header.size); if (contentSlice instanceof Promise) contentSlice = await contentSlice; assert(contentSlice); parseId3V2Tag(contentSlice, id3V2Header, this.metadataTags); currentPos = headerSlice.filePos + id3V2Header.size; } currentPos += 4; // Skip 'fLaC' while ( this.reader.fileSize === null || currentPos < this.reader.fileSize ) { let sizeSlice = this.reader.requestSlice(currentPos, 4); if (sizeSlice instanceof Promise) sizeSlice = await sizeSlice; currentPos += 4; if (sizeSlice === null) { throw new Error( `Metadata block at position ${currentPos} is too small! Corrupted file.`, ); } assert(sizeSlice); const byte = readU8(sizeSlice); // first bit: isLastMetadata, remaining 7 bits: metaBlockType const size = readU24Be(sizeSlice); const isLastMetadata = (byte & 0x80) !== 0; const metaBlockType = byte & 0x7f; switch (metaBlockType) { case FlacBlockType.STREAMINFO: { // Parse streaminfo block // https://www.rfc-editor.org/rfc/rfc9639.html#section-8.2 let streamInfoBlock = this.reader.requestSlice( currentPos, size, ); if (streamInfoBlock instanceof Promise) streamInfoBlock = await streamInfoBlock; assert(streamInfoBlock); if (streamInfoBlock === null) { throw new Error( `StreamInfo block at position ${currentPos} is too small! Corrupted file.`, ); } const streamInfoBytes = readBytes(streamInfoBlock, 34); const bitstream = new Bitstream(streamInfoBytes); const minimumBlockSize = bitstream.readBits(16); const maximumBlockSize = bitstream.readBits(16); const minimumFrameSize = bitstream.readBits(24); const maximumFrameSize = bitstream.readBits(24); const sampleRate = bitstream.readBits(20); const numberOfChannels = bitstream.readBits(3) + 1; bitstream.readBits(5); // bitsPerSample - 1 const totalSamples = bitstream.readBits(36); // https://www.w3.org/TR/webcodecs-flac-codec-registration/#audiodecoderconfig-description // description is required, and has to be the following: // 1. The bytes 0x66 0x4C 0x61 0x43 ("fLaC" in ASCII) // 2. A metadata block (called the STREAMINFO block) as described in section 7 of [FLAC] // 3. Optionaly (sic) other metadata blocks, that are not used by the specification bitstream.skipBits(16 * 8); // md5 hash const description = new Uint8Array(42); // 1. "fLaC" description.set(new Uint8Array([0x66, 0x4c, 0x61, 0x43]), 0); // 2. STREAMINFO block description.set(new Uint8Array([128, 0, 0, 34]), 4); // 3. Other metadata blocks description.set(streamInfoBytes, 8); this.audioInfo = { numberOfChannels, sampleRate, totalSamples, minimumBlockSize, maximumBlockSize, minimumFrameSize, maximumFrameSize, description, }; this.trackBacking = new FlacAudioTrackBacking(this); break; } case FlacBlockType.VORBIS_COMMENT: { // Parse vorbis comment block // https://www.rfc-editor.org/rfc/rfc9639.html#name-vorbis-comment let vorbisCommentBlock = this.reader.requestSlice( currentPos, size, ); if (vorbisCommentBlock instanceof Promise) vorbisCommentBlock = await vorbisCommentBlock; assert(vorbisCommentBlock); readVorbisComments( readBytes(vorbisCommentBlock, size), this.metadataTags, ); break; } case FlacBlockType.PICTURE: { // Parse picture block // https://www.rfc-editor.org/rfc/rfc9639.html#name-picture let pictureBlock = this.reader.requestSlice( currentPos, size, ); if (pictureBlock instanceof Promise) pictureBlock = await pictureBlock; assert(pictureBlock); const pictureType = readU32Be(pictureBlock); const mediaTypeLength = readU32Be(pictureBlock); const mediaType = textDecoder.decode( readBytes(pictureBlock, mediaTypeLength), ); const descriptionLength = readU32Be(pictureBlock); const description = textDecoder.decode( readBytes(pictureBlock, descriptionLength), ); pictureBlock.skip(4 + 4 + 4 + 4); // Skip width, height, color depth, number of indexed colors const dataLength = readU32Be(pictureBlock); const data = readBytes(pictureBlock, dataLength); this.metadataTags.images ??= []; this.metadataTags.images.push({ data, mimeType: mediaType, // https://www.rfc-editor.org/rfc/rfc9639.html#table13 kind: pictureType === 3 ? 'coverFront' : pictureType === 4 ? 'coverBack' : 'unknown', description, }); break; } default: break; } currentPos += size; if (isLastMetadata) { this.lastLoadedPos = currentPos; break; } } if (!this.audioInfo) { throw new Error('Missing STREAMINFO metadata block! Corrupted FLAC file.'); } })()); } async readNextFlacFrame({ startPos, isFirstPacket, }: { startPos: number; isFirstPacket: boolean; }): Promise { assert(this.audioInfo); // we expect that there are at least `minimumFrameSize` bytes left in the file // Ideally we also want to validate the next header is valid // to throw out an accidential sync word // The shortest valid FLAC header I can think of, based off the code // of readFlacFrameHeader: // 4 bytes used for bitstream from syncword to bit depth // 1 byte coded number // (uncommon values, no bytes read) // 1 byte crc // --> 6 bytes const minimumHeaderLength = 6; // If we read everything in readFlacFrameHeader, we read 16 bytes const maximumHeaderLength = 16; // The shortest valid FLAC frame per RFC 9639: // 6 bytes header (see minimumHeaderLength above) // 2 bytes subframe (constant subframe with minimum bit depth, // padded to byte boundary) // 2 bytes footer (CRC-16) // --> 10 bytes const minimumFrameLength = 10; // The longest valid FLAC frame per RFC 9639: // https://www.rfc-editor.org/rfc/rfc9639.html#name-prediction // https://www.rfc-editor.org/rfc/rfc9639.html#name-frame-structure // maximumBlockSize * numberOfChannels * 4 bytes (max 32 bps verbatim) // + 16 bytes header (see maximumHeaderSize above) // + 2 bytes footer (CRC-16) const maximumFrameLength = this.audioInfo.maximumBlockSize * this.audioInfo.numberOfChannels * 4 + maximumHeaderLength + 2; // Per RFC 9639, a value of 0 means "unknown" for frame sizes. const effectiveMinFrameSize = this.audioInfo.minimumFrameSize || minimumFrameLength; const effectiveMaxFrameSize = this.audioInfo.maximumFrameSize || maximumFrameLength; const maximumSliceLength = effectiveMaxFrameSize + maximumHeaderLength; const slice = await this.reader.requestSliceRange( startPos, maximumHeaderLength, maximumSliceLength, ); if (!slice) { return null; } const frameHeader = this.readFlacFrameHeader({ slice, isFirstPacket: isFirstPacket, }); if (!frameHeader) { return null; } // We don't know exactly how long the packet is, we only know the `minimumFrameSize` and `maximumFrameSize` // The packet is over if the next 2 bytes are the sync word followed by a valid header // or the end of the file is reached // The next sync word is expected at earliest when `minimumFrameSize` is reached, // we can skip over anything before that slice.filePos = startPos + effectiveMinFrameSize; while (true) { // Reached end of the file, packet is over if (slice.filePos > slice.end - minimumHeaderLength) { return { num: frameHeader.num, blockSize: frameHeader.blockSize, sampleRate: frameHeader.sampleRate, size: slice.end - startPos, isLastFrame: true, }; } const nextByte = readU8(slice); if (nextByte === 0xff) { const positionBeforeReading = slice.filePos; const byteAfterNextByte = readU8(slice); const expected = this.blockingBit === 1 ? 0b1111_1001 : 0b1111_1000; if (byteAfterNextByte !== expected) { slice.filePos = positionBeforeReading; continue; } slice.skip(-2); const lengthIfNextFlacFrameHeaderIsLegit = slice.filePos - startPos; const nextFrameHeader = this.readFlacFrameHeader({ slice, isFirstPacket: false, }); if (!nextFrameHeader) { slice.filePos = positionBeforeReading; continue; } // Ensure the frameOrSampleNum is consecutive. // https://github.com/Vanilagy/mediabunny/issues/194 if (this.blockingBit === 0) { // Case A: If the stream is fixed block size, this is the frame number, which increments by 1 if (nextFrameHeader.num - frameHeader.num !== 1) { slice.filePos = positionBeforeReading; continue; } } else { // Case B: If the stream is variable block size, this is the sample number, which increments by // amount of samples in a frame. if (nextFrameHeader.num - frameHeader.num !== frameHeader.blockSize) { slice.filePos = positionBeforeReading; continue; } } return { num: frameHeader.num, blockSize: frameHeader.blockSize, sampleRate: frameHeader.sampleRate, size: lengthIfNextFlacFrameHeaderIsLegit, isLastFrame: false, }; } } } readFlacFrameHeader({ slice, isFirstPacket, }: { slice: FileSlice; isFirstPacket: boolean; }) { // In this function, generally it is not safe to throw errors. // We might end up here because we stumbled upon a syncword, // but the data might not actually be a FLAC frame, it might be random bitstream // data, in that case we should return null and continue. const startOffset = slice.filePos; // https://www.rfc-editor.org/rfc/rfc9639.html#section-9.1 // Each frame MUST start on a byte boundary and start with the 15-bit frame // sync code 0b111111111111100. Following the sync code is the blocking strategy // bit, which MUST NOT change during the audio stream. const bytes = readBytes(slice, 4); const bitstream = new Bitstream(bytes); const bits = bitstream.readBits(15); if (bits !== 0b111111111111100) { // This cannot be a valid FLAC frame, must start with the syncword return null; } if (this.blockingBit === null) { assert(isFirstPacket); const newBlockingBit = bitstream.readBits(1); this.blockingBit = newBlockingBit; } else if (this.blockingBit === 1) { assert(!isFirstPacket); const newBlockingBit = bitstream.readBits(1); if (newBlockingBit !== 1) { // This cannot be a valid FLAC frame, expected 1 but got 0 return null; } } else if (this.blockingBit === 0) { assert(!isFirstPacket); const newBlockingBit = bitstream.readBits(1); if (newBlockingBit !== 0) { // This cannot be a valid FLAC frame, expected 0 but got 1 return null; } } else { throw new Error('Invalid blocking bit'); } const blockSizeOrUncommon = getBlockSizeOrUncommon(bitstream.readBits(4)); if (!blockSizeOrUncommon) { // This cannot be a valid FLAC frame, the syncword was just coincidental return null; } assert(this.audioInfo); const sampleRateOrUncommon = getSampleRateOrUncommon( bitstream.readBits(4), this.audioInfo.sampleRate, ); if (!sampleRateOrUncommon) { // This cannot be a valid FLAC frame, the syncword was just coincidental return null; } bitstream.readBits(4); // channel count bitstream.readBits(3); // bit depth const reservedZero = bitstream.readBits(1); // reserved zero if (reservedZero !== 0) { // This cannot be a valid FLAC frame, the syncword was just coincidental return null; } const num = readCodedNumber(slice); const blockSize = readBlockSize(slice, blockSizeOrUncommon); const sampleRate = readSampleRate(slice, sampleRateOrUncommon); if (sampleRate === null) { // This cannot be a valid FLAC frame, the syncword was just coincidental return null; } if (sampleRate !== this.audioInfo.sampleRate) { // This cannot be a valid FLAC frame, the sample rate is not the same as in the stream info return null; } const size = slice.filePos - startOffset; const crc = readU8(slice); slice.skip(-size); slice.skip(-1); const crcCalculated = calculateCrc8(readBytes(slice, size)); if (crc !== crcCalculated) { // Maybe this wasn't a FLAC frame at all, the syncword was just coincidentally // in the bitstream return null; } return { num, blockSize, sampleRate }; } async advanceReader() { await this.readMetadata(); assert(this.lastLoadedPos !== null); assert(this.audioInfo); const startPos = this.lastLoadedPos; const frame = await this.readNextFlacFrame({ startPos, isFirstPacket: this.loadedSamples.length === 0, }); if (!frame) { // Unexpected case, failed to read next FLAC frame // handling gracefully this.lastSampleLoaded = true; return; } const lastSample = this.loadedSamples[this.loadedSamples.length - 1]; const blockOffset = lastSample ? lastSample.blockOffset + lastSample.blockSize : 0; const sample: Sample = { blockOffset, blockSize: frame.blockSize, byteOffset: startPos, byteSize: frame.size, }; this.lastLoadedPos = this.lastLoadedPos + frame.size; this.loadedSamples.push(sample); if (frame.isLastFrame) { this.lastSampleLoaded = true; return; } } } class FlacAudioTrackBacking implements InputAudioTrackBacking { constructor(public demuxer: FlacDemuxer) {} getType() { return 'audio' as const; } getId() { return 1; } getNumber() { return 1; } getCodec() { return 'flac' as const; } getInternalCodecId(): string | number | Uint8Array | null { return null; } getNumberOfChannels() { assert(this.demuxer.audioInfo); return this.demuxer.audioInfo.numberOfChannels; } getSampleRate() { assert(this.demuxer.audioInfo); return this.demuxer.audioInfo.sampleRate; } getName(): string | null { return null; } getLanguageCode() { return UNDETERMINED_LANGUAGE; } getTimeResolution() { assert(this.demuxer.audioInfo); return this.demuxer.audioInfo.sampleRate; } isRelativeToUnixEpoch() { return false; } getUnixTimeForTimestamp() { return null; } getPairingMask() { return 1n; } getBitrate() { return null; } getAverageBitrate() { return null; } async getDurationFromMetadata() { assert(this.demuxer.audioInfo); if (this.demuxer.audioInfo.totalSamples === 0) { return null; } return this.demuxer.audioInfo.totalSamples / this.demuxer.audioInfo.sampleRate; } async getLiveRefreshInterval() { return null; } getDisposition() { return { ...DEFAULT_TRACK_DISPOSITION, }; } async getDecoderConfig(): Promise { assert(this.demuxer.audioInfo); return { codec: 'flac' as const, numberOfChannels: this.demuxer.audioInfo.numberOfChannels, sampleRate: this.demuxer.audioInfo.sampleRate, description: this.demuxer.audioInfo.description, }; } async getPacket( timestamp: number, options: PacketRetrievalOptions, ): Promise { assert(this.demuxer.audioInfo); if (timestamp < 0) { return null; } const release = await this.demuxer.readingMutex.acquire(); try { while (true) { const packetIndex = binarySearchLessOrEqual( this.demuxer.loadedSamples, timestamp, x => x.blockOffset / this.demuxer.audioInfo!.sampleRate, ); if (packetIndex === -1) { await this.demuxer.advanceReader(); continue; } const packet = this.demuxer.loadedSamples[packetIndex]!; const sampleTimestamp = packet.blockOffset / this.demuxer.audioInfo.sampleRate; const sampleDuration = packet.blockSize / this.demuxer.audioInfo.sampleRate; if (sampleTimestamp + sampleDuration <= timestamp) { if (this.demuxer.lastSampleLoaded) { return this.getPacketAtIndex( this.demuxer.loadedSamples.length - 1, options, ); } await this.demuxer.advanceReader(); continue; } return this.getPacketAtIndex(packetIndex, options); } } finally { release(); } } async getNextPacket( packet: EncodedPacket, options: PacketRetrievalOptions, ): Promise { const release = await this.demuxer.readingMutex.acquire(); try { const nextIndex = packet.sequenceNumber + 1; if ( this.demuxer.lastSampleLoaded && nextIndex >= this.demuxer.loadedSamples.length ) { return null; } // Ensure the next sample exists while ( nextIndex >= this.demuxer.loadedSamples.length && !this.demuxer.lastSampleLoaded ) { await this.demuxer.advanceReader(); } return this.getPacketAtIndex(nextIndex, options); } finally { release(); } } getKeyPacket( timestamp: number, options: PacketRetrievalOptions, ): Promise { return this.getPacket(timestamp, options); } getNextKeyPacket( packet: EncodedPacket, options: PacketRetrievalOptions, ): Promise { return this.getNextPacket(packet, options); } async getPacketAtIndex( sampleIndex: number, options: PacketRetrievalOptions, ): Promise { const rawSample = this.demuxer.loadedSamples[sampleIndex]; if (!rawSample) { return null; } let data: Uint8Array; if (options.metadataOnly) { data = PLACEHOLDER_DATA; } else { let slice = this.demuxer.reader.requestSlice( rawSample.byteOffset, rawSample.byteSize, ); if (slice instanceof Promise) slice = await slice; if (!slice) { return null; // Data didn't fit into the rest of the file } data = readBytes(slice, rawSample.byteSize); } assert(this.demuxer.audioInfo); const timestamp = rawSample.blockOffset / this.demuxer.audioInfo.sampleRate; const duration = rawSample.blockSize / this.demuxer.audioInfo.sampleRate; return new EncodedPacket( data, 'key', timestamp, duration, sampleIndex, rawSample.byteSize, ); } async getFirstPacket( options: PacketRetrievalOptions, ): Promise { // Ensure the next sample exists while ( this.demuxer.loadedSamples.length === 0 && !this.demuxer.lastSampleLoaded ) { await this.demuxer.advanceReader(); } return this.getPacketAtIndex(0, options); } } ===== src/flac/flac-misc.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Bitstream } from '../../shared/bitstream'; import { assert, assertNever } from '../misc'; import { FileSlice, readBytes, readU16Be, readU8 } from '../reader'; type BlockSizeOrUncommon = number | 'uncommon-u16' | 'uncommon-u8'; type SampleRateOrUncommon = | number | 'uncommon-u8' | 'uncommon-u16' | 'uncommon-u16-10'; // https://www.rfc-editor.org/rfc/rfc9639.html#name-block-size-bits export const getBlockSizeOrUncommon = (bits: number): BlockSizeOrUncommon | null => { if (bits === 0b0000) { return null; } else if (bits === 0b0001) { return 192; } else if (bits >= 0b0010 && bits <= 0b0101) { return 144 * 2 ** bits; } else if (bits === 0b0110) { return 'uncommon-u8'; } else if (bits === 0b0111) { return 'uncommon-u16'; } else if (bits >= 0b1000 && bits <= 0b1111) { return 2 ** bits; } else { return null; } }; // https://www.rfc-editor.org/rfc/rfc9639.html#name-sample-rate-bits export const getSampleRateOrUncommon = ( sampleRateBits: number, streamInfoSampleRate: number, ): SampleRateOrUncommon | null => { switch (sampleRateBits) { case 0b0000: return streamInfoSampleRate; case 0b0001: return 88200; case 0b0010: return 176400; case 0b0011: return 192000; case 0b0100: return 8000; case 0b0101: return 16000; case 0b0110: return 22050; case 0b0111: return 24000; case 0b1000: return 32000; case 0b1001: return 44100; case 0b1010: return 48000; case 0b1011: return 96000; case 0b1100: return 'uncommon-u8'; case 0b1101: return 'uncommon-u16'; case 0b1110: return 'uncommon-u16-10'; default: return null; } }; // https://www.rfc-editor.org/rfc/rfc9639.html#name-coded-number export const readCodedNumber = (fileSlice: FileSlice): number => { let ones = 0; const bitstream1 = new Bitstream(readBytes(fileSlice, 1)); while (bitstream1.readBits(1) === 1) { ones++; } if (ones === 0) { return bitstream1.readBits(7); } const bitArray: number[] = []; const extraBytes = ones - 1; const bitstream2 = new Bitstream(readBytes(fileSlice, extraBytes)); const firstByteBits = 8 - ones - 1; for (let i = 0; i < firstByteBits; i++) { bitArray.unshift(bitstream1.readBits(1)); } for (let i = 0; i < extraBytes; i++) { for (let j = 0; j < 8; j++) { const val = bitstream2.readBits(1); if (j < 2) { continue; } bitArray.unshift(val); } } const encoded = bitArray.reduce((acc, bit, index) => { return acc | (bit << index); }, 0); return encoded; }; export const readBlockSize = ( slice: FileSlice, blockSizeBits: BlockSizeOrUncommon, ) => { if (blockSizeBits === 'uncommon-u16') { return readU16Be(slice) + 1; } else if (blockSizeBits === 'uncommon-u8') { return readU8(slice) + 1; } else if (typeof blockSizeBits === 'number') { return blockSizeBits; } else { assertNever(blockSizeBits); assert(false); } }; export const readSampleRate = ( slice: FileSlice, sampleRateOrUncommon: SampleRateOrUncommon, ) => { if (sampleRateOrUncommon === 'uncommon-u16') { return readU16Be(slice); } if (sampleRateOrUncommon === 'uncommon-u16-10') { return readU16Be(slice) * 10; } if (sampleRateOrUncommon === 'uncommon-u8') { return readU8(slice); } if (typeof sampleRateOrUncommon === 'number') { return sampleRateOrUncommon; } return null; }; // https://www.rfc-editor.org/rfc/rfc9639.html#section-9.1.1 export const calculateCrc8 = (data: Uint8Array) => { const polynomial = 0x07; // x^8 + x^2 + x^1 + x^0 let crc = 0x00; // Initialize CRC to 0 for (const byte of data) { crc ^= byte; // XOR byte into least significant byte of crc for (let i = 0; i < 8; i++) { // For each bit in the byte if ((crc & 0x80) !== 0) { // If the leftmost bit (MSB) is set crc = (crc << 1) ^ polynomial; // Shift left and XOR with polynomial } else { crc <<= 1; // Just shift left } crc &= 0xff; // Ensure CRC remains 8-bit } } return crc; }; ===== src/flac/flac-muxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { validateAudioChunkMetadata } from '../codec'; import { createVorbisComments, FlacBlockType } from '../codec-data'; import { assert, textEncoder, toDataView, toUint8Array, } from '../misc'; import { Muxer } from '../muxer'; import { Output, OutputAudioTrack } from '../output'; import { FlacOutputFormat } from '../output-format'; import { EncodedPacket } from '../packet'; import { FileSlice, readBytes } from '../reader'; import { AttachedImage, metadataTagsAreEmpty } from '../metadata'; import { Writer } from '../writer'; import { readBlockSize, getBlockSizeOrUncommon, readCodedNumber, } from './flac-misc'; import { Bitstream } from '../../shared/bitstream'; const FLAC_HEADER = /* #__PURE__ */ new Uint8Array([0x66, 0x4c, 0x61, 0x43]); // 'fLaC' const STREAMINFO_SIZE = 38; const STREAMINFO_BLOCK_SIZE = 34; export class FlacMuxer extends Muxer { private writer!: Writer; private metadataWritten = false; private blockSizes: number[] = []; private frameSizes: number[] = []; private sampleRate: number | null = null; private channels: number | null = null; private bitsPerSample: number | null = null; private format: FlacOutputFormat; constructor(output: Output, format: FlacOutputFormat) { super(output); this.format = format; } async start() { const release = await this.mutex.acquire(); this.writer = await this.output._getRootWriter(!!this.format._options.appendOnly); this.writer.write(FLAC_HEADER); release(); } writeHeader({ bitsPerSample, minimumBlockSize, maximumBlockSize, minimumFrameSize, maximumFrameSize, sampleRate, channels, totalSamples, }: { minimumBlockSize: number; maximumBlockSize: number; minimumFrameSize: number; maximumFrameSize: number; sampleRate: number; channels: number; bitsPerSample: number; totalSamples: number; }) { assert(this.writer.getPos() === 4); const hasMetadata = !metadataTagsAreEmpty(this.output._metadataTags); const headerBitstream = new Bitstream(new Uint8Array(4)); headerBitstream.writeBits(1, Number(!hasMetadata)); // isLastMetadata headerBitstream.writeBits(7, FlacBlockType.STREAMINFO); // metaBlockType = streaminfo headerBitstream.writeBits(24, STREAMINFO_BLOCK_SIZE); // size this.writer.write(headerBitstream.bytes); const contentBitstream = new Bitstream(new Uint8Array(18)); contentBitstream.writeBits(16, minimumBlockSize); contentBitstream.writeBits(16, maximumBlockSize); contentBitstream.writeBits(24, minimumFrameSize); contentBitstream.writeBits(24, maximumFrameSize); contentBitstream.writeBits(20, sampleRate); contentBitstream.writeBits(3, channels - 1); contentBitstream.writeBits(5, bitsPerSample - 1); // Bitstream operations are only safe until 32bit, breaks when using 36 bits // Splitting up into writing 4 0 bits and then 32 bits is safe // This is safe for audio up to (2 ** 32 / 44100 / 3600) -> 27 hours // Not implementing support for more than 32 bits now if (totalSamples >= 2 ** 32) { throw new Error('This muxer only supports writing up to 2 ** 32 samples'); } contentBitstream.writeBits(4, 0); contentBitstream.writeBits(32, totalSamples); this.writer.write(contentBitstream.bytes); // The MD5 hash is calculated from decoded audio data, but we do not have access // to it here. We are allowed to set 0: // "A value of 0 signifies that the value is not known." // https://www.rfc-editor.org/rfc/rfc9639.html#name-streaminfo this.writer.write(new Uint8Array(16)); } writePictureBlock(picture: AttachedImage) { // Header size: // 4 bytes: picture type // 4 bytes: media type length // x bytes: media type // 4 bytes: description length // y bytes: description // 1 bytes: width // 1 bytes: height // 1 bytes: color depth // 1 bytes: number of indexed colors // 4 bytes: picture data length // z bytes: picture data // Total: 20 + x + y + z const headerSize = 32 + picture.mimeType.length + (picture.description?.length ?? 0) + picture.data.length; const header = new Uint8Array(headerSize); let offset = 0; const dataView = toDataView(header); dataView.setUint32( offset, picture.kind === 'coverFront' ? 3 : picture.kind === 'coverBack' ? 4 : 0, ); offset += 4; dataView.setUint32(offset, picture.mimeType.length); offset += 4; header.set(textEncoder.encode(picture.mimeType), 8); offset += picture.mimeType.length; dataView.setUint32(offset, picture.description?.length ?? 0); offset += 4; header.set(textEncoder.encode(picture.description ?? ''), offset); offset += picture.description?.length ?? 0; offset += 4 + 4 + 4 + 4; // setting width, height, color depth, number of indexed colors to 0 dataView.setUint32(offset, picture.data.length); offset += 4; header.set(picture.data, offset); offset += picture.data.length; assert(offset === headerSize); const headerBitstream = new Bitstream(new Uint8Array(4)); headerBitstream.writeBits(1, 0); // Last metadata block -> false, will be continued by vorbis comment headerBitstream.writeBits(7, FlacBlockType.PICTURE); // Type -> Picture headerBitstream.writeBits(24, headerSize); this.writer.write(headerBitstream.bytes); this.writer.write(header); } writeVorbisCommentAndPictureBlock() { if (!this.format._options.appendOnly) { this.writer.seek(STREAMINFO_SIZE + FLAC_HEADER.byteLength); } if (metadataTagsAreEmpty(this.output._metadataTags)) { this.metadataWritten = true; return; } const pictures = this.output._metadataTags.images ?? []; for (const picture of pictures) { this.writePictureBlock(picture); } const vorbisComment = createVorbisComments( new Uint8Array(0), this.output._metadataTags, false, ); const headerBitstream = new Bitstream(new Uint8Array(4)); headerBitstream.writeBits(1, 1); // Last metadata block -> true headerBitstream.writeBits(7, FlacBlockType.VORBIS_COMMENT); // Type -> Vorbis comment headerBitstream.writeBits(24, vorbisComment.length); this.writer.write(headerBitstream.bytes); this.writer.write(vorbisComment); this.metadataWritten = true; } async getMimeType() { return 'audio/flac'; } async addEncodedVideoPacket() { throw new Error('FLAC does not support video.'); } async addEncodedAudioPacket( track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata, ): Promise { const release = await this.mutex.acquire(); try { this.validateTimestamp( track, packet.timestamp, packet.type === 'key', ); if (this.sampleRate === null) { // It's the first packet validateAudioChunkMetadata(meta); assert(meta); assert(meta.decoderConfig); assert(meta.decoderConfig.description); this.sampleRate = meta.decoderConfig.sampleRate; this.channels = meta.decoderConfig.numberOfChannels; const descriptionBitstream = new Bitstream( toUint8Array(meta.decoderConfig.description), ); // skip 'fLaC' + block size + frame size + sample rate + number of channels // See demuxer for the exact structure descriptionBitstream.skipBits(103 + 64); const bitsPerSample = descriptionBitstream.readBits(5) + 1; this.bitsPerSample = bitsPerSample; if (this.format._options.appendOnly) { // Write STREAMINFO immediately since we can't seek back later. this.writeHeader({ // https://www.rfc-editor.org/rfc/rfc9639.html#name-streaminfo // Per RFC 9639, min/max block sizes can be looser than // actual values, so we use the full valid range (16–65535). // "The actual max block size MAY be smaller than what's // listed, and the actual min (excluding last block) MAY be // larger. This is because the encoder has to write these // fields before receiving any input audio data and cannot // know beforehand what block sizes it will use." minimumBlockSize: 16, maximumBlockSize: 65535, // https://www.rfc-editor.org/rfc/rfc9639.html#name-streaminfo // "A value of 0 signifies that the value is not known." minimumFrameSize: 0, maximumFrameSize: 0, sampleRate: this.sampleRate, channels: this.channels, bitsPerSample: this.bitsPerSample, totalSamples: 0, }); } } if (!this.metadataWritten) { this.writeVorbisCommentAndPictureBlock(); } const slice = FileSlice.tempFromBytes(packet.data); slice.skip(2); const bytes = readBytes(slice, 2); const bitstream = new Bitstream(bytes); const blockSizeOrUncommon = getBlockSizeOrUncommon(bitstream.readBits(4)); if (blockSizeOrUncommon === null) { throw new Error('Invalid FLAC frame: Invalid block size.'); } readCodedNumber(slice); // num const blockSize = readBlockSize(slice, blockSizeOrUncommon); if (!this.format._options.appendOnly) { this.blockSizes.push(blockSize); this.frameSizes.push(packet.data.length); } const startPos = this.writer.getPos(); this.writer.write(packet.data); if (this.format._options.onFrame) { this.format._options.onFrame(packet.data, startPos); } await this.writer.flush(); } finally { release(); } } override addSubtitleCue(): Promise { throw new Error('FLAC does not support subtitles.'); } async finalize(): Promise { const release = await this.mutex.acquire(); if (!this.format._options.appendOnly) { let minimumBlockSize = Infinity; let maximumBlockSize = 0; let minimumFrameSize = Infinity; let maximumFrameSize = 0; let totalSamples = 0; for (let i = 0; i < this.blockSizes.length; i++) { minimumFrameSize = Math.min(minimumFrameSize, this.frameSizes[i]!); maximumFrameSize = Math.max(maximumFrameSize, this.frameSizes[i]!); maximumBlockSize = Math.max(maximumBlockSize, this.blockSizes[i]!); totalSamples += this.blockSizes[i]!; // Excluding the last frame from block size calculation // https://www.rfc-editor.org/rfc/rfc9639.html#name-streaminfo // "The minimum block size (in samples) used in the stream, excluding the last block." const isLastFrame = i === this.blockSizes.length - 1; if (isLastFrame) { continue; } minimumBlockSize = Math.min(minimumBlockSize, this.blockSizes[i]!); } assert(this.sampleRate !== null); assert(this.channels !== null); assert(this.bitsPerSample !== null); this.writer.seek(4); this.writeHeader({ minimumBlockSize, maximumBlockSize, minimumFrameSize, maximumFrameSize, sampleRate: this.sampleRate, channels: this.channels, bitsPerSample: this.bitsPerSample, totalSamples, }); } release(); } } ===== src/conversion.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AUDIO_CODECS, AudioCodec, NON_PCM_AUDIO_CODECS, VIDEO_CODECS, VideoCodec, } from './codec'; import { AudioEncodingConfig, getEncodableAudioCodecs, getFirstEncodableVideoCodec, Quality, QUALITY_HIGH, VideoEncodingConfig, } from './encode'; import { Input } from './input'; import { InputAudioTrack, InputTrack, InputVideoTrack } from './input-track'; import { Logging } from './logging'; import { AudioSampleSink, EncodedPacketSink, VideoSampleSink, } from './media-sink'; import { AudioSource, EncodedVideoPacketSource, EncodedAudioPacketSource, VideoSource, VideoSampleSource, AudioSampleSource, } from './media-source'; import { assert, assertNever, ceilToMultipleOfTwo, clamp, isIso639Dash2LanguageCode, MaybePromise, normalizeRotation, promiseWithResolvers, Rotation, } from './misc'; import { Output, OutputTrackGroup, TrackType } from './output'; import { Mp4OutputFormat } from './output-format'; import { AudioSample, clampCropRectangle, CropRectangle, validateCropRectangle, VideoSample, VideoSampleResource, } from './sample'; import { MetadataTags, validateMetadataTags } from './metadata'; import { NullTarget } from './target'; /** * The options for media file conversion. * @group Conversion * @public */ export type ConversionOptions = { /** The input file. */ input: Input; /** The output file. */ output: Output; /** * Defines which input tracks are used for conversion. Defaults to `'all'` unless the input is an HLS input, in * which case it defaults to `'primary'`. * * - `'all'`: All input tracks are eligible for conversion. * - `'primary'`: Only the primary video and audio track from the input are eligible for conversion. */ tracks?: 'all' | 'primary'; /** * Video-specific options. When passing an object, the same options are applied to all video tracks. When passing a * function, it will be invoked for each video track and is expected to return or resolve to the options * for that specific track. The function is passed an instance of {@link InputVideoTrack} as well as a number `n`, * which is the 1-based index of the track in the list of all video tracks. Using `n` is deprecated, prefer the * identical `track.number` instead. * * When passing an array of a function that returns an array, one output track per array element will be created, * allowing for "fan-out". Useful for creating multiple variants from a single track, for example with different * resolutions. */ video?: ConversionVideoOptions | ConversionVideoOptions[] | ((track: InputVideoTrack, n: number) => MaybePromise< ConversionVideoOptions | ConversionVideoOptions[] | undefined >); /** * Audio-specific options. When passing an object, the same options are applied to all audio tracks. When passing a * function, it will be invoked for each audio track and is expected to return or resolve to the options * for that specific track. The function is passed an instance of {@link InputAudioTrack} as well as a number `n`, * which is the 1-based index of the track in the list of all audio tracks. Using `n` is deprecated, prefer the * identical `track.number` instead. * * When passing an array of a function that returns an array, one output track per array element will be created, * allowing for "fan-out". Useful for creating multiple variants from a single track, for example with different * bitrates. */ audio?: ConversionAudioOptions | ConversionAudioOptions[] | ((track: InputAudioTrack, n: number) => MaybePromise< ConversionAudioOptions | ConversionAudioOptions[] | undefined >); /** Options to trim the input file. */ trim?: { /** * The time in the input file in seconds at which the output file should start. Must be less than `end`. * When omitted, defaults to the earliest start timestamp of the non-discarded tracks, or to 0, whichever * is higher. */ start?: number; /** * The time in the input file in seconds at which the output file should end. Must be greater than `start`. * Defaults to the duration of the input when omitted. */ end?: number; }; /** * An object or a callback that returns or resolves to an object containing the descriptive metadata tags that * should be written to the output file. If a function is passed, it will be passed the tags of the input file as * its first argument, allowing you to modify, augment or extend them. * * If no function is set, the input's metadata tags will be copied to the output. */ tags?: MetadataTags | ((inputTags: MetadataTags) => MaybePromise); /** * Whether to show potential console warnings about discarded tracks after calling `Conversion.init()`, defaults to * `true`. Set this to `false` if you're properly handling the `discardedTracks` and `isValid` fields already and * want to keep the console output clean. */ showWarnings?: boolean; }; /** * Video-specific options. * @group Conversion * @public */ export type ConversionVideoOptions = { /** If `true`, all video tracks will be discarded and will not be present in the output. */ discard?: boolean; /** * The desired width of the output video in pixels, defaulting to the video's natural display width. If height * is not set, it will be deduced automatically based on aspect ratio. */ width?: number; /** * The desired height of the output video in pixels, defaulting to the video's natural display height. If width * is not set, it will be deduced automatically based on aspect ratio. */ height?: number; /** * The fitting algorithm in case both width and height are set, or if the input video changes its size over time. * * - `'fill'` will stretch the image to fill the entire box, potentially altering aspect ratio. * - `'contain'` will contain the entire image within the box while preserving aspect ratio. This may lead to * letterboxing. * - `'cover'` will scale the image until the entire box is filled, while preserving aspect ratio. */ fit?: 'fill' | 'contain' | 'cover'; /** * The angle in degrees to rotate the input video by, clockwise. Rotation is applied before cropping and resizing. * This rotation is _in addition to_ the natural rotation of the input video as specified in input file's metadata. */ rotate?: Rotation; /** * Defaults to `true`. When enabled, Mediabunny will use the rotation metadata in the output file to perform video * rotation whenever possible. Set this field to `false` if you want to ensure the output file does not make use of * rotation metadata and that any rotation is baked into the video frames directly. */ allowRotationMetadata?: boolean; /** * Specifies the rectangular region of the input video to crop to. The crop region will automatically be clamped to * the dimensions of the input video track. Cropping is performed after rotation but before resizing. */ crop?: CropRectangle; /** * The desired frame rate of the output video, in hertz. If not specified, the original input frame rate will * be used (which may be variable). */ frameRate?: number; /** The desired output video codec. */ codec?: VideoCodec; /** The desired bitrate of the output video. */ bitrate?: number | Quality; /** * Whether to discard or keep the transparency information of the input video. The default is `'discard'`. Note that * for `'keep'` to produce a transparent video, you must use an output config that supports it, such as WebM with * VP9. */ alpha?: 'discard' | 'keep'; /** * The interval, in seconds, of how often frames are encoded as a key frame. The default is 5 seconds. Frequent key * frames improve seeking behavior but increase file size. When using multiple video tracks, you should give them * all the same key frame interval. * * Setting this fields forces a transcode. */ keyFrameInterval?: number; /** * A hint that configures the hardware acceleration method used when transcoding. This is best left on * `'no-preference'`, the default. */ hardwareAcceleration?: 'no-preference' | 'prefer-hardware' | 'prefer-software'; /** When `true`, video will always be re-encoded instead of directly copying over the encoded samples. */ forceTranscode?: boolean; /** * Allows for custom user-defined processing of video frames, e.g. for applying overlays, color transformations, or * timestamp modifications. Will be called for each input video sample after transformations and frame rate * corrections. * * Must return a {@link VideoSample}, a {@link VideoSampleResource} or a `CanvasImageSource`, an array of them, or * `null` for dropping the frame. When non-timestamped data is returned, the timestamp and duration from the source * sample will be used. Rotation metadata of the returned sample will be ignored. * * This function can also be used to manually resize frames. When doing so, you should signal the post-process * dimensions using the `processedWidth` and `processedHeight` fields, which enables the encoder to better know what * to expect. If these fields aren't set, Mediabunny will assume you won't perform any resizing. */ process?: (sample: VideoSample) => MaybePromise< CanvasImageSource | VideoSample | VideoSampleResource | (CanvasImageSource | VideoSample | VideoSampleResource)[] | null >; /** * An optional hint specifying the width of video samples returned by the `process` function, for better * encoder configuration. */ processedWidth?: number; /** * An optional hint specifying the height of video samples returned by the `process` function, for better * encoder configuration. */ processedHeight?: number; /** * Defines the group(s) the output track is a part of. For more, see {@link BaseTrackMetadata.group}. * * If left blank, tracks will internally be assigned to groups such that the output track pairability graph exactly * matches the input track pairability graph. */ group?: OutputTrackGroup | OutputTrackGroup[]; }; /** * Audio-specific options. * @group Conversion * @public */ export type ConversionAudioOptions = { /** If `true`, all audio tracks will be discarded and will not be present in the output. */ discard?: boolean; /** The desired channel count of the output audio. */ numberOfChannels?: number; /** The desired sample rate of the output audio, in hertz. */ sampleRate?: number; /** * The desired sample format (and therefore bit depth) of the audio samples before they are passed to the encoder. * Can be used to control bit depth with certain output codecs such as FLAC. * * Setting this field forces audio transcoding. */ sampleFormat?: 'u8' | 's16' | 's32' | 'f32'; /** The desired output audio codec. */ codec?: AudioCodec; /** The desired bitrate of the output audio. */ bitrate?: number | Quality; /** When `true`, audio will always be re-encoded instead of directly copying over the encoded samples. */ forceTranscode?: boolean; /** * Allows for custom user-defined processing of audio samples, e.g. for applying audio effects, transformations, or * timestamp modifications. Will be called for each input audio sample after remixing and resampling. * * Must return an {@link AudioSample}, an array of them, or `null` for dropping the sample. * * This function can also be used to manually perform remixing or resampling. When doing so, you should signal the * post-process parameters using the `processedNumberOfChannels` and `processedSampleRate` fields, which enables the * encoder to better know what to expect. If these fields aren't set, Mediabunny will assume you won't perform * remixing or resampling. */ process?: (sample: AudioSample) => MaybePromise< AudioSample | AudioSample[] | null >; /** * An optional hint specifying the channel count of audio samples returned by the `process` function, for better * encoder configuration. */ processedNumberOfChannels?: number; /** * An optional hint specifying the sample rate of audio samples returned by the `process` function, for better * encoder configuration. */ processedSampleRate?: number; /** * Defines the group(s) the output track is a part of. For more, see {@link BaseTrackMetadata.group}. * * If left blank, tracks will internally be assigned to groups such that the output track pairability graph exactly * matches the input track pairability graph. */ group?: OutputTrackGroup | OutputTrackGroup[]; }; const validateVideoOptions = (videoOptions: ConversionVideoOptions) => { if (!videoOptions || typeof videoOptions !== 'object') { throw new TypeError('options.video, when provided, must be an object.'); } if (videoOptions?.discard !== undefined && typeof videoOptions.discard !== 'boolean') { throw new TypeError('options.video.discard, when provided, must be a boolean.'); } if (videoOptions?.forceTranscode !== undefined && typeof videoOptions.forceTranscode !== 'boolean') { throw new TypeError('options.video.forceTranscode, when provided, must be a boolean.'); } if (videoOptions?.codec !== undefined && !VIDEO_CODECS.includes(videoOptions.codec)) { throw new TypeError( `options.video.codec, when provided, must be one of: ${VIDEO_CODECS.join(', ')}.`, ); } if ( videoOptions?.bitrate !== undefined && !(videoOptions.bitrate instanceof Quality) && (!Number.isInteger(videoOptions.bitrate) || videoOptions.bitrate <= 0) ) { throw new TypeError('options.video.bitrate, when provided, must be a positive integer or a quality.'); } if ( videoOptions?.width !== undefined && (!Number.isInteger(videoOptions.width) || videoOptions.width <= 0) ) { throw new TypeError('options.video.width, when provided, must be a positive integer.'); } if ( videoOptions?.height !== undefined && (!Number.isInteger(videoOptions.height) || videoOptions.height <= 0) ) { throw new TypeError('options.video.height, when provided, must be a positive integer.'); } if (videoOptions?.fit !== undefined && !['fill', 'contain', 'cover'].includes(videoOptions.fit)) { throw new TypeError('options.video.fit, when provided, must be one of \'fill\', \'contain\', or \'cover\'.'); } if ( videoOptions?.width !== undefined && videoOptions.height !== undefined && videoOptions.fit === undefined ) { throw new TypeError( 'When both options.video.width and options.video.height are provided, options.video.fit must also be' + ' provided.', ); } if (videoOptions?.rotate !== undefined && ![0, 90, 180, 270].includes(videoOptions.rotate)) { throw new TypeError('options.video.rotate, when provided, must be 0, 90, 180 or 270.'); } if (videoOptions?.allowRotationMetadata !== undefined && typeof videoOptions.allowRotationMetadata !== 'boolean') { throw new TypeError('options.video.allowRotationMetadata, when provided, must be a boolean.'); } if (videoOptions?.crop !== undefined) { validateCropRectangle(videoOptions.crop, 'options.video.'); } if ( videoOptions?.frameRate !== undefined && (!Number.isFinite(videoOptions.frameRate) || videoOptions.frameRate <= 0) ) { throw new TypeError('options.video.frameRate, when provided, must be a finite positive number.'); } if (videoOptions?.alpha !== undefined && !['discard', 'keep'].includes(videoOptions.alpha)) { throw new TypeError('options.video.alpha, when provided, must be either \'discard\' or \'keep\'.'); } if ( videoOptions?.keyFrameInterval !== undefined && (!Number.isFinite(videoOptions.keyFrameInterval) || videoOptions.keyFrameInterval < 0) ) { throw new TypeError('options.video.keyFrameInterval, when provided, must be a non-negative number.'); } if (videoOptions?.process !== undefined && typeof videoOptions.process !== 'function') { throw new TypeError('options.video.process, when provided, must be a function.'); } if ( videoOptions?.processedWidth !== undefined && (!Number.isInteger(videoOptions.processedWidth) || videoOptions.processedWidth <= 0) ) { throw new TypeError('options.video.processedWidth, when provided, must be a positive integer.'); } if ( videoOptions?.processedHeight !== undefined && (!Number.isInteger(videoOptions.processedHeight) || videoOptions.processedHeight <= 0) ) { throw new TypeError('options.video.processedHeight, when provided, must be a positive integer.'); } if ( videoOptions?.hardwareAcceleration !== undefined && !['no-preference', 'prefer-hardware', 'prefer-software'].includes(videoOptions.hardwareAcceleration) ) { throw new TypeError( 'options.video.hardwareAcceleration, when provided, must be \'no-preference\', \'prefer-hardware\' or' + ' \'prefer-software\'.', ); } if ( videoOptions?.group !== undefined && !( videoOptions.group instanceof OutputTrackGroup || (Array.isArray(videoOptions.group) && videoOptions.group.every(x => x instanceof OutputTrackGroup)) ) ) { throw new TypeError( 'options.video.group, when provided, must be an OutputTrackGroup or an array of OutputTrackGroups.', ); } }; const validateAudioOptions = (audioOptions: ConversionAudioOptions) => { if (!audioOptions || typeof audioOptions !== 'object') { throw new TypeError('options.audio, when provided, must be an object.'); } if (audioOptions?.discard !== undefined && typeof audioOptions.discard !== 'boolean') { throw new TypeError('options.audio.discard, when provided, must be a boolean.'); } if (audioOptions?.forceTranscode !== undefined && typeof audioOptions.forceTranscode !== 'boolean') { throw new TypeError('options.audio.forceTranscode, when provided, must be a boolean.'); } if (audioOptions?.codec !== undefined && !AUDIO_CODECS.includes(audioOptions.codec)) { throw new TypeError( `options.audio.codec, when provided, must be one of: ${AUDIO_CODECS.join(', ')}.`, ); } if ( audioOptions?.bitrate !== undefined && !(audioOptions.bitrate instanceof Quality) && (!Number.isInteger(audioOptions.bitrate) || audioOptions.bitrate <= 0) ) { throw new TypeError('options.audio.bitrate, when provided, must be a positive integer or a quality.'); } if ( audioOptions?.numberOfChannels !== undefined && (!Number.isInteger(audioOptions.numberOfChannels) || audioOptions.numberOfChannels <= 0) ) { throw new TypeError('options.audio.numberOfChannels, when provided, must be a positive integer.'); } if ( audioOptions?.sampleRate !== undefined && (!Number.isInteger(audioOptions.sampleRate) || audioOptions.sampleRate <= 0) ) { throw new TypeError('options.audio.sampleRate, when provided, must be a positive integer.'); } if ( audioOptions?.sampleFormat !== undefined && !['u8', 's16', 's32', 'f32'].includes(audioOptions.sampleFormat) ) { throw new TypeError('options.audio.sampleFormat, when provided, must be one of: u8, s16, s32, f32.'); } if (audioOptions?.process !== undefined && typeof audioOptions.process !== 'function') { throw new TypeError('options.audio.process, when provided, must be a function.'); } if ( audioOptions?.processedNumberOfChannels !== undefined && (!Number.isInteger(audioOptions.processedNumberOfChannels) || audioOptions.processedNumberOfChannels <= 0) ) { throw new TypeError('options.audio.processedNumberOfChannels, when provided, must be a positive integer.'); } if ( audioOptions?.processedSampleRate !== undefined && (!Number.isInteger(audioOptions.processedSampleRate) || audioOptions.processedSampleRate <= 0) ) { throw new TypeError('options.audio.processedSampleRate, when provided, must be a positive integer.'); } if ( audioOptions?.group !== undefined && !( audioOptions.group instanceof OutputTrackGroup || (Array.isArray(audioOptions.group) && audioOptions.group.every(x => x instanceof OutputTrackGroup)) ) ) { throw new TypeError( 'options.audio.group, when provided, must be an OutputTrackGroup or an array of OutputTrackGroups.', ); } }; const FALLBACK_NUMBER_OF_CHANNELS = 2; const FALLBACK_SAMPLE_RATE = 48000; /** * An input track that was discarded (excluded) from a {@link Conversion} alongside the discard reason. * @group Conversion * @public */ export type DiscardedTrack = { /** The track that was discarded. */ track: InputTrack; /** * The reason for discarding the track. * * - `'discarded_by_user'`: You discarded this track by setting `discard: true`. * - `'max_track_count_reached'`: The output had no more room for another track. * - `'max_track_count_of_type_reached'`: The output had no more room for another track of this type, or the output * doesn't support this track type at all. * - `'unknown_source_codec'`: We don't know the codec of the input track and therefore don't know what to do * with it. * - `'undecodable_source_codec'`: The input track's codec is known, but we are unable to decode it. * - `'no_encodable_target_codec'`: We can't find a codec that we are able to encode and that can be contained * within the output format. This reason can be hit if the environment doesn't support the necessary encoders, or if * you requested a codec that cannot be contained within the output format. */ reason: | 'discarded_by_user' | 'max_track_count_reached' | 'max_track_count_of_type_reached' | 'unknown_source_codec' | 'undecodable_source_codec' | 'no_encodable_target_codec'; /** The options that were provided for this track, or `{}` if none were provided. */ trackOptions: ConversionVideoOptions | ConversionAudioOptions; }; /** * Represents a media file conversion process, used to convert one media file into another. In addition to conversion, * this class can be used to resize and rotate video, resample audio, drop tracks, or trim to a specific time range. * @group Conversion * @public */ export class Conversion { /** The input file. */ readonly input: Input; /** The output file. */ readonly output: Output; /** @internal */ _options: ConversionOptions; /** @internal */ _startTimestamp!: number; /** @internal */ _endTimestamp!: number; /** @internal */ _addedCounts: Record = { video: 0, audio: 0, subtitle: 0, }; /** @internal */ _totalTrackCount = 0; /** @internal */ _nextOutputTrackId = 0; /** @internal */ _outputTrackIds: number[] = []; /** @internal */ _outputOwnTrackGroups: (OutputTrackGroup | null)[] = []; /** @internal */ _trackPromises: Promise[] = []; /** @internal */ _started: Promise; /** @internal */ _start: () => void; /** @internal */ _executed = false; /** @internal */ _synchronizer = new TrackSynchronizer(); /** @internal */ _totalDuration: number | null = null; /** @internal */ _maxTimestamps = new Map(); // Track ID -> timestamp /** @internal */ _canceled = false; /** * A callback that is fired whenever the conversion progresses. Gets passed as first argument a number between * 0 and 1, indicating the completion of the conversion. Note that a progress of 1 doesn't necessarily mean the * conversion is complete; the conversion is complete once `execute()` resolves. * * As second argument, this callback receives the input time in seconds that has been processed. * * In order for progress to be computed, this property must be set before `execute` is called. */ onProgress?: (progress: number, processedTime: number) => unknown = undefined; /** @internal */ _computeProgress = false; /** @internal */ _lastProgress = 0; /** * Whether this conversion, as it has been configured, is valid and can be executed. If this field is `false`, check * the `discardedTracks` field for reasons. * * Note: a conversion having discarded tracks does not automatically mean it is invalid; if the remaining, utilized * tracks make for a valid output file, the conversion is still allowed. */ isValid = false; /** * The list of tracks that are included in the output file. When fan-out is used, the same track appears in this * array multiple times. */ readonly utilizedTracks: InputTrack[] = []; /** The list of tracks from the input file that have been discarded, alongside the discard reason. */ readonly discardedTracks: DiscardedTrack[] = []; /** Initializes a new conversion process without starting the conversion. */ static async init(options: ConversionOptions) { const conversion = new Conversion(options); await conversion._init(); return conversion; } /** Creates a new Conversion instance (duh). */ private constructor(options: ConversionOptions) { if (!options || typeof options !== 'object') { throw new TypeError('options must be an object.'); } if (!(options.input instanceof Input)) { throw new TypeError('options.input must be an Input.'); } if (!(options.output instanceof Output)) { throw new TypeError('options.output must be an Output.'); } if ( options.tracks !== undefined && options.tracks !== 'all' && options.tracks !== 'primary' ) { throw new TypeError( 'options.tracks, when provided, must be either \'all\' or \'primary\'.', ); } if ( options.output._tracks.length > 0 || Object.keys(options.output._metadataTags).length > 0 || options.output.state !== 'pending' ) { throw new TypeError('options.output must be fresh: no tracks or metadata tags added and not started.'); } if (options.video !== undefined && typeof options.video !== 'function') { if (Array.isArray(options.video)) { for (const obj of options.video) { validateVideoOptions(obj); } } else { validateVideoOptions(options.video); } } else { // We'll validate the return value later } if (options.audio !== undefined && typeof options.audio !== 'function') { if (Array.isArray(options.audio)) { for (const obj of options.audio) { validateAudioOptions(obj); } } else { validateAudioOptions(options.audio); } } else { // We'll validate the return value later } if (options.trim !== undefined && (!options.trim || typeof options.trim !== 'object')) { throw new TypeError('options.trim, when provided, must be an object.'); } if (options.trim?.start !== undefined && (!Number.isFinite(options.trim.start))) { throw new TypeError('options.trim.start, when provided, must be a finite number.'); } if (options.trim?.end !== undefined && (!Number.isFinite(options.trim.end))) { throw new TypeError('options.trim.end, when provided, must be a finite number.'); } if ( options.trim?.start !== undefined && options.trim.end !== undefined && options.trim.start >= options.trim.end) { throw new TypeError('options.trim.start must be less than options.trim.end.'); } if ( options.tags !== undefined && (typeof options.tags !== 'object' || !options.tags) && typeof options.tags !== 'function' ) { throw new TypeError('options.tags, when provided, must be an object or a function.'); } if (typeof options.tags === 'object') { validateMetadataTags(options.tags); } if (options.showWarnings !== undefined && typeof options.showWarnings !== 'boolean') { throw new TypeError('options.showWarnings, when provided, must be a boolean.'); } this._options = options; this.input = options.input; this.output = options.output; const { promise: started, resolve: start } = promiseWithResolvers(); this._started = started; this._start = start; } /** @internal */ async _init() { const inputFormat = await this.input.getFormat(); let tracks: InputTrack[]; let trackMode = this._options.tracks; if (trackMode === undefined) { // HACK to keep bundle size low, temp for now const defaultTrackMode = inputFormat.name.includes('(HLS)') ? 'primary' : 'all'; trackMode = defaultTrackMode; } if (trackMode === 'all') { tracks = await this.input.getTracks(); } else if (trackMode === 'primary') { const primaryVideoTrack = await this.input.getPrimaryVideoTrack(); const primaryAudioTrack = await this.input.getPrimaryAudioTrack(); tracks = [primaryVideoTrack, primaryAudioTrack].filter(x => x !== null); } else { assertNever(trackMode); assert(false); } const outputTrackCounts = this.output.format.getSupportedTrackCounts(); // Input track counters let nVideo = 1; let nAudio = 1; // All tracks that aren't discarded by the user const filteredTracks: InputTrack[] = []; const filteredTrackOptions: (ConversionVideoOptions | ConversionAudioOptions)[][] = []; for (const track of tracks) { let trackOptions: (ConversionVideoOptions | ConversionAudioOptions)[]; if (track.isVideoTrack()) { if (this._options.video) { if (typeof this._options.video === 'function') { const returnedTrackOptions = await this._options.video(track, nVideo) ?? {}; if (Array.isArray(returnedTrackOptions)) { for (const obj of returnedTrackOptions) { validateVideoOptions(obj); } } else { validateVideoOptions(returnedTrackOptions); } trackOptions = Array.isArray(returnedTrackOptions) ? returnedTrackOptions : [returnedTrackOptions]; nVideo++; } else { // Already validated trackOptions = Array.isArray(this._options.video) ? this._options.video : [this._options.video]; } } else { trackOptions = [{}]; } } else if (track.isAudioTrack()) { if (this._options.audio) { if (typeof this._options.audio === 'function') { const returnedTrackOptions = await this._options.audio(track, nAudio) ?? {}; if (Array.isArray(returnedTrackOptions)) { for (const obj of returnedTrackOptions) { validateAudioOptions(obj); } } else { validateAudioOptions(returnedTrackOptions); } trackOptions = Array.isArray(returnedTrackOptions) ? returnedTrackOptions : [returnedTrackOptions]; nAudio++; } else { // Already validated trackOptions = Array.isArray(this._options.audio) ? this._options.audio : [this._options.audio]; } } else { trackOptions = [{}]; } } else { assert(false); } const discardOptions = trackOptions.filter(x => x.discard); for (const discardOption of discardOptions) { this.discardedTracks.push({ track, reason: 'discarded_by_user', trackOptions: discardOption, }); } if (trackOptions.length === discardOptions.length) { if (trackOptions.length === 0) { this.discardedTracks.push({ track, reason: 'discarded_by_user', trackOptions: {}, }); } continue; } const nonDiscardOptions = trackOptions.filter(x => !x.discard); filteredTracks.push(track); filteredTrackOptions.push(nonDiscardOptions); } if (this._options.trim?.start !== undefined) { this._startTimestamp = this._options.trim.start; } else { // Compute the start timestamp from the set of filtered tracks. Techncially these can still be narrowed // down later due to discarded tracks, but we need to fix the start timestamp now due to track processing // depending on it. this._startTimestamp = Math.max( await this.input.getFirstTimestamp(filteredTracks), // Samples can also have negative timestamps, but the meaning typically is "don't present me", so let's // cut those out by default. 0, ); } this._endTimestamp = Math.max(this._options.trim?.end ?? Infinity, this._startTimestamp); // Run these sequentially so that output tracks have a deterministic order for (let i = 0; i < filteredTracks.length; i++) { const track = filteredTracks[i]!; const options = filteredTrackOptions[i]!; for (const option of options) { if (this._totalTrackCount === outputTrackCounts.total.max) { this.discardedTracks.push({ track, reason: 'max_track_count_reached', trackOptions: option, }); continue; } if (this._addedCounts[track.type] === outputTrackCounts[track.type].max) { this.discardedTracks.push({ track, reason: 'max_track_count_of_type_reached', trackOptions: option, }); continue; } const outputTrackId = this._nextOutputTrackId++; if (track.isVideoTrack()) { await this._processVideoTrack(track, option as ConversionVideoOptions, outputTrackId); } else if (track.isAudioTrack()) { await this._processAudioTrack(track, option as ConversionAudioOptions, outputTrackId); } else { assert(false); } } } // When no track groups are set by the user, then the output track pairability should be *identical* to the // input's. We do the naive algorithm to achieve this: assign each track to its own group, and pair groups with // each other based on input track pairability. for (let i = 0; i < this.utilizedTracks.length - 1; i++) { for (let j = i + 1; j < this.utilizedTracks.length; j++) { const trackA = this.utilizedTracks[i]!; const trackB = this.utilizedTracks[j]!; const ownGroupA = this._outputOwnTrackGroups[i]; const ownGroupB = this._outputOwnTrackGroups[j]; assert(ownGroupA !== undefined); assert(ownGroupB !== undefined); if (ownGroupA && ownGroupB && trackA.canBePairedWith(trackB)) { ownGroupA.pairWith(ownGroupB); } } } // Now, let's deal with metadata tags const inputTags = await this.input.getMetadataTags(); let outputTags: MetadataTags; if (this._options.tags) { const result = typeof this._options.tags === 'function' ? await this._options.tags(inputTags) : this._options.tags; validateMetadataTags(result); outputTags = result; } else { outputTags = inputTags; } // Somewhat dirty but pragmatic const inputAndOutputFormatMatch = inputFormat.mimeType === this.output.format.mimeType; const rawTagsAreUnchanged = inputTags.raw === outputTags.raw; if (inputTags.raw && rawTagsAreUnchanged && !inputAndOutputFormatMatch) { // If the input and output formats aren't the same, copying over raw metadata tags makes no sense and only // results in junk tags, so let's cut them out. delete outputTags.raw; } this.output.setMetadataTags(outputTags); // Let's check if the conversion can actually be executed this.isValid = this._totalTrackCount >= outputTrackCounts.total.min && this._addedCounts.video >= outputTrackCounts.video.min && this._addedCounts.audio >= outputTrackCounts.audio.min && this._addedCounts.subtitle >= outputTrackCounts.subtitle.min; if (this._options.showWarnings ?? true) { const warnElements: unknown[] = []; const unintentionallyDiscardedTracks = this.discardedTracks.filter(x => x.reason !== 'discarded_by_user'); if (unintentionallyDiscardedTracks.length > 0) { // Let's give the user a notice/warning about discarded tracks so they aren't confused warnElements.push( 'Some tracks had to be discarded from the conversion:', unintentionallyDiscardedTracks, ); } if (!this.isValid) { if (warnElements.length > 0) { warnElements.push('\n\n'); } warnElements.push(this._getInvalidityExplanation().join('')); } if (warnElements.length > 0) { Logging._warn(...warnElements); } } } /** @internal */ _getInvalidityExplanation() { const elements: string[] = []; if (this.discardedTracks.length === 0) { elements.push( 'Due to missing tracks, this conversion cannot be executed.', ); } else { const encodabilityIsTheProblem = this.discardedTracks.every(x => x.reason === 'discarded_by_user' || x.reason === 'no_encodable_target_codec', ) && this.discardedTracks.some(x => x.reason === 'no_encodable_target_codec'); elements.push( 'Due to discarded tracks, this conversion cannot be executed.', ); if (encodabilityIsTheProblem) { const codecs = this.discardedTracks.flatMap((x) => { if (x.reason === 'discarded_by_user') return []; if (x.track.type === 'video') { return this.output.format.getSupportedVideoCodecs(); } else if (x.track.type === 'audio') { return this.output.format.getSupportedAudioCodecs(); } else { return this.output.format.getSupportedSubtitleCodecs(); } }); const uniqueCodecs = [...new Set(codecs)]; if (uniqueCodecs.length === 1) { elements.push( `\nTracks were discarded because your environment is not able to encode '${uniqueCodecs[0]}'.`, ); } else { elements.push( '\nTracks were discarded because your environment is not able to encode any of the following' + ` codecs: ${uniqueCodecs.map(x => `'${x}'`).join(', ')}.`, ); } if (uniqueCodecs.includes('mp3')) { elements.push( `\nThe @mediabunny/mp3-encoder extension package provides support for encoding MP3.`, ); } if (uniqueCodecs.includes('aac')) { elements.push( '\nThe @mediabunny/aac-encoder extension package provides support for encoding AAC.', ); } if (uniqueCodecs.includes('ac3') || uniqueCodecs.includes('eac3')) { elements.push( '\nThe @mediabunny/ac3 extension package provides support' + ' for encoding and decoding AC-3/E-AC-3.', ); } if (uniqueCodecs.includes('flac')) { elements.push( '\nThe @mediabunny/flac-encoder extension package provides support for encoding FLAC.', ); } } else { elements.push('\nCheck the discardedTracks field for more info.'); } } return elements; } /** * Executes the conversion process. Resolves once conversion is complete. * * Will throw if `isValid` is `false`. */ async execute() { if (!this.isValid) { throw new Error( 'Cannot execute this conversion because its output configuration is invalid. Make sure to always check' + ' the isValid field before executing a conversion.\n' + this._getInvalidityExplanation().join(''), ); } if (this._executed) { throw new Error('Conversion cannot be executed twice.'); } this._executed = true; for (const id of this._outputTrackIds) { this._synchronizer.declareTrack(id); } if (this.onProgress) { // Compute duration using only the utilized tracks const uniqueUtilizedTracks = new Set(this.utilizedTracks); const durationPromises = [...uniqueUtilizedTracks].map(async (track) => { if (await track.isLive()) { return Infinity; // Upper bound (assuming no universe heat death) } return (await track.getDurationFromMetadata()) ?? (await track.computeDuration()); }); const duration = Math.max(0, ...await Promise.all(durationPromises)); this._computeProgress = true; this._totalDuration = Math.min( duration - this._startTimestamp, this._endTimestamp - this._startTimestamp, ); for (const id of this._outputTrackIds) { this._maxTimestamps.set(id, 0); } this.onProgress?.(0, 0); } await this.output.start(); this._start(); try { await Promise.all(this._trackPromises); } catch (error) { if (!this._canceled) { // Make sure to cancel to stop other encoding processes and clean up resources void this.cancel(); } throw error; } if (this._canceled) { throw new ConversionCanceledError(); } await this.output.finalize(); if (this._computeProgress) { const minTimestamp = Math.min(...this._maxTimestamps.values()); this.onProgress?.(1, minTimestamp); } } /** * Cancels the conversion process, causing any ongoing `execute` call to throw a `ConversionCanceledError`. * Does nothing if the conversion is already complete. */ async cancel() { if (this.output.state === 'finalizing' || this.output.state === 'finalized') { return; } if (this._canceled) { Logging._warn('Conversion already canceled.'); return; } this._canceled = true; await this.output.cancel(); } /** @internal */ async _processVideoTrack(track: InputVideoTrack, trackOptions: ConversionVideoOptions, outputTrackId: number) { const sourceCodec = await track.getCodec(); if (!sourceCodec) { this.discardedTracks.push({ track, reason: 'unknown_source_codec', trackOptions, }); return; } let videoSource: VideoSource; const innateRotation = await track.getRotation(); const totalRotation = normalizeRotation(innateRotation + (trackOptions.rotate ?? 0)); let outputTrackRotation = totalRotation; const canUseRotationMetadata = this.output.format.supportsVideoRotationMetadata && (trackOptions.allowRotationMetadata ?? true); const squarePixelWidth = await track.getSquarePixelWidth(); const squarePixelHeight = await track.getSquarePixelHeight(); const [rotatedWidth, rotatedHeight] = totalRotation % 180 === 0 ? [squarePixelWidth, squarePixelHeight] : [squarePixelHeight, squarePixelWidth]; let crop = trackOptions.crop; if (crop) { crop = clampCropRectangle(crop, rotatedWidth, rotatedHeight); } const [originalWidth, originalHeight] = crop ? [crop.width, crop.height] : [rotatedWidth, rotatedHeight]; let width = originalWidth; let height = originalHeight; const aspectRatio = width / height; // A lot of video encoders require that the dimensions be multiples of 2 if (trackOptions.width !== undefined && trackOptions.height === undefined) { width = ceilToMultipleOfTwo(trackOptions.width); height = ceilToMultipleOfTwo(Math.round(width / aspectRatio)); } else if (trackOptions.width === undefined && trackOptions.height !== undefined) { height = ceilToMultipleOfTwo(trackOptions.height); width = ceilToMultipleOfTwo(Math.round(height * aspectRatio)); } else if (trackOptions.width !== undefined && trackOptions.height !== undefined) { width = ceilToMultipleOfTwo(trackOptions.width); height = ceilToMultipleOfTwo(trackOptions.height); } const firstTimestamp = await track.getFirstTimestamp(); let videoCodecs = this.output.format.getSupportedVideoCodecs(); const needsTranscode = !!trackOptions.forceTranscode || firstTimestamp < this._startTimestamp || !!trackOptions.frameRate || trackOptions.keyFrameInterval !== undefined || trackOptions.process !== undefined || trackOptions.bitrate !== undefined || !videoCodecs.includes(sourceCodec) || (trackOptions.codec && trackOptions.codec !== sourceCodec) || width !== originalWidth || height !== originalHeight // TODO This is suboptimal: Forcing a rerender when both rotation and process are set is not // performance-optimal, but right now there's no other way because we can't change the track rotation // metadata after the output has already started. Should be possible with API changes in v2, though! || (totalRotation !== 0 && !canUseRotationMetadata) || !!crop; const alpha = trackOptions.alpha ?? 'discard'; if (!needsTranscode) { // Fast path, we can simply copy over the encoded packets const source = new EncodedVideoPacketSource(sourceCodec); videoSource = source; this._trackPromises.push((async () => { await this._started; const sink = new EncodedPacketSink(track); const decoderConfig = await track.getDecoderConfig(); const meta: EncodedVideoChunkMetadata = { decoderConfig: decoderConfig ?? undefined }; for await (const packet of sink.packets(undefined, undefined, { verifyKeyPackets: true })) { if (this._canceled) { return; } if (packet.timestamp >= this._endTimestamp) { break; } const modifiedPacket = packet.clone({ timestamp: packet.timestamp - this._startTimestamp, sideData: alpha === 'discard' ? {} // Remove alpha side data : packet.sideData, }); assert(modifiedPacket.timestamp >= 0); this._reportProgress(outputTrackId, modifiedPacket.timestamp + modifiedPacket.duration); await source.add(modifiedPacket, meta); if (this._synchronizer.shouldWait(outputTrackId, modifiedPacket.timestamp)) { await this._synchronizer.wait(modifiedPacket.timestamp); } } source.close(); this._synchronizer.closeTrack(outputTrackId); })()); } else { // We need to decode & reencode the video const canDecode = await track.canDecode(); if (!canDecode) { this.discardedTracks.push({ track, reason: 'undecodable_source_codec', trackOptions, }); return; } if (trackOptions.codec) { videoCodecs = videoCodecs.filter(codec => codec === trackOptions.codec); } const bitrate = trackOptions.bitrate ?? QUALITY_HIGH; const encodableCodec = await getFirstEncodableVideoCodec(videoCodecs, { width: trackOptions.process && trackOptions.processedWidth ? trackOptions.processedWidth : width, height: trackOptions.process && trackOptions.processedHeight ? trackOptions.processedHeight : height, bitrate, }); if (!encodableCodec) { this.discardedTracks.push({ track, reason: 'no_encodable_target_codec', trackOptions, }); return; } const encodingConfig: VideoEncodingConfig = { codec: encodableCodec, bitrate, keyFrameInterval: trackOptions.keyFrameInterval, sizeChangeBehavior: trackOptions.fit ?? 'passThrough', alpha, hardwareAcceleration: trackOptions.hardwareAcceleration, transform: {}, }; assert(encodingConfig.transform); let needsRerender = width !== originalWidth || height !== originalHeight || (totalRotation !== 0 && (!canUseRotationMetadata || trackOptions.process !== undefined)) || !!crop // Don't expect encoders to reliably handle non-square pixels: || squarePixelWidth !== await track.getCodedWidth() || squarePixelHeight !== await track.getCodedHeight(); if (!needsRerender) { // If we're directly passing decoded samples back to the encoder, sometimes the encoder may error due // to lack of support of certain video frame formats, like when HDR is at play. To check for this, we // first try to pass a single frame to the encoder to see how it behaves. If it throws, we then fall // back to the rerender path. // // Creating a new temporary Output is sort of hacky, but due to a lack of an isolated encoder API right // now, this is the simplest way. Will refactor in the future! TODO const tempOutput = new Output({ format: new Mp4OutputFormat(), // Supports all video codecs target: new NullTarget(), }); const tempSource = new VideoSampleSource(encodingConfig); tempOutput.addVideoTrack(tempSource); await tempOutput.start(); const sink = new VideoSampleSink(track); const firstSample = await sink.getSample(firstTimestamp); // Let's just use the first sample if (firstSample) { try { await tempSource.add(firstSample); firstSample.close(); await tempOutput.finalize(); } catch (error) { Logging._info('Error when probing encoder support. Falling back to rerender path.', error); needsRerender = true; void tempOutput.cancel(); } } else { await tempOutput.cancel(); } } if (trackOptions.frameRate) { encodingConfig.transform.frameRate = trackOptions.frameRate; } if (trackOptions.process) { encodingConfig.transform.process = trackOptions.process; } if (needsRerender) { outputTrackRotation = 0; // Since the rotation is baked into the output encodingConfig.transform.width = width; encodingConfig.transform.height = height; encodingConfig.transform.fit = trackOptions.fit ?? 'fill'; encodingConfig.transform.rotate = normalizeRotation(totalRotation - innateRotation); encodingConfig.transform.crop = crop; encodingConfig.transform.alpha = alpha; } // We need to do this because `process` can emit new timestamps let lastSampleTimestamp: number | null = null; encodingConfig.onEncodedSample = (sample) => { lastSampleTimestamp = sample.timestamp; }; const source = new VideoSampleSource(encodingConfig); videoSource = source; this._trackPromises.push((async () => { await this._started; const sink = new VideoSampleSink(track); for await (const sample of sink.samples(this._startTimestamp, this._endTimestamp)) { if (this._canceled) { sample.close(); return; } const adjustedSampleTimestamp = Math.max(sample.timestamp - this._startTimestamp, 0); sample.setTimestamp(adjustedSampleTimestamp); this._reportProgress(outputTrackId, sample.timestamp + sample.duration); await source.add(sample); if (lastSampleTimestamp !== null) { if (this._synchronizer.shouldWait(outputTrackId, lastSampleTimestamp)) { await this._synchronizer.wait(lastSampleTimestamp); } } sample.close(); } source.close(); this._synchronizer.closeTrack(outputTrackId); })()); } let ownGroup: OutputTrackGroup | null = null; if (!trackOptions.group) { ownGroup = new OutputTrackGroup(); } const videoTrackLanguageCode = await track.getLanguageCode(); this.output.addVideoTrack(videoSource, { frameRate: trackOptions.frameRate, // TODO: This condition can be removed when all demuxers properly homogenize to BCP47 in v2 languageCode: isIso639Dash2LanguageCode(videoTrackLanguageCode) ? videoTrackLanguageCode : undefined, name: await track.getName() ?? undefined, disposition: await track.getDisposition(), rotation: outputTrackRotation, group: ownGroup ?? trackOptions.group, }); this._addedCounts.video++; this._totalTrackCount++; this.utilizedTracks.push(track); this._outputTrackIds.push(outputTrackId); this._outputOwnTrackGroups.push(ownGroup); } /** @internal */ async _processAudioTrack(track: InputAudioTrack, trackOptions: ConversionAudioOptions, outputTrackId: number) { const sourceCodec = await track.getCodec(); if (!sourceCodec) { this.discardedTracks.push({ track, reason: 'unknown_source_codec', trackOptions, }); return; } let audioSource: AudioSource; const originalNumberOfChannels = await track.getNumberOfChannels(); const originalSampleRate = await track.getSampleRate(); const firstTimestamp = await track.getFirstTimestamp(); let numberOfChannels = trackOptions.numberOfChannels ?? originalNumberOfChannels; let sampleRate = trackOptions.sampleRate ?? originalSampleRate; const needsTrimming = firstTimestamp < this._startTimestamp; const needsPadding = firstTimestamp > this._startTimestamp && !this.output.format.supportsTimestampedMediaData; let audioCodecs = this.output.format.getSupportedAudioCodecs(); if ( !trackOptions.forceTranscode && !trackOptions.bitrate && numberOfChannels === originalNumberOfChannels && sampleRate === originalSampleRate && !needsTrimming && !needsPadding && audioCodecs.includes(sourceCodec) && (!trackOptions.codec || trackOptions.codec === sourceCodec) && !trackOptions.process && trackOptions.sampleFormat === undefined ) { // Fast path, we can simply copy over the encoded packets const source = new EncodedAudioPacketSource(sourceCodec); audioSource = source; this._trackPromises.push((async () => { await this._started; const sink = new EncodedPacketSink(track); const decoderConfig = await track.getDecoderConfig(); const meta: EncodedAudioChunkMetadata = { decoderConfig: decoderConfig ?? undefined }; for await (const packet of sink.packets()) { if (this._canceled) { return; } if (packet.timestamp >= this._endTimestamp) { break; } const modifiedPacket = packet.clone({ timestamp: packet.timestamp - this._startTimestamp, }); assert(modifiedPacket.timestamp >= 0); this._reportProgress(outputTrackId, modifiedPacket.timestamp + modifiedPacket.duration); await source.add(modifiedPacket, meta); if (this._synchronizer.shouldWait(outputTrackId, modifiedPacket.timestamp)) { await this._synchronizer.wait(modifiedPacket.timestamp); } } source.close(); this._synchronizer.closeTrack(outputTrackId); })()); } else { // We need to decode & reencode the audio const canDecode = await track.canDecode(); if (!canDecode) { this.discardedTracks.push({ track, reason: 'undecodable_source_codec', trackOptions, }); return; } let codecOfChoice: AudioCodec | null = null; if (trackOptions.codec) { audioCodecs = audioCodecs.filter(codec => codec === trackOptions.codec); } const bitrate = trackOptions.bitrate ?? QUALITY_HIGH; const encodableCodecs = await getEncodableAudioCodecs(audioCodecs, { numberOfChannels: trackOptions.process && trackOptions.processedNumberOfChannels ? trackOptions.processedNumberOfChannels : numberOfChannels, sampleRate: trackOptions.process && trackOptions.processedSampleRate ? trackOptions.processedSampleRate : sampleRate, bitrate, }); if ( !encodableCodecs.some(codec => (NON_PCM_AUDIO_CODECS as readonly string[]).includes(codec)) && audioCodecs.some(codec => (NON_PCM_AUDIO_CODECS as readonly string[]).includes(codec)) && (numberOfChannels !== FALLBACK_NUMBER_OF_CHANNELS || sampleRate !== FALLBACK_SAMPLE_RATE) ) { // We could not find a compatible non-PCM codec despite the container supporting them. This can be // caused by strange channel count or sample rate configurations. Therefore, let's try again but with // fallback parameters. const encodableCodecsWithDefaultParams = await getEncodableAudioCodecs(audioCodecs, { numberOfChannels: FALLBACK_NUMBER_OF_CHANNELS, sampleRate: FALLBACK_SAMPLE_RATE, bitrate, }); const nonPcmCodec = encodableCodecsWithDefaultParams .find(codec => (NON_PCM_AUDIO_CODECS as readonly string[]).includes(codec)); if (nonPcmCodec) { // We are able to encode using a non-PCM codec, but it'll require resampling codecOfChoice = nonPcmCodec; numberOfChannels = FALLBACK_NUMBER_OF_CHANNELS; sampleRate = FALLBACK_SAMPLE_RATE; } } else { codecOfChoice = encodableCodecs[0] ?? null; } if (codecOfChoice === null) { this.discardedTracks.push({ track, reason: 'no_encodable_target_codec', trackOptions, }); return; } const encodingConfig: AudioEncodingConfig = { codec: codecOfChoice, bitrate, transform: { sampleFormat: trackOptions.sampleFormat, process: trackOptions.process, }, }; assert(encodingConfig.transform); if (numberOfChannels !== originalNumberOfChannels) { encodingConfig.transform.numberOfChannels = numberOfChannels; } if (sampleRate !== originalSampleRate) { encodingConfig.transform.sampleRate = sampleRate; } let lastSampleTimestamp: number | null = null; encodingConfig.onEncodedSample = (sample) => { lastSampleTimestamp = sample.timestamp; }; const source = new AudioSampleSource(encodingConfig); audioSource = source; this._trackPromises.push((async () => { await this._started; if (needsPadding) { const paddingLength = firstTimestamp - this._startTimestamp; const paddingLengthSamples = Math.round(paddingLength * originalSampleRate); const silentSample = new AudioSample({ data: new Float32Array(paddingLengthSamples * originalNumberOfChannels), format: 'f32-planar', numberOfChannels: originalNumberOfChannels, sampleRate: originalSampleRate, timestamp: 0, }); await this._registerAudioSample(silentSample, source, outputTrackId, () => lastSampleTimestamp); } const sink = new AudioSampleSink(track); for await (let sample of sink.samples(this._startTimestamp, this._endTimestamp)) { if (this._canceled) { sample.close(); return; } let startFrame = 0; let endFrame = sample.numberOfFrames; if (sample.timestamp < this._startTimestamp) { startFrame = Math.round((this._startTimestamp - sample.timestamp) * sample.sampleRate); } if (sample.timestamp + sample.duration > this._endTimestamp) { endFrame = Math.round((this._endTimestamp - sample.timestamp) * sample.sampleRate); } if (startFrame > 0 || endFrame < sample.numberOfFrames) { // Trim the sample if it sticks out of the trim region on either end const trimmedSample = sample.trim(startFrame, endFrame); sample.close(); sample = trimmedSample; if (sample.numberOfFrames === 0) { sample.close(); continue; } } // Offset the timestamp as needed sample.setTimestamp(sample.timestamp - this._startTimestamp); await this._registerAudioSample(sample, source, outputTrackId, () => lastSampleTimestamp); } source.close(); this._synchronizer.closeTrack(outputTrackId); })()); } let ownGroup: OutputTrackGroup | null = null; if (!trackOptions.group) { ownGroup = new OutputTrackGroup(); } const audioTrackLanguageCode = await track.getLanguageCode(); this.output.addAudioTrack(audioSource, { // TODO: This condition can be removed when all demuxers properly homogenize to BCP47 in v2 languageCode: isIso639Dash2LanguageCode(audioTrackLanguageCode) ? audioTrackLanguageCode : undefined, name: await track.getName() ?? undefined, disposition: await track.getDisposition(), group: ownGroup ?? trackOptions.group, }); this._addedCounts.audio++; this._totalTrackCount++; this.utilizedTracks.push(track); this._outputTrackIds.push(outputTrackId); this._outputOwnTrackGroups.push(ownGroup); } /** @internal */ async _registerAudioSample( sample: AudioSample, source: AudioSampleSource, outputTrackId: number, getLastSampleTimestamp: () => number | null, ) { this._reportProgress(outputTrackId, sample.timestamp + sample.duration); await source.add(sample); sample.close(); const lastSampleTimestamp = getLastSampleTimestamp(); if (lastSampleTimestamp !== null) { if (this._synchronizer.shouldWait(outputTrackId, lastSampleTimestamp)) { await this._synchronizer.wait(lastSampleTimestamp); } } } /** @internal */ _reportProgress(trackId: number, endTimestamp: number) { if (!this._computeProgress) { return; } assert(this._totalDuration !== null); this._maxTimestamps.set( trackId, Math.max(endTimestamp, this._maxTimestamps.get(trackId)!), ); const minTimestamp = Math.min(...this._maxTimestamps.values()); const newProgress = clamp(minTimestamp / this._totalDuration, 0, 1); if (newProgress !== this._lastProgress) { this._lastProgress = newProgress; this.onProgress?.(newProgress, minTimestamp); } } } /** * Thrown when a conversion couldn't complete due to being canceled. * @group Conversion * @public */ export class ConversionCanceledError extends Error { /** Creates a new {@link ConversionCanceledError}. */ constructor(message = 'Conversion has been canceled.') { super(message); this.name = 'ConversionCanceledError'; } } const MAX_TIMESTAMP_GAP = 1; // in seconds /** * Utility class for synchronizing multiple track packet consumers with one another. We don't want one consumer to get * too out-of-sync with the others, as that may lead to a large number of packets that need to be internally buffered * before they can be written. Therefore, we use this class to slow down a consumer if it is too far ahead of the * slowest consumer. */ class TrackSynchronizer { maxTimestamps = new Map(); // Track ID -> timestamp resolvers: { timestamp: number; resolve: () => void; }[] = []; declareTrack(trackId: number) { this.maxTimestamps.set(trackId, 0); } shouldWait(trackId: number, timestamp: number) { const currentValue = this.maxTimestamps.get(trackId); assert(currentValue !== undefined); this.maxTimestamps.set(trackId, Math.max(timestamp, currentValue)); const newMin = this.computeMinAndMaybeResolve(); return timestamp - newMin > MAX_TIMESTAMP_GAP; // Should wait if it is too far ahead of the slowest consumer } wait(timestamp: number) { const { promise, resolve } = promiseWithResolvers(); this.resolvers.push({ timestamp, resolve, }); return promise; } closeTrack(trackId: number) { this.maxTimestamps.delete(trackId); this.computeMinAndMaybeResolve(); } computeMinAndMaybeResolve() { let newMin = Infinity; for (const [, timestamp] of this.maxTimestamps) { newMin = Math.min(newMin, timestamp); } for (let i = 0; i < this.resolvers.length; i++) { const entry = this.resolvers[i]!; if (entry.timestamp - newMin < MAX_TIMESTAMP_GAP) { // The gap has gotten small enough again, the consumer can continue again entry.resolve(); this.resolvers.splice(i, 1); i--; } } return newMin; } } ===== src/mpeg-ts/mpeg-ts-muxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { buildAdtsHeaderTemplate, parseAacAudioSpecificConfig, writeAdtsFrameLength } from '../../shared/aac-misc'; import { validateAudioChunkMetadata, validateVideoChunkMetadata } from '../codec'; import { AC3_REGISTRATION_DESCRIPTOR, AvcDecoderConfigurationRecord, AvcNalUnitType, concatNalUnitsInAnnexB, deserializeAvcDecoderConfigurationRecord, deserializeHevcDecoderConfigurationRecord, EAC3_REGISTRATION_DESCRIPTOR, extractNalUnitTypeForAvc, extractNalUnitTypeForHevc, HevcDecoderConfigurationRecord, HevcNalUnitType, iterateNalUnitsInAnnexB, iterateNalUnitsInLengthPrefixed, } from '../codec-data'; import { Bitstream } from '../../shared/bitstream'; import { assert, promiseWithResolvers, setUint24, toDataView, toUint8Array } from '../misc'; import { Muxer } from '../muxer'; import { Output, OutputAudioTrack, OutputTrack, OutputVideoTrack } from '../output'; import { MpegTsOutputFormat } from '../output-format'; import { EncodedPacket } from '../packet'; import { Writer } from '../writer'; import { buildMpegTsMimeType, MpegTsStreamType, TIMESCALE, TS_PACKET_SIZE } from './mpeg-ts-misc'; // Resources: // ISO/IEC 13818-1 const PAT_PID = 0x0000; const PMT_PID = 0x1000; const FIRST_TRACK_PID = 0x0100; const VIDEO_STREAM_ID_BASE = 0xE0; const AUDIO_STREAM_ID_BASE = 0xC0; const AVC_AUD_NAL = new Uint8Array([0x09, 0xF0]); const HEVC_AUD_NAL = new Uint8Array([0x46, 0x01]); type MpegTsTrackData = { track: OutputVideoTrack | OutputAudioTrack; pid: number; streamType: MpegTsStreamType; streamId: number; codecString: string; timestampProcessingQueue: QueuedPacket[]; packetQueue: QueuedPacket[]; inputIsAnnexB: boolean | null; inputIsAdts: boolean | null; avcDecoderConfig: AvcDecoderConfigurationRecord | null; hevcDecoderConfig: HevcDecoderConfigurationRecord | null; adtsHeader: Uint8Array | null; adtsHeaderBitstream: Bitstream | null; firstPacketWritten: boolean; closed: boolean; }; type QueuedPacket = { data: Uint8Array; presentationTimestamp: number; decodeTimestamp: number | null; isKeyframe: boolean; }; export class MpegTsMuxer extends Muxer { private format: MpegTsOutputFormat; private writer!: Writer; private trackDatas: MpegTsTrackData[] = []; private tablesWritten = false; private continuityCounters = new Map(); private packetBuffer = new Uint8Array(TS_PACKET_SIZE); private packetView = toDataView(this.packetBuffer); private allTracksKnown = promiseWithResolvers(); private videoTrackIndex = 0; private audioTrackIndex = 0; private adaptationFieldBuffer = new Uint8Array(184); private payloadBuffer = new Uint8Array(184); constructor(output: Output, format: MpegTsOutputFormat) { super(output); this.format = format; } async start() { const release = await this.mutex.acquire(); this.writer = await this.output._getRootWriter(true); release(); } async getMimeType() { await this.allTracksKnown.promise; return buildMpegTsMimeType(this.trackDatas.map(x => x.codecString)); } private getVideoTrackData(track: OutputVideoTrack, meta?: EncodedVideoChunkMetadata) { const existingTrackData = this.trackDatas.find(x => x.track === track); if (existingTrackData) { return existingTrackData; } validateVideoChunkMetadata(meta); assert(meta?.decoderConfig); const codec = track.source._codec; assert(codec === 'avc' || codec === 'hevc'); const streamType = codec === 'avc' ? MpegTsStreamType.AVC : MpegTsStreamType.HEVC; const pid = FIRST_TRACK_PID + this.trackDatas.length; const streamId = VIDEO_STREAM_ID_BASE + this.videoTrackIndex++; const newTrackData: MpegTsTrackData = { track, pid, streamType, streamId, codecString: meta.decoderConfig.codec, timestampProcessingQueue: [], packetQueue: [], inputIsAnnexB: null, inputIsAdts: null, avcDecoderConfig: null, hevcDecoderConfig: null, adtsHeader: null, adtsHeaderBitstream: null, firstPacketWritten: false, closed: false, }; this.trackDatas.push(newTrackData); if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } return newTrackData; } private getAudioTrackData(track: OutputAudioTrack, meta?: EncodedAudioChunkMetadata) { const existingTrackData = this.trackDatas.find(x => x.track === track); if (existingTrackData) { return existingTrackData; } validateAudioChunkMetadata(meta); assert(meta?.decoderConfig); const codec = track.source._codec; assert(codec === 'aac' || codec === 'mp3' || codec === 'ac3' || codec === 'eac3'); let streamType: MpegTsStreamType; let streamId: number; switch (codec) { case 'aac': { streamType = MpegTsStreamType.AAC; streamId = AUDIO_STREAM_ID_BASE + this.audioTrackIndex++; }; break; case 'mp3': { streamType = MpegTsStreamType.MP3_MPEG1; streamId = AUDIO_STREAM_ID_BASE + this.audioTrackIndex++; }; break; case 'ac3': { streamType = MpegTsStreamType.AC3_SYSTEM_A; streamId = 0xbd; }; break; case 'eac3': { streamType = MpegTsStreamType.EAC3_SYSTEM_A; streamId = 0xbd; }; break; } const pid = FIRST_TRACK_PID + this.trackDatas.length; const newTrackData: MpegTsTrackData = { track, pid, streamType, streamId, codecString: meta.decoderConfig.codec, timestampProcessingQueue: [], packetQueue: [], inputIsAnnexB: null, inputIsAdts: null, avcDecoderConfig: null, hevcDecoderConfig: null, adtsHeader: null, adtsHeaderBitstream: null, firstPacketWritten: false, closed: false, }; this.trackDatas.push(newTrackData); if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } return newTrackData; } async addEncodedVideoPacket( track: OutputVideoTrack, packet: EncodedPacket, meta?: EncodedVideoChunkMetadata, ) { const release = await this.mutex.acquire(); try { const trackData = this.getVideoTrackData(track, meta); this.validateTimestamp( trackData.track, packet.timestamp, packet.type === 'key', ); const preparedData = this.prepareVideoPacket(trackData, packet, meta); if (packet.type === 'key') { await this.flushTimestampQueue(trackData); } trackData.timestampProcessingQueue.push({ data: preparedData, presentationTimestamp: packet.timestamp, decodeTimestamp: null, isKeyframe: packet.type === 'key', }); } finally { release(); } } async addEncodedAudioPacket( track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata, ) { const release = await this.mutex.acquire(); try { const trackData = this.getAudioTrackData(track, meta); this.validateTimestamp( trackData.track, packet.timestamp, packet.type === 'key', ); const preparedData = this.prepareAudioPacket(trackData, packet, meta); if (packet.type === 'key') { await this.flushTimestampQueue(trackData); } trackData.timestampProcessingQueue.push({ data: preparedData, presentationTimestamp: packet.timestamp, decodeTimestamp: null, isKeyframe: packet.type === 'key', }); } finally { release(); } } async addSubtitleCue(): Promise { throw new Error('MPEG-TS does not support subtitles.'); } private prepareVideoPacket( trackData: MpegTsTrackData, packet: EncodedPacket, meta?: EncodedVideoChunkMetadata, ): Uint8Array { const codec = (trackData.track as OutputVideoTrack).source._codec; if (trackData.inputIsAnnexB === null) { // This is the first packet const description = meta?.decoderConfig?.description; trackData.inputIsAnnexB = !description; if (!trackData.inputIsAnnexB) { const bytes = toUint8Array(description!); if (codec === 'avc') { trackData.avcDecoderConfig = deserializeAvcDecoderConfigurationRecord(bytes); } else { trackData.hevcDecoderConfig = deserializeHevcDecoderConfigurationRecord(bytes); } } } if (trackData.inputIsAnnexB) { return this.prepareAnnexBVideoPacket(packet.data, codec as 'avc' | 'hevc'); } else { return this.prepareLengthPrefixedVideoPacket(trackData, packet, codec as 'avc' | 'hevc'); } } private prepareAnnexBVideoPacket(data: Uint8Array, codec: 'avc' | 'hevc'): Uint8Array { const nalUnits: Uint8Array[] = []; for (const loc of iterateNalUnitsInAnnexB(data)) { const nalUnit = data.subarray(loc.offset, loc.offset + loc.length); const isAud = codec === 'avc' ? extractNalUnitTypeForAvc(nalUnit[0]!) === AvcNalUnitType.AUD : extractNalUnitTypeForHevc(nalUnit[0]!) === HevcNalUnitType.AUD_NUT; if (!isAud) { nalUnits.push(nalUnit); } } // Pretend the AUD const aud = codec === 'avc' ? AVC_AUD_NAL : HEVC_AUD_NAL; nalUnits.unshift(aud); return concatNalUnitsInAnnexB(nalUnits); } private prepareLengthPrefixedVideoPacket( trackData: MpegTsTrackData, packet: EncodedPacket, codec: 'avc' | 'hevc', ): Uint8Array { const data = packet.data; const lengthSize = codec === 'avc' ? (trackData.avcDecoderConfig!.lengthSizeMinusOne + 1) as 1 | 2 | 3 | 4 : (trackData.hevcDecoderConfig!.lengthSizeMinusOne + 1) as 1 | 2 | 3 | 4; const nalUnits: Uint8Array[] = []; for (const loc of iterateNalUnitsInLengthPrefixed(data, lengthSize)) { const nalUnit = data.subarray(loc.offset, loc.offset + loc.length); const isAud = codec === 'avc' ? extractNalUnitTypeForAvc(nalUnit[0]!) === AvcNalUnitType.AUD : extractNalUnitTypeForHevc(nalUnit[0]!) === HevcNalUnitType.AUD_NUT; if (!isAud) { nalUnits.push(nalUnit); } } if (packet.type === 'key') { // Add whichever NALUs are missing if (codec === 'avc') { const config = trackData.avcDecoderConfig!; for (const pps of config.pictureParameterSets) { nalUnits.unshift(pps); } for (const sps of config.sequenceParameterSets) { nalUnits.unshift(sps); } } else { const config = trackData.hevcDecoderConfig!; for (const arr of config.arrays) { if (arr.nalUnitType === HevcNalUnitType.PPS_NUT) { for (const nal of arr.nalUnits) { nalUnits.unshift(nal); } } } for (const arr of config.arrays) { if (arr.nalUnitType === HevcNalUnitType.SPS_NUT) { for (const nal of arr.nalUnits) { nalUnits.unshift(nal); } } } for (const arr of config.arrays) { if (arr.nalUnitType === HevcNalUnitType.VPS_NUT) { for (const nal of arr.nalUnits) { nalUnits.unshift(nal); } } } } } // Prepend the AUD const aud = codec === 'avc' ? AVC_AUD_NAL : HEVC_AUD_NAL; nalUnits.unshift(aud); return concatNalUnitsInAnnexB(nalUnits); } private prepareAudioPacket( trackData: MpegTsTrackData, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata, ): Uint8Array { const codec = (trackData.track as OutputAudioTrack).source._codec; if (codec === 'mp3' || codec === 'ac3' || codec === 'eac3') { // We're good return packet.data; } if (trackData.inputIsAdts === null) { // It's the first packet const description = meta?.decoderConfig?.description; trackData.inputIsAdts = !description; if (!trackData.inputIsAdts) { const config = parseAacAudioSpecificConfig(toUint8Array(description!)); const template = buildAdtsHeaderTemplate(config); trackData.adtsHeader = template.header; trackData.adtsHeaderBitstream = template.bitstream; } } if (trackData.inputIsAdts) { return packet.data; } assert(trackData.adtsHeader); assert(trackData.adtsHeaderBitstream); const header = trackData.adtsHeader; const frameLength = packet.data.byteLength + header.byteLength; writeAdtsFrameLength(trackData.adtsHeaderBitstream, frameLength); const result = new Uint8Array(frameLength); result.set(header, 0); result.set(packet.data, header.byteLength); return result; } private allTracksAreKnown() { for (const track of this.output._tracks) { if (!track.source._closed && !this.trackDatas.some(x => x.track === track)) { return false; } } return true; } private async flushTimestampQueue(trackData: MpegTsTrackData, alsoInterleave = true) { if (trackData.timestampProcessingQueue.length === 0) { return; } const sortedTimestamps = trackData.timestampProcessingQueue .map(packet => packet.presentationTimestamp) .sort((a, b) => a - b); for (let i = 0; i < trackData.timestampProcessingQueue.length; i++) { const queuedPacket = trackData.timestampProcessingQueue[i]!; queuedPacket.decodeTimestamp = sortedTimestamps[i]!; trackData.packetQueue.push(queuedPacket); } trackData.timestampProcessingQueue.length = 0; if (alsoInterleave) { await this.interleavePackets(); } } private async interleavePackets(isFinalCall = false) { if (!this.tablesWritten) { if (!this.allTracksAreKnown() && !isFinalCall) { return; } this.writeTables(); } outer: while (true) { let trackWithMinTimestamp: MpegTsTrackData | null = null; let minTimestamp = Infinity; for (const trackData of this.trackDatas) { if ( !isFinalCall && trackData.packetQueue.length === 0 && !trackData.closed ) { break outer; } if ( trackData.packetQueue.length > 0 && trackData.packetQueue[0]!.presentationTimestamp < minTimestamp ) { trackWithMinTimestamp = trackData; minTimestamp = trackData.packetQueue[0]!.presentationTimestamp; } } if (!trackWithMinTimestamp) { break; } const queuedPacket = trackWithMinTimestamp.packetQueue.shift()!; this.writePesPacket(trackWithMinTimestamp, queuedPacket); } if (!isFinalCall) { await this.writer.flush(); } } private writeTables() { assert(!this.tablesWritten); this.writePsiSection(PAT_PID, PAT_SECTION); this.writePsiSection(PMT_PID, buildPmt(this.trackDatas)); this.tablesWritten = true; } private writePsiSection(pid: number, section: Uint8Array) { let offset = 0; let isFirst = true; // Long PSI sections might span more than one TS packet while (offset < section.length) { const pointerFieldSize = isFirst ? 1 : 0; const availablePayload = 184 - pointerFieldSize; const remainingData = section.length - offset; const chunkSize = Math.min(availablePayload, remainingData); let payload: Uint8Array; if (isFirst) { payload = this.payloadBuffer.subarray(0, 1 + chunkSize); payload[0] = 0x00; // pointer_field payload.set(section.subarray(offset, offset + chunkSize), 1); } else { payload = section.subarray(offset, offset + chunkSize); } this.writeTsPacket(pid, isFirst, null, payload); offset += chunkSize; isFirst = false; } } private writePesPacket(trackData: MpegTsTrackData, queuedPacket: QueuedPacket) { const includeDts = trackData.track.type === 'video'; const headerDataLength = includeDts ? 10 : 5; const pesHeaderBuffer = new Uint8Array(9 + headerDataLength); const pesView = toDataView(pesHeaderBuffer); const ptsDtsBitstream = new Bitstream(pesHeaderBuffer.subarray(9)); setUint24(pesView, 0, 0x000001, false); // packet_start_code_prefix pesHeaderBuffer[3] = trackData.streamId; // stream_id const pesPacketLength = trackData.track.type === 'video' ? 0 // Unbounded : Math.min(8 + queuedPacket.data.length, 0xFFFF); // Required for audio for some reason pesView.setUint16(4, pesPacketLength, false); // '10' marker, PES_scrambling_control=0, PES_priority=0, // data_alignment_indicator=1, copyright=0, original_or_copy=0 pesView.setUint8(6, 0x84); pesView.setUint8(7, includeDts ? 0xC0 : 0x80); // PTS_DTS_flags, other flags=0 pesView.setUint8(8, headerDataLength); // PES_header_data_length const pts = Math.round(queuedPacket.presentationTimestamp * TIMESCALE); ptsDtsBitstream.pos = 0; ptsDtsBitstream.writeBits(4, includeDts ? 0b0011 : 0b0010); // marker ptsDtsBitstream.writeBits(3, (pts >>> 30) & 0x7); // PTS[32:30] ptsDtsBitstream.writeBits(1, 1); // marker_bit ptsDtsBitstream.writeBits(15, (pts >>> 15) & 0x7FFF); // PTS[29:15] ptsDtsBitstream.writeBits(1, 1); // marker_bit ptsDtsBitstream.writeBits(15, pts & 0x7FFF); // PTS[14:0] ptsDtsBitstream.writeBits(1, 1); // marker_bit if (includeDts) { assert(queuedPacket.decodeTimestamp !== null); const dts = Math.round(queuedPacket.decodeTimestamp * TIMESCALE); ptsDtsBitstream.writeBits(4, 0b0001); ptsDtsBitstream.writeBits(3, (dts >>> 30) & 0x7); // DTS[32:30] ptsDtsBitstream.writeBits(1, 1); // marker_bit ptsDtsBitstream.writeBits(15, (dts >>> 15) & 0x7FFF); // DTS[29:15] ptsDtsBitstream.writeBits(1, 1); // marker_bit ptsDtsBitstream.writeBits(15, dts & 0x7FFF); // DTS[14:0] ptsDtsBitstream.writeBits(1, 1); // marker_bit } const totalLength = pesHeaderBuffer.length + queuedPacket.data.length; let offset = 0; let isFirstTsPacket = true; while (offset < totalLength) { const pusi = isFirstTsPacket; const remainingData = totalLength - offset; const randomAccessIndicator = isFirstTsPacket && queuedPacket.isKeyframe; const discontinuityIndicator = isFirstTsPacket && !trackData.firstPacketWritten; const basePaddingNeeded = Math.max(0, 184 - remainingData); let adaptationFieldSize: number; if (randomAccessIndicator || discontinuityIndicator) { // We need at least two bytes adaptationFieldSize = Math.max(2, basePaddingNeeded); } else { adaptationFieldSize = basePaddingNeeded; } let adaptationField: Uint8Array | null = null; if (adaptationFieldSize > 0) { const buf = this.adaptationFieldBuffer; if (adaptationFieldSize === 1) { buf[0] = 0; // adaptation_field_length } else { buf[0] = adaptationFieldSize - 1; // adaptation_field_length buf[1] = (Number(discontinuityIndicator) << 7) // discontinuity_indicator | (Number(randomAccessIndicator) << 6); // random_access_indicator buf.fill(0xFF, 2, adaptationFieldSize); // stuffing_bytes } adaptationField = buf.subarray(0, adaptationFieldSize); } const payloadSize = Math.min(184 - adaptationFieldSize, remainingData); const payload = this.payloadBuffer.subarray(0, payloadSize); let payloadOffset = 0; if (offset < pesHeaderBuffer.length) { const headerBytes = Math.min(pesHeaderBuffer.length - offset, payloadSize); payload.set(pesHeaderBuffer.subarray(offset, offset + headerBytes), 0); payloadOffset = headerBytes; } const dataStart = Math.max(0, offset - pesHeaderBuffer.length); const dataEnd = dataStart + (payloadSize - payloadOffset); if (payloadOffset < payloadSize) { payload.set(queuedPacket.data.subarray(dataStart, dataEnd), payloadOffset); } this.writeTsPacket(trackData.pid, pusi, adaptationField, payload); offset += payloadSize; isFirstTsPacket = false; } trackData.firstPacketWritten = true; } private writeTsPacket( pid: number, pusi: boolean, adaptationField: Uint8Array | null, payload: Uint8Array, ) { const cc = this.continuityCounters.get(pid) ?? 0; const hasPayload = payload.length > 0; const adaptCtrl = adaptationField ? (hasPayload ? 0b11 : 0b10) : (hasPayload ? 0b01 : 0b00); this.packetBuffer[0] = 0x47; // sync_byte this.packetView.setUint16(1, (pusi ? 0x4000 : 0) | (pid & 0x1FFF), false); // TEI=0, PUSI, priority=0, PID // scrambling=0, adaptation_field_control, continuity_counter this.packetBuffer[3] = (adaptCtrl << 4) | (cc & 0x0F); if (hasPayload) { this.continuityCounters.set(pid, (cc + 1) & 0x0F); } let offset = 4; if (adaptationField) { this.packetBuffer.set(adaptationField, offset); offset += adaptationField.length; } this.packetBuffer.set(payload, offset); offset += payload.length; if (offset < TS_PACKET_SIZE) { this.packetBuffer.fill(0xFF, offset); // stuffing_bytes } const startPos = this.writer.getPos(); this.writer.write(this.packetBuffer); if (this.format._options.onPacket) { this.format._options.onPacket(this.packetBuffer.slice(), startPos); } } // eslint-disable-next-line @typescript-eslint/no-misused-promises override async onTrackClose(track: OutputTrack) { const release = await this.mutex.acquire(); const trackData = this.trackDatas.find(x => x.track === track); if (trackData) { trackData.closed = true; await this.flushTimestampQueue(trackData, false); } if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } await this.interleavePackets(); release(); } async finalize() { const release = await this.mutex.acquire(); this.allTracksKnown.resolve(); for (const trackData of this.trackDatas) { trackData.closed = true; await this.flushTimestampQueue(trackData, false); } await this.interleavePackets(true); release(); } } // CRC-32 for MPEG-TS (polynomial 0x04C11DB7, initial value 0xFFFFFFFF) const MPEG_TS_CRC_POLYNOMIAL = 0x04c11db7; const MPEG_TS_CRC_TABLE = new Uint32Array(256); for (let n = 0; n < 256; n++) { let crc = n << 24; for (let k = 0; k < 8; k++) { crc = (crc & 0x80000000) ? ((crc << 1) ^ MPEG_TS_CRC_POLYNOMIAL) : (crc << 1); } MPEG_TS_CRC_TABLE[n] = (crc >>> 0) & 0xffffffff; } const computeMpegTsCrc32 = (data: Uint8Array) => { let crc = 0xFFFFFFFF; for (let i = 0; i < data.length; i++) { const byte = data[i]!; crc = ((crc << 8) ^ MPEG_TS_CRC_TABLE[(crc >>> 24) ^ byte]!) >>> 0; } return crc; }; const PAT_SECTION = new Uint8Array(16); { const view = toDataView(PAT_SECTION); PAT_SECTION[0] = 0x00; // table_id view.setUint16(1, 0xB00D, false); // section_syntax_indicator=1, '0', reserved=11, section_length=13 view.setUint16(3, 0x0001, false); // transport_stream_id PAT_SECTION[5] = 0xC1; // reserved=11, version_number=0, current_next_indicator=1 PAT_SECTION[6] = 0x00; // section_number PAT_SECTION[7] = 0x00; // last_section_number view.setUint16(8, 0x0001, false); // program_number view.setUint16(10, 0xE000 | (PMT_PID & 0x1FFF), false); // reserved=111, program_map_PID view.setUint32(12, computeMpegTsCrc32(PAT_SECTION.subarray(0, 12)), false); // CRC_32 } const buildPmt = (trackDatas: MpegTsTrackData[]) => { let totalEsBytes = 0; for (const trackData of trackDatas) { totalEsBytes += 5; if (trackData.streamType === MpegTsStreamType.AC3_SYSTEM_A) { totalEsBytes += AC3_REGISTRATION_DESCRIPTOR.length; } else if (trackData.streamType === MpegTsStreamType.EAC3_SYSTEM_A) { totalEsBytes += EAC3_REGISTRATION_DESCRIPTOR.length; } } const sectionLength = 9 + totalEsBytes + 4; const section = new Uint8Array(3 + sectionLength - 4); const view = toDataView(section); section[0] = 0x02; // table_id // section_syntax_indicator=1, '0', reserved=11, section_length view.setUint16(1, 0xB000 | (sectionLength & 0x0FFF), false); view.setUint16(3, 0x0001, false); // program_number section[5] = 0xC1; // reserved=11, version_number=0, current_next_indicator=1 section[6] = 0x00; // section_number section[7] = 0x00; // last_section_number view.setUint16(8, 0xE000 | 0x1FFF, false); // reserved=111, PCR_PID=0x1FFF (none) view.setUint16(10, 0xF000, false); // reserved=1111, program_info_length=0 let offset = 12; for (const trackData of trackDatas) { section[offset++] = trackData.streamType; // stream_type view.setUint16(offset, 0xE000 | (trackData.pid & 0x1FFF), false); // reserved=111, elementary_PID offset += 2; if (trackData.streamType === MpegTsStreamType.AC3_SYSTEM_A) { view.setUint16(offset, 0xF000 | AC3_REGISTRATION_DESCRIPTOR.length, false); offset += 2; section.set(AC3_REGISTRATION_DESCRIPTOR, offset); offset += AC3_REGISTRATION_DESCRIPTOR.length; } else if (trackData.streamType === MpegTsStreamType.EAC3_SYSTEM_A) { view.setUint16(offset, 0xF000 | EAC3_REGISTRATION_DESCRIPTOR.length, false); offset += 2; section.set(EAC3_REGISTRATION_DESCRIPTOR, offset); offset += EAC3_REGISTRATION_DESCRIPTOR.length; } else { view.setUint16(offset, 0xF000, false); // reserved=1111, ES_info_length=0 offset += 2; } } const crc = computeMpegTsCrc32(section); const result = new Uint8Array(section.length + 4); result.set(section, 0); toDataView(result).setUint32(section.length, crc, false); // CRC_32 return result; }; ===== src/mpeg-ts/mpeg-ts-misc.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ export const TIMESCALE = 90_000; // MPEG-TS timestamps run on a 90 kHz clock export const TS_PACKET_SIZE = 188; export const enum MpegTsStreamType { MP3_MPEG1 = 0x03, MP3_MPEG2 = 0x04, AAC = 0x0f, AC3_SYSTEM_A = 0x81, EAC3_SYSTEM_A = 0x87, PRIVATE_DATA = 0x06, AVC = 0x1b, HEVC = 0x24, } export const buildMpegTsMimeType = (codecStrings: (string | null)[]) => { let string = 'video/MP2T'; const uniqueCodecStrings = [...new Set(codecStrings.filter(Boolean))]; if (uniqueCodecStrings.length > 0) { string += `; codecs="${uniqueCodecStrings.join(', ')}"`; } return string; }; ===== src/mpeg-ts/mpeg-ts-demuxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { TrackType } from '../output'; import { SAMPLES_PER_AAC_FRAME } from '../adts/adts-demuxer'; import { MAX_ADTS_FRAME_HEADER_SIZE, readAdtsFrameHeader } from '../adts/adts-reader'; import { aacChannelMap, aacFrequencyTable } from '../../shared/aac-misc'; import { AacCodecInfo, AudioCodec, extractAudioCodecString, extractVideoCodecString, MediaCodec, VideoCodec, } from '../codec'; import { AC3_ACMOD_CHANNEL_COUNTS, AC3_SAMPLES_PER_FRAME, AvcDecoderConfigurationRecord, AvcNalUnitType, determineVideoPacketType, extractAvcDecoderConfigurationRecord, extractHevcDecoderConfigurationRecord, EAC3_NUMBLKS_TABLE, getEac3ChannelCount, getEac3SampleRate, HevcDecoderConfigurationRecord, HevcNalUnitType, parseAc3SyncFrame, parseAvcSps, parseEac3SyncFrame, parseHevcSps, AC3_FRAME_SIZES, extractNalUnitTypeForAvc, extractNalUnitTypeForHevc, } from '../codec-data'; import { Demuxer } from '../demuxer'; import { Input } from '../input'; import { Logging } from '../logging'; import { InputAudioTrackBacking, InputTrackBacking, InputVideoTrackBacking, } from '../input-track'; import { PacketRetrievalOptions } from '../media-sink'; import { DEFAULT_TRACK_DISPOSITION, MetadataTags, TrackDisposition } from '../metadata'; import { assert, binarySearchExact, binarySearchLessOrEqual, COLOR_PRIMARIES_MAP_INVERSE, findLastIndex, floorToMultiple, last, MATRIX_COEFFICIENTS_MAP_INVERSE, readExpGolomb, Rotation, roundIfAlmostInteger, toDataView, TRANSFER_CHARACTERISTICS_MAP_INVERSE, UNDETERMINED_LANGUAGE, } from '../misc'; import { MP3_FRAME_HEADER_SIZE, getMp3ChannelCount, readMp3FrameHeader, } from '../../shared/mp3-misc'; import { EncodedPacket, PacketType, PLACEHOLDER_DATA } from '../packet'; import { FileSlice, readBytes, Reader, readU16Be, readU32Be, readU8 } from '../reader'; import { buildMpegTsMimeType, MpegTsStreamType, TIMESCALE, TS_PACKET_SIZE } from './mpeg-ts-misc'; import { AC3_SAMPLE_RATES } from '../../shared/ac3-misc'; import { Bitstream } from '../../shared/bitstream'; // Resources: // ISO/IEC 13818-1 const MISSING_PTS_ERROR_MESSAGE = 'PES packet is missing PTS where it was expected. PES packets without PTS are not' + ' currently supported. If you think this file should be supported, please report it.'; type ElementaryStream = { demuxer: MpegTsDemuxer; pid: number; streamType: number; initialized: boolean; firstSection: Section | null; /** * Some muxers suck ass and don't correctly label key frames, meaning we'll need to use our skill to * compensate for another programmer's skill issue. */ canBeTrustedWithKeyPackets: boolean; info: { type: 'video'; codec: VideoCodec; decoderConfig: VideoDecoderConfig | null; avcCodecInfo: AvcDecoderConfigurationRecord | null; hevcCodecInfo: HevcDecoderConfigurationRecord | null; colorSpace: VideoColorSpaceInit; width: number; height: number; squarePixelWidth: number; squarePixelHeight: number; reorderSize: number; } | { type: 'audio'; codec: AudioCodec; decoderConfig: AudioDecoderConfig | null; aacCodecInfo: AacCodecInfo | null; numberOfChannels: number; sampleRate: number; }; /** * Reference PES packets, spread throughout the file, to be used to speed up repeated random access. Sorted by both * byte offset and PTS. */ referencePesPackets: TimestampedPesPacketHeader[]; }; type ElementaryVideoStream = ElementaryStream & { info: { type: 'video' } }; type ElementaryAudioStream = ElementaryStream & { info: { type: 'audio' } }; type TsPacketHeader = { payloadUnitStartIndicator: number; pid: number; adaptationFieldControl: number; }; type TsPacket = TsPacketHeader & { body: Uint8Array; }; type Section = { startPos: number; endPos: number | null; // null if the section was not read fully pid: number; payload: Uint8Array; randomAccessIndicator: number; }; // Remember them so the warning doesn't get spammed const ignoredStreamTypes = new Set(); export class MpegTsDemuxer extends Demuxer { reader: Reader; metadataPromise: Promise | null = null; elementaryStreams: ElementaryStream[] = []; trackBackingEntries: InputTrackBacking[] = []; packetOffset = 0; packetStride = -1; sectionEndPositions: number[] = []; seekChunkSize = 5 * 1024 * 1024; // 5 MiB, picked because most HLS segments are below this size minReferencePointByteDistance = -1; constructor(input: Input) { super(input); this.reader = input._reader; } async readMetadata() { return this.metadataPromise ??= (async () => { const lengthToCheck = TS_PACKET_SIZE + 16 + 1; let startingSlice = this.reader.requestSlice(0, lengthToCheck); if (startingSlice instanceof Promise) startingSlice = await startingSlice; assert(startingSlice); const startingBytes = readBytes(startingSlice, lengthToCheck); if (startingBytes[0] === 0x47 && startingBytes[TS_PACKET_SIZE] === 0x47) { // Regular MPEG-TS this.packetOffset = 0; this.packetStride = TS_PACKET_SIZE; } else if (startingBytes[0] === 0x47 && startingBytes[TS_PACKET_SIZE + 16] === 0x47) { // MPEG-TS with Forward Error Correction this.packetOffset = 0; this.packetStride = TS_PACKET_SIZE + 16; } else if (startingBytes[4] === 0x47 && startingBytes[4 + TS_PACKET_SIZE + 4] === 0x47) { // MPEG-2-TS (DVHS) this.packetOffset = 4; this.packetStride = TS_PACKET_SIZE + 4; } else { throw new Error('Unreachable.'); } const MIN_REFERENCE_POINT_PACKET_DISTANCE = 256; this.minReferencePointByteDistance = MIN_REFERENCE_POINT_PACKET_DISTANCE * this.packetStride; let currentPos = this.packetOffset; let programMapPid: number | null = null; // Some files contain these multiple times, but we only care about their first appearance let hasProgramAssociationTable = false; let hasProgramMap = false; while (true) { const packetHeader = await this.readPacketHeader(currentPos); if (!packetHeader) { break; } if (packetHeader.payloadUnitStartIndicator === 0) { // Not the start of a section currentPos += this.packetStride; continue; } if (hasProgramMap && !this.elementaryStreams.some(x => x.pid === packetHeader.pid)) { // Don't care about this PID currentPos += this.packetStride; continue; } const section = await this.readSection( currentPos, true, !hasProgramMap, // Expect contiguous sections as long as we don't have the PMT ); if (!section) { break; } const BYTES_BEFORE_SECTION_LENGTH = 3; const BITS_IN_CRC_32 = 32; // Duh // Some streams don't contain a PAT for some reason, so we must do some guesswork to figure out where // the PMT is. let isProbablyProgramMap = false; if (!hasProgramMap && section.pid !== 0) { const isPesPacket = section.payload[0] === 0x00 && section.payload[1] === 0x00 && section.payload[2] === 0x01; if (!isPesPacket) { // Assume it's a PSI const bitstream = new Bitstream(section.payload); const pointerField = bitstream.readAlignedByte(); bitstream.skipBits(8 * pointerField); const tableId = bitstream.readBits(8); isProbablyProgramMap = tableId === 0x02; // 0x02 == TS_program_map_section } } if (section.pid === 0 && !hasProgramAssociationTable) { const bitstream = new Bitstream(section.payload); const pointerField = bitstream.readAlignedByte(); bitstream.skipBits(8 * pointerField); bitstream.skipBits(14); const sectionLength = bitstream.readBits(10); bitstream.skipBits(40); while (8 * (sectionLength + BYTES_BEFORE_SECTION_LENGTH) - bitstream.pos > BITS_IN_CRC_32) { const programNumber = bitstream.readBits(16); bitstream.skipBits(3); // Reserved const id = bitstream.readBits(13); if (programNumber !== 0) { if (programMapPid !== null) { throw new Error('Only files with a single program are supported.'); } else { programMapPid = id; } } } if (programMapPid === null) { throw new Error('Program Association Table must link to a Program Map Table.'); } hasProgramAssociationTable = true; } else if ((section.pid === programMapPid || isProbablyProgramMap) && !hasProgramMap) { const bitstream = new Bitstream(section.payload); const pointerField = bitstream.readAlignedByte(); bitstream.skipBits(8 * pointerField); bitstream.skipBits(12); const sectionLength = bitstream.readBits(12); bitstream.skipBits(43); // eslint-disable-next-line @typescript-eslint/no-unused-vars const pcrPid = bitstream.readBits(13); bitstream.skipBits(6); // "The remaining 10 bits specify the number of bytes of the descriptors immediately following the // program_info_length field" const programInfoLength = bitstream.readBits(10); bitstream.skipBits(8 * programInfoLength); while (8 * (sectionLength + BYTES_BEFORE_SECTION_LENGTH) - bitstream.pos > BITS_IN_CRC_32) { const streamType = bitstream.readBits(8); bitstream.skipBits(3); const elementaryPid = bitstream.readBits(13); bitstream.skipBits(6); const esInfoLength = bitstream.readBits(10); // Check ES descriptors to detect AC-3/E-AC-3 in System B const esInfoEndPos = bitstream.pos + 8 * esInfoLength; let hasAc3Descriptor = false; let hasEac3Descriptor = false; while (bitstream.pos < esInfoEndPos) { const descriptorTag = bitstream.readBits(8); const descriptorLength = bitstream.readBits(8); if (descriptorTag === 0x6a) { hasAc3Descriptor = true; } else if (descriptorTag === 0x7a || descriptorTag === 0xcc) { hasEac3Descriptor = true; } bitstream.skipBits(8 * descriptorLength); } let info: ElementaryStream['info'] | null = null; switch (streamType) { case MpegTsStreamType.AVC: case MpegTsStreamType.HEVC: { const codec = streamType === MpegTsStreamType.AVC ? 'avc' : 'hevc'; info = { type: 'video', codec, decoderConfig: null, avcCodecInfo: null, hevcCodecInfo: null, colorSpace: { primaries: null, transfer: null, matrix: null, fullRange: null, }, width: -1, height: -1, squarePixelWidth: -1, squarePixelHeight: -1, reorderSize: -1, }; }; break; case MpegTsStreamType.MP3_MPEG1: case MpegTsStreamType.MP3_MPEG2: case MpegTsStreamType.AAC: case MpegTsStreamType.AC3_SYSTEM_A: case MpegTsStreamType.EAC3_SYSTEM_A: { let codec: AudioCodec; if ( streamType === MpegTsStreamType.MP3_MPEG1 || streamType === MpegTsStreamType.MP3_MPEG2 ) { codec = 'mp3'; } else if (streamType === MpegTsStreamType.AAC) { codec = 'aac'; } else if (streamType === MpegTsStreamType.AC3_SYSTEM_A) { codec = 'ac3'; } else if (streamType === MpegTsStreamType.EAC3_SYSTEM_A) { codec = 'eac3'; } else { throw new Error('Unreachable.'); } info = { type: 'audio', codec, decoderConfig: null, aacCodecInfo: null, numberOfChannels: -1, sampleRate: -1, }; }; break; case MpegTsStreamType.PRIVATE_DATA: { if (hasEac3Descriptor) { info = { type: 'audio', codec: 'eac3', decoderConfig: null, aacCodecInfo: null, numberOfChannels: -1, sampleRate: -1, }; } else if (hasAc3Descriptor) { info = { type: 'audio', codec: 'ac3', decoderConfig: null, aacCodecInfo: null, numberOfChannels: -1, sampleRate: -1, }; } }; break; default: { // If we don't recognize the codec, we don't surface the track at all. This is because // we can't determine its metadata and also have no idea how to packetize its data. if (!ignoredStreamTypes.has(streamType)) { Logging._warn( `Note: MPEG-TS streams with stream_type 0x${streamType.toString(16)} are not` + ` currently supported.`, ); ignoredStreamTypes.add(streamType); } } } if (info) { this.elementaryStreams.push({ demuxer: this, pid: elementaryPid, streamType, initialized: false, firstSection: null, canBeTrustedWithKeyPackets: false, info, referencePesPackets: [], }); } } hasProgramMap = true; } else { const elementaryStream = this.elementaryStreams.find(x => x.pid === section.pid); outer: if (elementaryStream && !elementaryStream.initialized) { const pesPacket = readPesPacket(section, true); if (!pesPacket) { throw new Error( `Couldn't read first PES packet for Elementary Stream with PID ${elementaryStream.pid}`, ); } elementaryStream.firstSection = section; elementaryStream.canBeTrustedWithKeyPackets = section.randomAccessIndicator === 1; if (this.input._initInput) { const initDemuxer = (await this.input._initInput._getDemuxer()) as MpegTsDemuxer; const matchingStream = initDemuxer.elementaryStreams.find(x => ( x.pid === section.pid && x.info.codec === elementaryStream.info.codec )); if (matchingStream) { elementaryStream.info = matchingStream.info; elementaryStream.initialized = true; break outer; // We have the stream info, we're done } } const context = new PacketReadingContext(elementaryStream, pesPacket); if (elementaryStream.info.type === 'video') { // We loop because in some files, the video parameters are not in the first packet while (true) { const contextAlias = context; // TyyyyypeScript 😩 contextAlias.suppliedPacket = null; await context.markNextPacket(); if (elementaryStream.info.codec === 'avc') { if (!context.suppliedPacket) { throw new Error( 'Invalid AVC video stream; could not extract AVCDecoderConfigurationRecord' + ' from any packet.', ); } elementaryStream.info.avcCodecInfo = extractAvcDecoderConfigurationRecord(context.suppliedPacket.data); if (!elementaryStream.info.avcCodecInfo) { continue; // Search the next packet for it } const spsUnit = elementaryStream.info.avcCodecInfo.sequenceParameterSets[0]; assert(spsUnit); const spsInfo = parseAvcSps(spsUnit)!; elementaryStream.info.width = spsInfo.displayWidth; elementaryStream.info.height = spsInfo.displayHeight; const num = spsInfo.pixelAspectRatio.num; const den = spsInfo.pixelAspectRatio.den; if (num > 0 && den > 0) { if (num > den) { elementaryStream.info.squarePixelWidth = Math.round( elementaryStream.info.width * num / den, ); elementaryStream.info.squarePixelHeight = elementaryStream.info.height; } else { elementaryStream.info.squarePixelWidth = elementaryStream.info.width; elementaryStream.info.squarePixelHeight = Math.round( elementaryStream.info.height * den / num, ); } } elementaryStream.info.colorSpace = { primaries: COLOR_PRIMARIES_MAP_INVERSE[spsInfo.colourPrimaries] as VideoColorPrimaries | undefined, transfer: TRANSFER_CHARACTERISTICS_MAP_INVERSE[spsInfo.transferCharacteristics] as VideoTransferCharacteristics | undefined, matrix: MATRIX_COEFFICIENTS_MAP_INVERSE[spsInfo.matrixCoefficients] as VideoMatrixCoefficients | undefined, fullRange: !!spsInfo.fullRangeFlag, }; elementaryStream.info.reorderSize = spsInfo.maxDecFrameBuffering; break; } else if (elementaryStream.info.codec === 'hevc') { if (!context.suppliedPacket) { throw new Error( 'Invalid HEVC video stream; could not extract HVCDecoderConfigurationRecord' + ' from first packet.', ); } elementaryStream.info.hevcCodecInfo = extractHevcDecoderConfigurationRecord(context.suppliedPacket.data); if (!elementaryStream.info.hevcCodecInfo) { continue; // Search the next packet for it } const spsArray = elementaryStream.info.hevcCodecInfo.arrays.find( a => a.nalUnitType === HevcNalUnitType.SPS_NUT, )!; const spsUnit = spsArray.nalUnits[0]; assert(spsUnit); const spsInfo = parseHevcSps(spsUnit)!; elementaryStream.info.width = spsInfo.displayWidth; elementaryStream.info.height = spsInfo.displayHeight; if (spsInfo.pixelAspectRatio.num > spsInfo.pixelAspectRatio.den) { elementaryStream.info.squarePixelWidth = Math.round( elementaryStream.info.width * spsInfo.pixelAspectRatio.num / spsInfo.pixelAspectRatio.den, ); elementaryStream.info.squarePixelHeight = elementaryStream.info.height; } else { elementaryStream.info.squarePixelWidth = elementaryStream.info.width; elementaryStream.info.squarePixelHeight = Math.round( elementaryStream.info.height * spsInfo.pixelAspectRatio.den / spsInfo.pixelAspectRatio.num, ); } elementaryStream.info.colorSpace = { primaries: COLOR_PRIMARIES_MAP_INVERSE[spsInfo.colourPrimaries] as VideoColorPrimaries | undefined, transfer: TRANSFER_CHARACTERISTICS_MAP_INVERSE[spsInfo.transferCharacteristics] as VideoTransferCharacteristics | undefined, matrix: MATRIX_COEFFICIENTS_MAP_INVERSE[spsInfo.matrixCoefficients] as VideoMatrixCoefficients | undefined, fullRange: !!spsInfo.fullRangeFlag, }; elementaryStream.info.reorderSize = spsInfo.maxDecFrameBuffering; break; } else { throw new Error('Unhandled.'); } } elementaryStream.info.decoderConfig = { codec: extractVideoCodecString({ width: elementaryStream.info.width, height: elementaryStream.info.height, codec: elementaryStream.info.codec, codecDescription: null, colorSpace: elementaryStream.info.colorSpace, avcType: 1, avcCodecInfo: elementaryStream.info.avcCodecInfo, hevcCodecInfo: elementaryStream.info.hevcCodecInfo, vp9CodecInfo: null, av1CodecInfo: null, }), codedWidth: elementaryStream.info.width, codedHeight: elementaryStream.info.height, colorSpace: elementaryStream.info.colorSpace, }; if ( elementaryStream.info.width !== elementaryStream.info.squarePixelWidth || elementaryStream.info.height !== elementaryStream.info.squarePixelHeight ) { elementaryStream.info.decoderConfig.displayAspectWidth = elementaryStream.info.squarePixelWidth; elementaryStream.info.decoderConfig.displayAspectHeight = elementaryStream.info.squarePixelHeight; } elementaryStream.initialized = true; } else { await context.markNextPacket(); if (!context.suppliedPacket) { throw new Error( `Couldn't parse first media packet for Elementary Stream with` + ` PID ${elementaryStream.pid}`, ); } if (elementaryStream.info.codec === 'aac') { const slice = FileSlice.tempFromBytes(context.suppliedPacket.data); const header = readAdtsFrameHeader(slice); if (!header) { throw new Error( 'Invalid AAC audio stream; could not read ADTS frame header from first packet.', ); } elementaryStream.info.aacCodecInfo = { isMpeg2: false, objectType: header.objectType, }; elementaryStream.info.numberOfChannels = aacChannelMap[header.channelConfiguration]!; elementaryStream.info.sampleRate = aacFrequencyTable[header.samplingFrequencyIndex]!; } else if (elementaryStream.info.codec === 'mp3') { const word = readU32Be(FileSlice.tempFromBytes(context.suppliedPacket.data)); const result = readMp3FrameHeader(word, context.suppliedPacket.data.byteLength); if (!result.header) { throw new Error( 'Invalid MP3 audio stream; could not read frame header from first packet.', ); } elementaryStream.info.numberOfChannels = getMp3ChannelCount(result.header.channel); elementaryStream.info.sampleRate = result.header.sampleRate; } else if (elementaryStream.info.codec === 'ac3') { const frameInfo = parseAc3SyncFrame(context.suppliedPacket.data); if (!frameInfo) { throw new Error( 'Invalid AC-3 audio stream; could not read sync frame from first packet.', ); } if (frameInfo.fscod === 3) { throw new Error( 'Invalid AC-3 audio stream; reserved sample rate code found in first packet.', ); } elementaryStream.info.numberOfChannels = AC3_ACMOD_CHANNEL_COUNTS[frameInfo.acmod]! + frameInfo.lfeon; elementaryStream.info.sampleRate = AC3_SAMPLE_RATES[frameInfo.fscod]!; } else if (elementaryStream.info.codec === 'eac3') { const frameInfo = parseEac3SyncFrame(context.suppliedPacket.data); if (!frameInfo) { throw new Error( 'Invalid E-AC-3 audio stream; could not read sync frame from first packet.', ); } const sampleRate = getEac3SampleRate(frameInfo); if (sampleRate === null) { throw new Error( 'Invalid E-AC-3 audio stream; reserved sample rate code found in first packet.', ); } elementaryStream.info.numberOfChannels = getEac3ChannelCount(frameInfo); elementaryStream.info.sampleRate = sampleRate; } else { throw new Error('Unhandled.'); } elementaryStream.info.decoderConfig = { codec: extractAudioCodecString({ codec: elementaryStream.info.codec, codecDescription: null, aacCodecInfo: elementaryStream.info.aacCodecInfo, }), numberOfChannels: elementaryStream.info.numberOfChannels, sampleRate: elementaryStream.info.sampleRate, }; elementaryStream.initialized = true; } } } const isDone = hasProgramMap && this.elementaryStreams.every(x => x.initialized); if (isDone) { break; } currentPos += this.packetStride; } if (!hasProgramMap) { if (!hasProgramAssociationTable) { throw new Error('No Program Association Table found in the file.'); } throw new Error('No Program Map Table found in the file.'); } for (const stream of this.elementaryStreams) { if (stream.info.type === 'video') { this.trackBackingEntries.push( new MpegTsVideoTrackBacking(stream as ElementaryVideoStream), ); } else { this.trackBackingEntries.push( new MpegTsAudioTrackBacking(stream as ElementaryAudioStream), ); } } })(); } async getTrackBackings() { await this.readMetadata(); return this.trackBackingEntries; } async getMetadataTags(): Promise { return {}; // Nothing for now } async getMimeType(): Promise { await this.readMetadata(); const codecStrings = await Promise.all(this.trackBackingEntries.map( x => x.getDecoderConfig().then(c => c?.codec ?? null), )); return buildMpegTsMimeType(codecStrings); } async readSection(startPos: number, full: boolean, contiguous = false): Promise
{ let endPos = startPos; let currentPos = startPos; const chunks: Uint8Array[] = []; let chunksByteLength = 0; let firstPacket: TsPacket | null = null; let mustAddSectionEnd = true; let randomAccessIndicator = 0; while (true) { const packet = await this.readPacket(currentPos); currentPos += this.packetStride; if (!packet) { break; } if (!firstPacket) { if (packet.payloadUnitStartIndicator === 0) { break; } firstPacket = packet; } else { if (packet.pid !== firstPacket.pid) { if (contiguous) { break; // End of section } else { continue; // Ignore this packet } } if (packet.payloadUnitStartIndicator === 1) { break; } } const hasAdaptationField = !!(packet.adaptationFieldControl & 0b10); const hasPayload = !!(packet.adaptationFieldControl & 0b01); let adaptationFieldLength = 0; if (hasAdaptationField) { adaptationFieldLength = 1 + packet.body[0]!; // Extract random_access_indicator from first packet's adaptation field if (packet === firstPacket && adaptationFieldLength > 1) { randomAccessIndicator = (packet.body[1]! >> 6) & 1; } } if (hasPayload) { if (adaptationFieldLength === 0) { chunks.push(packet.body); chunksByteLength += packet.body.byteLength; } else { chunks.push(packet.body.subarray(adaptationFieldLength)); chunksByteLength += packet.body.byteLength - adaptationFieldLength; } } endPos = currentPos; // 64 is just "a bit of data", enough for the PES packet header if (!full && chunksByteLength >= 64) { mustAddSectionEnd = false; // Not the actual section end break; } // Check if we already know this is a section end const isKnownSectionEnd = binarySearchExact(this.sectionEndPositions, endPos, x => x) !== -1; if (isKnownSectionEnd) { mustAddSectionEnd = false; break; } } if (mustAddSectionEnd) { const index = binarySearchLessOrEqual(this.sectionEndPositions, endPos, x => x); this.sectionEndPositions.splice(index + 1, 0, endPos); } if (!firstPacket) { return null; } let merged: Uint8Array; if (chunks.length === 1) { merged = chunks[0]!; } else { const totalLength = chunks.reduce((sum, chunk) => sum + chunk.length, 0); merged = new Uint8Array(totalLength); let offset = 0; for (const chunk of chunks) { merged.set(chunk, offset); offset += chunk.length; } } return { startPos, endPos: full ? endPos : null, pid: firstPacket.pid, payload: merged, randomAccessIndicator, }; } async readPacketHeader(pos: number): Promise { let slice = this.reader.requestSlice(pos, 4); if (slice instanceof Promise) slice = await slice; if (!slice) { return null; } const syncByte = readU8(slice); if (syncByte !== 0x47) { throw new Error('Invalid TS packet sync byte. Likely an internal bug, please report this file.'); } const nextTwoBytes = readU16Be(slice); // eslint-disable-next-line @typescript-eslint/no-unused-vars const transportErrorIndicator = nextTwoBytes >> 15; const payloadUnitStartIndicator = (nextTwoBytes >> 14) & 0x1; // eslint-disable-next-line @typescript-eslint/no-unused-vars const transportPriority = (nextTwoBytes >> 13) & 0x1; const pid = nextTwoBytes & 0x1FFF; const nextByte = readU8(slice); // eslint-disable-next-line @typescript-eslint/no-unused-vars const transportScramblingControl = nextByte >> 6; const adaptationFieldControl = (nextByte >> 4) & 0x3; // eslint-disable-next-line @typescript-eslint/no-unused-vars const continuityCounter = nextByte & 0xF; return { payloadUnitStartIndicator, pid, adaptationFieldControl, }; } async readPacket(pos: number): Promise { // Code in here is duplicated from readPacketHeader for performance reasons let slice = this.reader.requestSlice(pos, TS_PACKET_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) { return null; } const bytes = readBytes(slice, TS_PACKET_SIZE); const syncByte = bytes[0]!; if (syncByte !== 0x47) { throw new Error('Invalid TS packet sync byte. Likely an internal bug, please report this file.'); } const nextTwoBytes = (bytes[1]! << 8) + bytes[2]!; // eslint-disable-next-line @typescript-eslint/no-unused-vars const transportErrorIndicator = nextTwoBytes >> 15; const payloadUnitStartIndicator = (nextTwoBytes >> 14) & 0x1; // eslint-disable-next-line @typescript-eslint/no-unused-vars const transportPriority = (nextTwoBytes >> 13) & 0x1; const pid = nextTwoBytes & 0x1FFF; const nextByte = bytes[3]!; // eslint-disable-next-line @typescript-eslint/no-unused-vars const transportScramblingControl = nextByte >> 6; const adaptationFieldControl = (nextByte >> 4) & 0x3; // eslint-disable-next-line @typescript-eslint/no-unused-vars const continuityCounter = nextByte & 0xF; return { payloadUnitStartIndicator, pid, adaptationFieldControl, body: bytes.subarray(4), }; } } type PesPacketHeader = { sectionStartPos: number; sectionEndPos: number | null; // null if the section wasn't read fully pts: number | null; randomAccessIndicator: number; }; type TimestampedPesPacketHeader = PesPacketHeader & { pts: number; }; type PesPacket = PesPacketHeader & { data: Uint8Array; }; type TimestampedPesPacket = PesPacket & { pts: number; }; const readPesPacketHeader = ( section: Section, expectPts: T, ): (T extends true ? TimestampedPesPacketHeader : PesPacketHeader) | null => { if (section.payload.byteLength < 3) { return null; } const bitstream = new Bitstream(section.payload); const startCodePrefix = bitstream.readBits(24); if (startCodePrefix !== 0x000001) { return null; } const streamId = bitstream.readBits(8); bitstream.skipBits(16); if ( streamId === 0b10111100 // program_stream_map || streamId === 0b10111110 // padding_stream || streamId === 0b10111111 // private_stream_2 || streamId === 0b11110000 // ECM || streamId === 0b11110001 // EMM || streamId === 0b11111111 // program_stream_directory || streamId === 0b11110010 // DSMCC_stream || streamId === 0b11111000 // ITU-T Rec. H.222.1 type E stream ) { return null; } bitstream.skipBits(8); const ptsDtsFlags = bitstream.readBits(2); bitstream.skipBits(14); let pts: number | null = null; if (ptsDtsFlags === 0b10 || ptsDtsFlags === 0b11) { pts = 0; bitstream.skipBits(4); pts += bitstream.readBits(3) * (1 << 30); bitstream.skipBits(1); pts += bitstream.readBits(15) * (1 << 15); bitstream.skipBits(1); pts += bitstream.readBits(15); } else { if (expectPts) { throw new Error(MISSING_PTS_ERROR_MESSAGE); } } return { sectionStartPos: section.startPos, sectionEndPos: section.endPos, pts, randomAccessIndicator: section.randomAccessIndicator, } as T extends true ? TimestampedPesPacketHeader : PesPacketHeader; }; const readPesPacket = ( section: Section, expectPts: T, ): (T extends true ? TimestampedPesPacket : PesPacket) | null => { assert(section.endPos !== null); // Can only read full PES packets from fully read sections const header = readPesPacketHeader(section, expectPts); if (!header) { return null; } const bitstream = new Bitstream(section.payload); bitstream.skipBits(32); const pesPacketLength = bitstream.readBits(16); const BYTES_UNTIL_END_OF_PES_PACKET_LENGTH = 6; bitstream.skipBits(16); const pesHeaderDataLength = bitstream.readBits(8); const pesHeaderEndPos = bitstream.pos + 8 * pesHeaderDataLength; bitstream.pos = pesHeaderEndPos; const bytePos = pesHeaderEndPos / 8; assert(Number.isInteger(bytePos)); const data = section.payload.subarray( bytePos, // "A value of 0 indicates that the PES packet length is neither specified nor bounded and is allowed only in // PES packets whose payload consists of bytes from a video elementary stream contained in // transport stream packets." pesPacketLength > 0 ? BYTES_UNTIL_END_OF_PES_PACKET_LENGTH + pesPacketLength : section.payload.byteLength, ); return { ...header, data, } as T extends true ? TimestampedPesPacket : PesPacket; }; abstract class MpegTsTrackBacking implements InputTrackBacking { packetBuffers = new WeakMap(); /** Used for recreating PacketBuffers if necessary. */ packetSectionStarts = new WeakMap(); constructor(public elementaryStream: ElementaryStream) {} abstract getType(): TrackType; abstract getDecoderConfig(): Promise; getId() { return this.elementaryStream.pid; } getNumber() { const demuxer = this.elementaryStream.demuxer; const trackType = this.elementaryStream.info.type; let number = 0; for (const backing of demuxer.trackBackingEntries) { if (backing.getType() === trackType) { number++; } assert(backing instanceof MpegTsTrackBacking); if (backing.elementaryStream === this.elementaryStream) { break; } } return number; } getCodec(): MediaCodec | null { throw new Error('Not implemented on base class.'); } getInternalCodecId() { return this.elementaryStream.streamType; } getName() { return null; } getLanguageCode() { return UNDETERMINED_LANGUAGE; } getDisposition(): TrackDisposition { return { ...DEFAULT_TRACK_DISPOSITION, primary: false, }; } getTimeResolution() { return TIMESCALE; } isRelativeToUnixEpoch() { return false; } getUnixTimeForTimestamp() { return null; } getPairingMask() { return 1n; } getBitrate() { return null; } getAverageBitrate() { return null; } async getDurationFromMetadata() { return null; } async getLiveRefreshInterval() { return null; } abstract allPacketsAreKeyPackets(): boolean; abstract getReorderSize(): number; createEncodedPacket( suppliedPacket: SuppliedPacket, duration: number, options: PacketRetrievalOptions, ) { let packetType: PacketType; if (this.allPacketsAreKeyPackets()) { packetType = 'key'; } else { packetType = suppliedPacket.randomAccessIndicator === 1 ? 'key' : 'delta'; } return new EncodedPacket( options.metadataOnly ? PLACEHOLDER_DATA : suppliedPacket.data, packetType, suppliedPacket.pts / TIMESCALE, Math.max(duration / TIMESCALE, 0), suppliedPacket.sequenceNumber, suppliedPacket.data.byteLength, ); } async getFirstPacket(options: PacketRetrievalOptions): Promise { const section = this.elementaryStream.firstSection; assert(section); const pesPacket = readPesPacket(section, true); assert(pesPacket); const context = new PacketReadingContext(this.elementaryStream, pesPacket); const buffer = new PacketBuffer(this, context); const result = await buffer.readNext(); if (!result) { return null; } const packet = this.createEncodedPacket(result.packet, result.duration, options); this.packetBuffers.set(packet, buffer); this.packetSectionStarts.set(packet, result.packet.sectionStartPos); return packet; } async getNextPacket(packet: EncodedPacket, options: PacketRetrievalOptions): Promise { let buffer = this.packetBuffers.get(packet); if (buffer) { // Fast path const result = await buffer.readNext(); if (!result) { return null; } // Remove PacketBuffer access from the old packet, it belongs to the next packet now this.packetBuffers.delete(packet); const newPacket = this.createEncodedPacket(result.packet, result.duration, options); this.packetBuffers.set(newPacket, buffer); this.packetSectionStarts.set(newPacket, result.packet.sectionStartPos); return newPacket; } // No buffer, we gotta do some rereading const sectionStartPos = this.packetSectionStarts.get(packet); if (sectionStartPos === undefined) { throw new Error('Packet was not created from this track.'); } const demuxer = this.elementaryStream.demuxer; const section = await demuxer.readSection(sectionStartPos, true); assert(section); const pesPacket = readPesPacket(section, true); assert(pesPacket); const context = new PacketReadingContext(this.elementaryStream, pesPacket); buffer = new PacketBuffer(this, context); // Advance until we pass the current packet's sequence number const targetSequenceNumber = packet.sequenceNumber; while (true) { const result = await buffer.readNext(); if (!result) { return null; } if (result.packet.sequenceNumber > targetSequenceNumber) { // We found the next packet! const newPacket = this.createEncodedPacket(result.packet, result.duration, options); this.packetBuffers.set(newPacket, buffer); this.packetSectionStarts.set(newPacket, result.packet.sectionStartPos); return newPacket; } } } async getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions): Promise { let currentPacket: EncodedPacket | null = packet; // Just loop until we hit one while (true) { currentPacket = await this.getNextPacket(currentPacket, options); if (!currentPacket) { return null; } if (currentPacket.type === 'key') { return currentPacket; } } } getPacket(timestamp: number, options: PacketRetrievalOptions): Promise { return this.doPacketLookup(timestamp, false, options); } getKeyPacket(timestamp: number, options: PacketRetrievalOptions): Promise { return this.doPacketLookup(timestamp, true, options); } /** * Searches for the packet with the largest timestamp not larger than `timestamp` in the file, using a combination * of chunk-based binary search and linear refinement. The reason the coarse search is done in large chunks is to * make it more performant for small files and over high-latency readers such as the network. */ async doPacketLookup( timestamp: number, keyframesOnly: boolean, options: PacketRetrievalOptions, ): Promise { const searchPts = roundIfAlmostInteger(timestamp * TIMESCALE); const demuxer = this.elementaryStream.demuxer; const { reader, seekChunkSize } = demuxer; const pid = this.elementaryStream.pid; const findFirstPesPacketHeaderInChunk = async ( startPos: number, endPos: number, readSectionInFull: boolean, ) => { let currentPos = startPos; while (currentPos < endPos) { const packetHeader = await demuxer.readPacketHeader(currentPos); if (!packetHeader) { return null; } if (packetHeader.pid === pid && packetHeader.payloadUnitStartIndicator === 1) { const section = await demuxer.readSection(currentPos, readSectionInFull); if (!section) { return null; } const pesPacketHeader = readPesPacketHeader(section, false); if (pesPacketHeader && pesPacketHeader.pts !== null) { return { pesPacketHeader: pesPacketHeader as TimestampedPesPacketHeader, section, }; } } currentPos += demuxer.packetStride; } return null; }; // Get the first PES packet of the track const firstSection = this.elementaryStream.firstSection; assert(firstSection); const firstPesPacketHeader = readPesPacketHeader(firstSection, true); assert(firstPesPacketHeader); if (searchPts < firstPesPacketHeader.pts) { // We're before the first packet, definitely nothing here return null; } let scanStartPos: number; const referencePesPackets = this.elementaryStream.referencePesPackets; const referencePointIndex = binarySearchLessOrEqual(referencePesPackets, searchPts, x => x.pts); const referencePoint = referencePointIndex !== -1 ? referencePesPackets[referencePointIndex]! : null; if (referencePoint && searchPts - referencePoint.pts < TIMESCALE / 2) { // Reference point ain't too far away, prefer it over the chunk search scanStartPos = referencePoint.sectionStartPos; } else { let startChunkIndex = 0; if (reader.fileSize !== null) { const numChunks = Math.ceil(reader.fileSize / seekChunkSize); if (numChunks > 1) { // Binary search to find the chunk with highest index whose first PES has pts <= searchPts let low = 0; let high = numChunks - 1; startChunkIndex = low; while (low <= high) { const mid = Math.floor((low + high) / 2); const chunkStartPos = floorToMultiple(mid * seekChunkSize, demuxer.packetStride) + firstPesPacketHeader.sectionStartPos; const chunkEndPos = chunkStartPos + seekChunkSize; const result = await findFirstPesPacketHeaderInChunk(chunkStartPos, chunkEndPos, false); if (!result) { // No PES packet found in this chunk, search left high = mid - 1; continue; } if (result.pesPacketHeader.pts <= searchPts) { // This chunk's first PES is <= searchPts, it's a candidate startChunkIndex = mid; low = mid + 1; // Search right } else { // Search left high = mid - 1; } } } } scanStartPos = floorToMultiple( startChunkIndex * seekChunkSize, demuxer.packetStride, ) + firstPesPacketHeader.sectionStartPos; } // Find the first PES packet at or after scanStartPos const result = await findFirstPesPacketHeaderInChunk( scanStartPos, reader.fileSize ?? Infinity, false, ); let currentPesHeader = result?.pesPacketHeader ?? null; if (!currentPesHeader) { // Fall back to first packet currentPesHeader = firstPesPacketHeader; } const reorderSize = this.getReorderSize(); const retrieveEncodedPacket = async ( sectionStartPos: number, predicate: (packet: SuppliedPacket) => boolean, ) => { // Load the relevant section in full const section = await demuxer.readSection(sectionStartPos, true); assert(section); const pesPacket = readPesPacket(section, true); assert(pesPacket); const context = new PacketReadingContext(this.elementaryStream, pesPacket); const buffer = new PacketBuffer(this, context); // Advance until the top-most presentation timestamp crosses or equals searchPts while (true) { const topPts = last(buffer.presentationOrderPackets)?.pts ?? -Infinity; if (topPts >= searchPts) { break; } const didRead = await buffer.readNextPacket(); if (!didRead) { break; } } const targetIndex = findLastIndex(buffer.presentationOrderPackets, predicate); if (targetIndex === -1) { return null; } const targetPacket = buffer.presentationOrderPackets[targetIndex]!; const lastDuration = targetIndex === 0 ? 0 : targetPacket.pts - buffer.presentationOrderPackets[targetIndex - 1]!.pts; // Pop packets in decode order until we hit the target packet while (buffer.decodeOrderPackets[0] !== targetPacket) { buffer.decodeOrderPackets.shift(); } buffer.lastDuration = lastDuration; // Kinda ugly but necessary fix const result = await buffer.readNext(); assert(result); const packet = this.createEncodedPacket(result.packet, result.duration, options); this.packetBuffers.set(packet, buffer); this.packetSectionStarts.set(packet, result.packet.sectionStartPos); return packet; }; if (!keyframesOnly || this.allPacketsAreKeyPackets()) { // Normat packet lookup case. Slightly easier since we just need to search (mostly) forward to find the // packet. // Linear scan to find the PES packet with largest pts <= searchPts. This will be used as the "midpoint" // of the next refinement step (which is needed because of B-frames). outer: while (true) { let currentPos = currentPesHeader.sectionStartPos + demuxer.packetStride; while (true) { const packetHeader = await demuxer.readPacketHeader(currentPos); if (!packetHeader) { break outer; // End of file } if (packetHeader.pid === pid && packetHeader.payloadUnitStartIndicator === 1) { const section = await demuxer.readSection(currentPos, false); if (section) { const nextPesHeader = readPesPacketHeader(section, false); if (nextPesHeader && nextPesHeader.pts !== null) { if (nextPesHeader.pts > searchPts) { break outer; } currentPesHeader = nextPesHeader as TimestampedPesPacketHeader; maybeInsertReferencePacket(this.elementaryStream, currentPesHeader); break; } } } currentPos += demuxer.packetStride; } } // Rewind by reorderSize + 1 PES packets (even for audio! To ensure proper durations) outer: for (let i = 0; i < reorderSize + 1; i++) { let pos = currentPesHeader.sectionStartPos - demuxer.packetStride; while (pos >= demuxer.packetOffset) { const packetHeader = await demuxer.readPacketHeader(pos); if (!packetHeader) { break outer; } if (packetHeader.pid === pid && packetHeader.payloadUnitStartIndicator === 1) { const section = await demuxer.readSection(pos, false); if (section) { const header = readPesPacketHeader(section, false); if (header && header.pts !== null) { currentPesHeader = header as TimestampedPesPacketHeader; break; } } } pos -= demuxer.packetStride; } } return retrieveEncodedPacket(currentPesHeader.sectionStartPos, p => p.pts <= searchPts); } else { // Key packet lookup case. Slightly harder since the starting chunk may not have a key packet at all, which // means we might need to search the previous chunks until we find something. let currentChunkStartPos = scanStartPos; let nextChunkStartPos: number | null = null; // "next" as in later in the file, even tho we scan backwards const readSectionsInFull = !this.elementaryStream.canBeTrustedWithKeyPackets; while (true) { let bestKeyPesHeader: TimestampedPesPacketHeader | null = null; const isFirstChunk = currentChunkStartPos <= firstPesPacketHeader.sectionStartPos; let pesHeader: TimestampedPesPacketHeader | null; let pesHeaderSection: Section | null = null; if (isFirstChunk) { pesHeader = firstPesPacketHeader; pesHeaderSection = firstSection; } else { const result = await findFirstPesPacketHeaderInChunk( currentChunkStartPos, reader.fileSize ?? Infinity, readSectionsInFull, ); pesHeader = result?.pesPacketHeader ?? null; pesHeaderSection = result?.section ?? null; } let passedSearchPts = false; let lookaheadCount = 0; outer: while (pesHeader) { if (nextChunkStartPos !== null && pesHeader.sectionStartPos >= nextChunkStartPos) { // Stop at the next chunk boundary break; } if (pesHeader.pts <= searchPts) { let isKeyPacket: boolean; if (this.elementaryStream.canBeTrustedWithKeyPackets) { isKeyPacket = pesHeader.randomAccessIndicator === 1; } else { assert(pesHeaderSection); const pesPacket = readPesPacket(pesHeaderSection, true); assert(pesPacket); const context = new PacketReadingContext(this.elementaryStream, pesPacket); await context.markNextPacket(); isKeyPacket = context.suppliedPacket?.randomAccessIndicator === 1; } if (isKeyPacket) { bestKeyPesHeader = pesHeader; } } if (pesHeader.pts > searchPts) { passedSearchPts = true; } // If we've passed searchPts, do lookahead for reorderSize more packets just to be sure if (passedSearchPts) { lookaheadCount++; if (lookaheadCount > reorderSize) { break; } } // Find next PES packet let currentPos = pesHeader.sectionStartPos + demuxer.packetStride; while (true) { const packetHeader = await demuxer.readPacketHeader(currentPos); if (!packetHeader) { break outer; // End of file } if (packetHeader.pid === pid && packetHeader.payloadUnitStartIndicator === 1) { const section = await demuxer.readSection(currentPos, readSectionsInFull); if (section) { const nextPesHeader = readPesPacketHeader(section, false); if (nextPesHeader && nextPesHeader.pts !== null) { pesHeader = nextPesHeader as TimestampedPesPacketHeader; pesHeaderSection = section; maybeInsertReferencePacket(this.elementaryStream, pesHeader); break; } } } currentPos += demuxer.packetStride; } } if (bestKeyPesHeader) { let startPesHeader = bestKeyPesHeader; if (lookaheadCount === 0) { // Packet is at the end of stream, let's rewind a little to obtain the correct packet duration outer: for (let i = 0; i < reorderSize; i++) { let pos = startPesHeader.sectionStartPos - demuxer.packetStride; while (pos >= demuxer.packetOffset) { const packetHeader = await demuxer.readPacketHeader(pos); if (!packetHeader) { break outer; } if (packetHeader.pid === pid && packetHeader.payloadUnitStartIndicator === 1) { const section = await demuxer.readSection(pos, readSectionsInFull); if (section) { const header = readPesPacketHeader(section, false); if (header && header.pts !== null) { startPesHeader = header as TimestampedPesPacketHeader; break; } } } pos -= demuxer.packetStride; } } } const encodedPacket = await retrieveEncodedPacket( startPesHeader.sectionStartPos, p => p.pts <= searchPts && p.randomAccessIndicator === 1, ); assert(encodedPacket); // There must be one return encodedPacket; } if (isFirstChunk) { return null; } // No key frame found in this chunk, move one chunk to the left nextChunkStartPos = currentChunkStartPos; currentChunkStartPos = Math.max( floorToMultiple( currentChunkStartPos - firstPesPacketHeader.sectionStartPos - seekChunkSize, demuxer.packetStride, ) + firstPesPacketHeader.sectionStartPos, firstPesPacketHeader.sectionStartPos, ); } } } } class MpegTsVideoTrackBacking extends MpegTsTrackBacking implements InputVideoTrackBacking { override elementaryStream!: ElementaryVideoStream; getType() { return 'video' as const; } override getCodec(): VideoCodec { return this.elementaryStream.info.codec; } getCodedWidth() { return this.elementaryStream.info.width; } getCodedHeight() { return this.elementaryStream.info.height; } getSquarePixelWidth() { return this.elementaryStream.info.squarePixelWidth; } getSquarePixelHeight() { return this.elementaryStream.info.squarePixelHeight; } getRotation(): Rotation { return 0; } async getColorSpace(): Promise { return this.elementaryStream.info.colorSpace; } async canBeTransparent() { return false; } async getDecoderConfig(): Promise { assert(this.elementaryStream.info.decoderConfig); return this.elementaryStream.info.decoderConfig; } override allPacketsAreKeyPackets(): boolean { return false; } override getReorderSize(): number { return this.elementaryStream.info.reorderSize; } } class MpegTsAudioTrackBacking extends MpegTsTrackBacking implements InputAudioTrackBacking { override elementaryStream!: ElementaryAudioStream; getType() { return 'audio' as const; } override getCodec(): AudioCodec { return this.elementaryStream.info.codec; } getNumberOfChannels() { return this.elementaryStream.info.numberOfChannels; } getSampleRate() { return this.elementaryStream.info.sampleRate; } async getDecoderConfig(): Promise { assert(this.elementaryStream.info.decoderConfig); return this.elementaryStream.info.decoderConfig; } override allPacketsAreKeyPackets(): boolean { return true; } override getReorderSize(): number { return 0; // No reordering, since no B-frames because goated } } const maybeInsertReferencePacket = ( elementaryStream: ElementaryStream, pesPacketHeader: TimestampedPesPacketHeader, ) => { const referencePesPackets = elementaryStream.referencePesPackets; const index = binarySearchLessOrEqual( referencePesPackets, pesPacketHeader.sectionStartPos, x => x.sectionStartPos, ); if (index >= 0) { // Since pts and file position don't necessarily have a monotonic relationship (since pts can go crazy), // let's see if inserting at the given index would violate the pts order. If so, return. const entry = referencePesPackets[index]!; if (pesPacketHeader.pts <= entry.pts) { return false; } const minByteDistance = elementaryStream.demuxer.minReferencePointByteDistance; if (pesPacketHeader.sectionStartPos - entry.sectionStartPos < minByteDistance) { // Too close return false; } if (index < referencePesPackets.length - 1) { const nextEntry = referencePesPackets[index + 1]!; if (nextEntry.pts < pesPacketHeader.pts) { // Out of order return false; } if (nextEntry.sectionStartPos - pesPacketHeader.sectionStartPos < minByteDistance) { // Too close return false; } } } referencePesPackets.splice(index + 1, 0, pesPacketHeader); return true; }; type SuppliedPacket = { pts: number; data: Uint8Array; sequenceNumber: number; sectionStartPos: number; randomAccessIndicator: number; }; /** Stateful context used to extract exact encoded packets from the underlying data stream. */ class PacketReadingContext { elementaryStream: ElementaryStream; pid: number; demuxer: MpegTsDemuxer; startingPesPacket: TimestampedPesPacket; currentPos = 0; // Relative to the data in startingPesPacket pesPackets: PesPacket[] = []; currentPesPacketIndex = 0; currentPesPacketPos = 0; endPos = 0; lastSuppliedPesPacket: PesPacket | null = null; nextPts: number | null = null; suppliedPacket: SuppliedPacket | null = null; constructor(elementaryStream: ElementaryStream, startingPesPacket: TimestampedPesPacket) { this.elementaryStream = elementaryStream; this.pid = elementaryStream.pid; this.demuxer = elementaryStream.demuxer; this.startingPesPacket = startingPesPacket; } ensureBuffered(length: number) { const remaining = this.endPos - this.currentPos; if (remaining >= length) { return length; } return this.bufferData(length - remaining) .then(() => Math.min(this.endPos - this.currentPos, length)); } getCurrentPesPacket() { const packet = this.pesPackets[this.currentPesPacketIndex]; assert(packet); return packet; } async bufferData(length: number): Promise { const targetEndPos = this.endPos + length; while (this.endPos < targetEndPos) { let pesPacket: PesPacket; if (this.pesPackets.length === 0) { pesPacket = this.startingPesPacket; } else { // Find the next PES packet let currentPos = last(this.pesPackets)!.sectionEndPos; assert(currentPos !== null); while (true) { const packetHeader = await this.demuxer.readPacketHeader(currentPos); if (!packetHeader) { return; } if (packetHeader.pid === this.pid) { const nextSection = await this.demuxer.readSection(currentPos, true); if (!nextSection) { return; } const nextPesPacket = readPesPacket(nextSection, false); if (nextPesPacket) { pesPacket = nextPesPacket; break; } } currentPos += this.demuxer.packetStride; } } this.pesPackets.push(pesPacket); this.endPos += pesPacket.data.byteLength; } } readBytes(length: number) { const currentPesPacket = this.getCurrentPesPacket(); const relativeStartOffset = this.currentPos - this.currentPesPacketPos; const relativeEndOffset = relativeStartOffset + length; this.currentPos += length; if (relativeEndOffset <= currentPesPacket.data.byteLength) { // Request can be satisfied with one PES packet return currentPesPacket.data.subarray(relativeStartOffset, relativeEndOffset); } // Data spans multiple PES packets, we must do some merging const result = new Uint8Array(length); result.set(currentPesPacket.data.subarray(relativeStartOffset)); let offset = currentPesPacket.data.byteLength - relativeStartOffset; while (true) { this.advanceCurrentPacket(); const currentPesPacket = this.getCurrentPesPacket(); const relativeEndOffset = length - offset; if (relativeEndOffset <= currentPesPacket.data.byteLength) { result.set(currentPesPacket.data.subarray(0, relativeEndOffset), offset); break; } result.set(currentPesPacket.data, offset); offset += currentPesPacket.data.byteLength; } return result; } readU8() { let currentPesPacket = this.getCurrentPesPacket(); const relativeOffset = this.currentPos - this.currentPesPacketPos; this.currentPos++; if (relativeOffset < currentPesPacket.data.byteLength) { return currentPesPacket.data[relativeOffset]!; } this.advanceCurrentPacket(); currentPesPacket = this.getCurrentPesPacket(); return currentPesPacket.data[0]!; } seekTo(pos: number) { if (pos === this.currentPos) { return; } if (pos < this.currentPos) { while (pos < this.currentPesPacketPos) { // Move to the previous PES packet this.currentPesPacketIndex--; const currentPacket = this.getCurrentPesPacket(); this.currentPesPacketPos -= currentPacket.data.byteLength; } } else { while (true) { // Move to the next PES packet const currentPesPacket = this.getCurrentPesPacket(); const currentEndPos = this.currentPesPacketPos + currentPesPacket.data.byteLength; if (pos < currentEndPos) { break; } this.currentPesPacketPos += currentPesPacket.data.byteLength; this.currentPesPacketIndex++; } } this.currentPos = pos; } skip(n: number) { this.seekTo(this.currentPos + n); } advanceCurrentPacket() { this.currentPesPacketPos += this.getCurrentPesPacket().data.byteLength; this.currentPesPacketIndex++; } async markNextPacket() { assert(!this.suppliedPacket); const elementaryStream = this.elementaryStream; if (elementaryStream.info.type === 'video') { // Our job here is to separate the video stream into access units. Sometimes this is easy (like when AUDs // are present), sometimes it's a little harder. const codec = elementaryStream.info.codec; const CHUNK_SIZE = 1024; if (codec !== 'avc' && codec !== 'hevc') { throw new Error('Unhandled.'); } const nalHeaderSize = codec === 'avc' ? 1 : 2; let packetStartPos: number | null = null; let frameStartFound = false; let lastFirstMacroblockInSlice = 0; while (true) { let remaining = this.ensureBuffered(CHUNK_SIZE); if (remaining instanceof Promise) remaining = await remaining; if (remaining === 0) { break; } const chunkStartPos = this.currentPos; const chunk = this.readBytes(remaining); const length = chunk.byteLength; let i = 0; while (i < length) { const zeroIndex = chunk.indexOf(0, i); if (zeroIndex === -1 || zeroIndex >= length) { break; } i = zeroIndex; const posBeforeZero = chunkStartPos + i; // Need 3 more bytes after the 0x00 to recognize a start code prefix if (i + 3 >= length) { // Not enough data in current chunk, seek back and let the next iteration handle it this.seekTo(posBeforeZero); break; } const b1 = chunk[i + 1]!; const b2 = chunk[i + 2]!; const b3 = chunk[i + 3]!; let startCodeLength = 0; // Check for 4-byte start code (0x00000001) if (b1 === 0x00 && b2 === 0x00 && b3 === 0x01) { startCodeLength = 4; } else if (b1 === 0x00 && b2 === 0x01) { // 3-byte start code (0x000001) startCodeLength = 3; } if (startCodeLength === 0) { // Not a start code, continue i++; continue; } const startCodePos = posBeforeZero; // The packet only really begins at the first NAL unit; anything before it isn't usable packetStartPos ??= startCodePos; const nalHeaderStart = i + startCodeLength; const payloadStart = nalHeaderStart + nalHeaderSize; // Bytes peeked from the start of a slice header to decode first_mb_in_slice. Six bytes (48 bits) // comfortably covers the exp-Golomb code for any realistic macroblock count const AVC_SLICE_HEADER_PEEK_SIZE = 6; // We read the NAL header plus, for slices, the start of the slice header to decode the // first_mb_in_slice / first_slice_segment_in_pic_flag. Make sure all of it is buffered. const bytesNeeded = payloadStart + (codec === 'avc' ? AVC_SLICE_HEADER_PEEK_SIZE : 1); if (bytesNeeded > length) { this.seekTo(posBeforeZero); break; } const headerByte0 = chunk[nalHeaderStart]!; let nalUnitType: number; let isSlice: boolean; let isAccessUnitStart: boolean; if (codec === 'avc') { nalUnitType = extractNalUnitTypeForAvc(headerByte0); isSlice = nalUnitType === AvcNalUnitType.NON_IDR_SLICE || nalUnitType === AvcNalUnitType.SLICE_DPA || nalUnitType === AvcNalUnitType.IDR; isAccessUnitStart = nalUnitType === AvcNalUnitType.SEI || nalUnitType === AvcNalUnitType.SPS || nalUnitType === AvcNalUnitType.PPS || nalUnitType === AvcNalUnitType.AUD; } else { nalUnitType = extractNalUnitTypeForHevc(headerByte0); const layerId = ((headerByte0 & 1) << 5) | (chunk[nalHeaderStart + 1]! >> 3); if (layerId > 0) { // Higher layers don't delimit the base-layer frames we care about i += startCodeLength; continue; } // VCL slices: 0..RASL_R, plus the IRAP range BLA_W_LP..CRA_NUT isSlice = nalUnitType <= HevcNalUnitType.RASL_R || (nalUnitType >= HevcNalUnitType.BLA_W_LP && nalUnitType <= 21); // VPS..FD, prefix SEI, and the reserved/unspecified non-VCL ranges isAccessUnitStart = (nalUnitType >= HevcNalUnitType.VPS_NUT && nalUnitType <= 37) || nalUnitType === HevcNalUnitType.PREFIX_SEI_NUT || (nalUnitType >= 41 && nalUnitType <= 44) || (nalUnitType >= 48 && nalUnitType <= 55); } let isFrameBoundary = false; if (isSlice) { let startsNewPicture: boolean; if (codec === 'avc') { const headerBytes = chunk.subarray(payloadStart, payloadStart + AVC_SLICE_HEADER_PEEK_SIZE); const firstMacroblockInSlice = readExpGolomb(new Bitstream(headerBytes)); startsNewPicture = !frameStartFound || firstMacroblockInSlice <= lastFirstMacroblockInSlice; lastFirstMacroblockInSlice = firstMacroblockInSlice; } else { startsNewPicture = (chunk[payloadStart]! >> 7) === 1; } if (startsNewPicture) { if (frameStartFound) { isFrameBoundary = true; } else { frameStartFound = true; } } } else if (isAccessUnitStart && frameStartFound) { isFrameBoundary = true; } if (isFrameBoundary) { // End the packet at this start code (the next frame begins here) const packetLength = startCodePos - packetStartPos; this.seekTo(packetStartPos); return this.supplyPacket(packetLength, 0); } i += startCodeLength; } if (remaining < CHUNK_SIZE) { // End of stream break; } } // End of stream - emit whatever's left as the final packet if (packetStartPos !== null && this.endPos > packetStartPos) { const packetLength = this.endPos - packetStartPos; this.seekTo(packetStartPos); return this.supplyPacket(packetLength, 0); } } else { const codec = elementaryStream.info.codec; const CHUNK_SIZE = 128; while (true) { let remaining = this.ensureBuffered(CHUNK_SIZE); if (remaining instanceof Promise) remaining = await remaining; const startPos = this.currentPos; while (this.currentPos - startPos < remaining) { const byte = this.readU8(); if (codec === 'aac') { if (byte !== 0xff) { continue; } this.skip(-1); const possibleHeaderStartPos = this.currentPos; let remaining = this.ensureBuffered(MAX_ADTS_FRAME_HEADER_SIZE); if (remaining instanceof Promise) remaining = await remaining; if (remaining < MAX_ADTS_FRAME_HEADER_SIZE) { return; } const headerBytes = this.readBytes(MAX_ADTS_FRAME_HEADER_SIZE); const header = readAdtsFrameHeader(FileSlice.tempFromBytes(headerBytes)); if (header) { this.seekTo(possibleHeaderStartPos); let remaining = this.ensureBuffered(header.frameLength); if (remaining instanceof Promise) remaining = await remaining; return this.supplyPacket( remaining, Math.round(SAMPLES_PER_AAC_FRAME * TIMESCALE / elementaryStream.info.sampleRate), ); } else { this.seekTo(possibleHeaderStartPos + 1); } } else if (codec === 'mp3') { if (byte !== 0xff) { continue; } this.skip(-1); const possibleHeaderStartPos = this.currentPos; let remaining = this.ensureBuffered(MP3_FRAME_HEADER_SIZE); if (remaining instanceof Promise) remaining = await remaining; if (remaining < MP3_FRAME_HEADER_SIZE) { return; } const headerBytes = this.readBytes(MP3_FRAME_HEADER_SIZE); const word = toDataView(headerBytes).getUint32(0); const result = readMp3FrameHeader(word, null); if (result.header) { this.seekTo(possibleHeaderStartPos); let remaining = this.ensureBuffered(result.header.totalSize); if (remaining instanceof Promise) remaining = await remaining; const duration = result.header.audioSamplesInFrame * TIMESCALE / elementaryStream.info.sampleRate; return this.supplyPacket(remaining, Math.round(duration)); } else { this.seekTo(possibleHeaderStartPos + 1); } } else if (codec === 'ac3') { if (byte !== 0x0b) { continue; } this.skip(-1); const possibleSyncPos = this.currentPos; // Need at least 5 bytes for sync word + CRC + fscod/frmsizecod let remaining = this.ensureBuffered(5); if (remaining instanceof Promise) remaining = await remaining; if (remaining < 5) { return; } const headerBytes = this.readBytes(5); // Verify sync word (0x0B77) if (headerBytes[0] !== 0x0b || headerBytes[1] !== 0x77) { this.seekTo(possibleSyncPos + 1); continue; } const fscod = headerBytes[4]! >> 6; const frmsizecod = headerBytes[4]! & 0x3f; if (fscod === 3 || frmsizecod > 37) { // Invalid this.seekTo(possibleSyncPos + 1); continue; } const frameSize = AC3_FRAME_SIZES[3 * frmsizecod + fscod]; assert(frameSize !== undefined); this.seekTo(possibleSyncPos); remaining = this.ensureBuffered(frameSize); if (remaining instanceof Promise) remaining = await remaining; const duration = Math.round( AC3_SAMPLES_PER_FRAME * TIMESCALE / elementaryStream.info.sampleRate, ); return this.supplyPacket(remaining, duration); } else if (codec === 'eac3') { if (byte !== 0x0b) { continue; } this.skip(-1); const possibleSyncPos = this.currentPos; // Need at least 5 bytes for E-AC-3 header parsing (sync word + frmsiz + fscod/numblkscod) let remaining = this.ensureBuffered(5); if (remaining instanceof Promise) remaining = await remaining; if (remaining < 5) { return; } const headerBytes = this.readBytes(5); if (headerBytes[0] !== 0x0b || headerBytes[1] !== 0x77) { this.seekTo(possibleSyncPos + 1); continue; } const frmsiz = ((headerBytes[2]! & 0x07) << 8) | headerBytes[3]!; const frameSize = (frmsiz + 1) * 2; const fscod = headerBytes[4]! >> 6; const numblkscod = fscod === 3 ? 3 : (headerBytes[4]! >> 4) & 0x03; const numblks = EAC3_NUMBLKS_TABLE[numblkscod]!; this.seekTo(possibleSyncPos); remaining = this.ensureBuffered(frameSize); if (remaining instanceof Promise) remaining = await remaining; // Duration = numblks * 256 samples per block const samplesPerFrame = numblks * 256; const duration = Math.round( samplesPerFrame * TIMESCALE / elementaryStream.info.sampleRate, ); return this.supplyPacket(remaining, duration); } else { throw new Error('Unhandled.'); } } if (remaining < CHUNK_SIZE) { break; } } } } /** Supplies the context with a new encoded packet, beginning at the current position. */ supplyPacket(packetLength: number, intrinsicDuration: number) { const currentPesPacket = this.getCurrentPesPacket(); let pts: number; if (this.lastSuppliedPesPacket === currentPesPacket) { assert(this.nextPts !== null); pts = this.nextPts; } else { if (currentPesPacket.pts === null) { throw new Error(MISSING_PTS_ERROR_MESSAGE); } pts = currentPesPacket.pts; maybeInsertReferencePacket(this.elementaryStream, currentPesPacket as TimestampedPesPacket); } this.lastSuppliedPesPacket = currentPesPacket; this.nextPts = pts + intrinsicDuration; const sectionStartPos = currentPesPacket.sectionStartPos; // The sequence number is the starting position of the section the PES packet is in, PLUS the offset within the // PES packet where the packet starts. const sequenceNumber = sectionStartPos + (this.currentPos - this.currentPesPacketPos); const data = this.readBytes(packetLength); let randomAccessIndicator = currentPesPacket.randomAccessIndicator; if (randomAccessIndicator === 0 && !this.elementaryStream.canBeTrustedWithKeyPackets) { if (this.elementaryStream.info.type === 'audio') { randomAccessIndicator = 1; } else { if (this.elementaryStream.info.decoderConfig) { const isKey = determineVideoPacketType( this.elementaryStream.info.codec, this.elementaryStream.info.decoderConfig, data, ) === 'key'; randomAccessIndicator = Number(isKey); } else { // We're reading packets before the decoder config is determined } } } this.suppliedPacket = { pts, data, sequenceNumber, sectionStartPos, randomAccessIndicator, }; this.pesPackets.splice(0, this.currentPesPacketIndex); this.currentPesPacketIndex = 0; } } /** * A buffer that simulates decoder frame reordering to compute packet durations. Packets arrive in decode order but * durations are based on presentation order. */ class PacketBuffer { backing: MpegTsTrackBacking; context: PacketReadingContext; decodeOrderPackets: SuppliedPacket[] = []; reorderSize: number; reorderBuffer: SuppliedPacket[] = []; presentationOrderPackets: SuppliedPacket[] = []; reachedEnd = false; lastDuration = 0; constructor(backing: MpegTsTrackBacking, context: PacketReadingContext) { this.backing = backing; this.context = context; this.reorderSize = backing.getReorderSize(); assert(this.reorderSize >= 0); } async readNext(): Promise<{ packet: SuppliedPacket; duration: number } | null> { if (this.decodeOrderPackets.length === 0) { // We need the next packet const didRead = await this.readNextPacket(); if (!didRead) { return null; } } // Ensure we know the next packet in presentation order so we can compute the current packet's duration await this.ensureCurrentPacketHasNext(); const packet = this.decodeOrderPackets[0]!; // Let's compute the duration const presentationIndex = this.presentationOrderPackets.indexOf(packet); assert(presentationIndex !== -1); let duration: number; if (presentationIndex === this.presentationOrderPackets.length - 1) { duration = this.lastDuration; // Reasonable heuristic } else { const nextPacket = this.presentationOrderPackets[presentationIndex + 1]!; duration = nextPacket.pts - packet.pts; this.lastDuration = duration; } this.decodeOrderPackets.shift(); // Shrink the presentation array as much as possible while (this.presentationOrderPackets.length > 0) { const first = this.presentationOrderPackets[0]!; if (this.decodeOrderPackets.includes(first)) { break; } this.presentationOrderPackets.shift(); } return { packet, duration }; } async readNextPacket() { if (this.reachedEnd) { return false; } let suppliedPacket: SuppliedPacket | null; if (this.context.suppliedPacket) { // Small optimization: there was already a supplied packet in the context, so let's first use that one suppliedPacket = this.context.suppliedPacket; } else { await this.context.markNextPacket(); suppliedPacket = this.context.suppliedPacket; } this.context.suppliedPacket = null; if (!suppliedPacket) { this.reachedEnd = true; this.flushReorderBuffer(); return false; } this.decodeOrderPackets.push(suppliedPacket); this.processPacketThroughReorderBuffer(suppliedPacket); return true; } async ensureCurrentPacketHasNext() { const current = this.decodeOrderPackets[0]; assert(current); while (true) { const presentationIndex = this.presentationOrderPackets.indexOf(current); // Check if current packet has a next packet if (presentationIndex !== -1 && presentationIndex <= this.presentationOrderPackets.length - 2) { break; } const didRead = await this.readNextPacket(); if (!didRead) { break; } } } processPacketThroughReorderBuffer(packet: SuppliedPacket) { this.reorderBuffer.push(packet); // If buffer is now overfull, output the packet with smallest PTS if (this.reorderBuffer.length > this.reorderSize) { let minIndex = 0; for (let i = 1; i < this.reorderBuffer.length; i++) { if (this.reorderBuffer[i]!.pts < this.reorderBuffer[minIndex]!.pts) { minIndex = i; } } const packet = this.reorderBuffer[minIndex]!; this.presentationOrderPackets.push(packet); this.reorderBuffer.splice(minIndex, 1); } } flushReorderBuffer() { this.reorderBuffer.sort((a, b) => a.pts - b.pts); this.presentationOrderPackets.push(...this.reorderBuffer); this.reorderBuffer.length = 0; } } ===== src/hls/hls-segmented-input.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AES_128_BLOCK_SIZE, createAes128CbcDecryptStream } from '../aes'; import { ENCRYPTION_KEY_CACHE_GROUP, Input } from '../input'; import { Segment, SegmentedInput, SegmentedInputTrackDeclaration, SegmentRetrievalOptions } from '../segmented-input'; import { toDataView, joinPaths, last, assert, binarySearchLessOrEqual, arrayArgmin, wait, base64ToBytes, } from '../misc'; import { readAllLines, readBytes, Reader } from '../reader'; import { CustomPathedSource, PathedSource, ReadableStreamSource, SourceRef, SourceRequest } from '../source'; import { HlsDemuxer } from './hls-demuxer'; import { AttributeList, canIgnoreLine, TAG_BYTERANGE, TAG_DISCONTINUITY, TAG_ENDLIST, TAG_EXTINF, TAG_KEY, TAG_MAP, TAG_MEDIA_SEQUENCE, TAG_PLAYLIST_TYPE, TAG_PROGRAM_DATE_TIME, TAG_TARGETDURATION, } from './hls-misc'; import { HlsInputFormat, type InputFormatOptions } from '../input-format'; import { parsePsshBoxContents, psshBoxesAreEqual, type PsshBox } from '../isobmff/isobmff-misc'; const IV_STRING_REGEX = /^0[xX][0-9a-fA-F]+$/; const BASE64_DATA_URI_REGEX = /^data:.*;base64,/i; export type HlsSegment = Segment & { sequenceNumber: number | null; location: HlsSegmentLocation; encryption: HlsEncryptionInfo | null; firstSegment: HlsSegment | null; initSegment: HlsSegment | null; lastProgramDateTimeSeconds: number | null; }; export type HlsEncryptionInfo = { method: 'AES-128'; keyUri: string; iv: Uint8Array | null; keyFormat: string; } | { method: 'SAMPLE-AES' | 'SAMPLE-AES-CTR'; psshBox: PsshBox | null; }; export type HlsSegmentLocation = { path: string; offset: number; length: number | null; }; export class HlsSegmentedInput extends SegmentedInput { rootPath: string; demuxer: HlsDemuxer; segments: HlsSegment[] = []; nextLines: string[] | null = null; currentUpdateSegmentsPromise: Promise | null = null; streamHasEnded = false; lastSegmentUpdateTime = -Infinity; refreshInterval = 5; // Reasonable default in case the playlist doesn't specify it constructor( demuxer: HlsDemuxer, path: string, trackDeclarations: SegmentedInputTrackDeclaration[] | null, lines: string[] | null, ) { super(demuxer.input, path, trackDeclarations); this.rootPath = path; this.demuxer = demuxer; this.nextLines = lines; } runUpdateSegments() { return this.currentUpdateSegmentsPromise ??= (async () => { try { const remainingWaitTimeMs = this.getRemainingWaitTimeMs(); if (remainingWaitTimeMs > 0) { await wait(remainingWaitTimeMs); } this.lastSegmentUpdateTime = performance.now(); await this.updateSegments(); } finally { this.currentUpdateSegmentsPromise = null; } })(); } getRemainingWaitTimeMs() { const elapsed = performance.now() - this.lastSegmentUpdateTime; const result = Math.max(0, 1000 * this.refreshInterval - elapsed); if (result <= 50) { // If only a little bit of time is left, don't wait at all; this removes the chance for timing race // conditions when running a task every `refreshInterval` seconds return 0; } return result; } /** * Reads and parses the segment info from the playlist file. When called more than one, it updates the existing * segments by appending the new ones. Existing segments are never removed. */ async updateSegments() { let lines = this.nextLines; this.nextLines = null; if (!lines) { using ref = await this.demuxer.input._getSourceUncached({ path: this.rootPath, isRoot: false }); const reader = new Reader(ref.source); const slice = await reader.requestEntireFile(); assert(slice); lines = readAllLines(slice, slice.length, { ignore: canIgnoreLine }); if (ref.source instanceof PathedSource) { // Copy back the source's path to become aware of potential redirects this.rootPath = ref.source.rootPath; } } const offsetTimestampsByDateTime = this.input._formatOptions.hls?.offsetTimestampsByDateTime !== false; let headerRead = false; let accumulatedTime = 0; let accumulatedUnixTime: number | null = null; let nextSegmentDuration: number | null = null; let currentKey: HlsEncryptionInfo | null = null; let nextSequenceNumber = 0; let currentFirstSegment: HlsSegment | null = null; let currentInitSegment: HlsSegment | null = null; let lastByteRangeEnd: number | null = null; let nextByteRange: { offset: number; length: number } | null = null; let lastProgramDateTimeSeconds: number | null = null; let targetDuration: number | null = null; let segmentSeen = false; // Used for repeated parses where our job it is to only add the new segments let prevLastSegment = last(this.segments) ?? null; const parseByteRange = (content: string) => { const atIndex = content.indexOf('@'); const length = Number(atIndex === -1 ? content : content.slice(0, atIndex)); if (!Number.isInteger(length) || length < 0) { throw new Error(`Invalid #EXT-X-BYTERANGE length '${content}'.`); } let offset: number | null = null; if (atIndex !== -1) { offset = Number(content.slice(atIndex + 1)); if (!Number.isInteger(offset) || offset < 0) { throw new Error(`Invalid #EXT-X-BYTERANGE offset '${content}'.`); } } return { length, offset }; }; const setNextSequenceNumber = (number: number) => { nextSequenceNumber = number; if (prevLastSegment) { assert(prevLastSegment.sequenceNumber !== null); if (prevLastSegment.sequenceNumber < number) { // The sequence number has finally exceeded the last sequence number we knew, meaning we can now // continue the segment list from there. Set some data to continue where we left off. accumulatedTime = prevLastSegment.timestamp + prevLastSegment.duration; currentFirstSegment = prevLastSegment.firstSegment; currentInitSegment = prevLastSegment.initSegment; lastProgramDateTimeSeconds = prevLastSegment.lastProgramDateTimeSeconds; accumulatedUnixTime = prevLastSegment.unixEpochTimestamp !== null ? prevLastSegment.unixEpochTimestamp + prevLastSegment.duration : null; prevLastSegment = null; } } }; for (let i = 0; i < lines.length; i++) { const line = lines[i]!; if (!headerRead) { if (line !== '#EXTM3U') { throw new Error('Invalid M3U8 file; expected first line to be #EXTM3U.'); } headerRead = true; continue; } if (!line.startsWith('#')) { if (!prevLastSegment) { if (nextSegmentDuration === null) { throw new Error('Invalid M3U8 file; a segment must be preceded by an #EXTINF tag.'); } let key = currentKey; if (key && key.method === 'AES-128' && !key.iv) { // "the Media Sequence Number is to be used as the IV when decrypting a Media Segment, by // putting its big-endian binary representation into a 16-octet (128-bit) buffer and padding // (on the left) with zeros" const iv = new Uint8Array(AES_128_BLOCK_SIZE); const view = toDataView(iv); view.setUint32(8, Math.floor(nextSequenceNumber / (2 ** 32))); view.setUint32(12, nextSequenceNumber); key = { ...key, iv }; } const fullPath = joinPaths(this.rootPath, line); const location: HlsSegmentLocation = { path: fullPath, offset: nextByteRange?.offset ?? 0, length: nextByteRange?.length ?? null, }; const segment: HlsSegment = { timestamp: accumulatedTime, unixEpochTimestamp: accumulatedUnixTime, firstSegment: currentFirstSegment, sequenceNumber: nextSequenceNumber, location, duration: nextSegmentDuration, encryption: key, initSegment: currentInitSegment, lastProgramDateTimeSeconds, }; currentFirstSegment ??= segment; accumulatedTime += nextSegmentDuration; if (accumulatedUnixTime !== null) { accumulatedUnixTime += nextSegmentDuration; } this.segments.push(segment); } else { // We're still seeing segments we already know about } nextSegmentDuration = null; if (nextByteRange === null) { lastByteRangeEnd = null; } else { nextByteRange = null; } setNextSequenceNumber(nextSequenceNumber + 1); } if (line.startsWith(TAG_EXTINF)) { if (prevLastSegment) { segmentSeen = true; continue; } if (!segmentSeen) { if (lastProgramDateTimeSeconds === null && nextSequenceNumber > 0 && targetDuration !== null) { // Offset the first segment's start timestamp by the following: accumulatedTime = nextSequenceNumber * targetDuration; } segmentSeen = true; } const extinfContent = line.slice(TAG_EXTINF.length); const commaIndex = extinfContent.indexOf(','); const durationStr = commaIndex === -1 ? extinfContent : extinfContent.slice(0, commaIndex); const duration = Number(durationStr); if (!Number.isFinite(duration) || duration < 0) { throw new Error(`Invalid #EXTINF tag duration '${durationStr}'.`); } nextSegmentDuration = duration; } else if (line.startsWith(TAG_MAP)) { const attributes = new AttributeList(line.slice(TAG_MAP.length)); const uri = attributes.get('uri'); if (!uri) { throw new Error('Invalid #EXT-X-MAP tag; missing URI attribute.'); } const byteRange = attributes.get('byterange'); let parsedByteRange: ReturnType | null = null; if (byteRange !== null) { parsedByteRange = parseByteRange(byteRange); } if (parsedByteRange && parsedByteRange.offset === null) { throw new Error('Invalid #EXT-X-MAP tag; BYTERANGE attribute must have a specified offset.'); } if (!prevLastSegment) { const fullPath = joinPaths(this.rootPath, uri); const location: HlsSegmentLocation = { path: fullPath, offset: parsedByteRange?.offset ?? 0, length: parsedByteRange?.length ?? null, }; if (currentKey?.method === 'AES-128' && !currentKey.iv) { // Required by the spec throw new Error('IV attribute must be set on #EXT-X-KEY tag preceding the #EXT-X-MAP tag.'); } const segment: HlsSegment = { timestamp: accumulatedTime, unixEpochTimestamp: accumulatedUnixTime, firstSegment: null, sequenceNumber: null, location, duration: 0, encryption: currentKey, initSegment: null, lastProgramDateTimeSeconds, }; // Accumulated time and sequence number are not updated in this case currentInitSegment = segment; } else { // We're still seeing segments we already know about } nextSegmentDuration = null; if (nextByteRange === null) { lastByteRangeEnd = null; } else { nextByteRange = null; } } else if (line.startsWith(TAG_KEY)) { const attributes = new AttributeList(line.slice(TAG_KEY.length)); const method = attributes.get('method'); if (method === 'NONE') { currentKey = null; } else if (method === 'AES-128') { const uri = attributes.get('uri'); if (!uri) { throw new Error('Invalid #EXT-X-KEY: AES-128 requires a URI attribute.'); } let iv: Uint8Array | null = null; const ivString = attributes.get('iv'); if (ivString) { if (!IV_STRING_REGEX.test(ivString)) { throw new Error(`Unsupported IV format '${ivString}'.`); } let hex = ivString.slice(2); hex = hex.padStart(AES_128_BLOCK_SIZE * 2, '0'); iv = new Uint8Array(AES_128_BLOCK_SIZE); for (let i = 0; i < AES_128_BLOCK_SIZE; i++) { const startIndex = -AES_128_BLOCK_SIZE * 2 + i; iv[i] = parseInt(hex.slice(startIndex, startIndex + 2), 16); } } const keyFormat = attributes.get('keyformat') ?? 'identity'; if (keyFormat !== 'identity') { throw new Error( 'For AES-128 encryption, only the \'identity\' KEYFORMAT is currently supported. If you' + ' think other formats should be supported, please raise an issue.', ); } currentKey = { method: 'AES-128', keyUri: joinPaths(this.rootPath, uri), iv, keyFormat, }; } else if (method === 'SAMPLE-AES' || method === 'SAMPLE-AES-CTR') { const uri = attributes.get('uri'); if (!uri) { throw new Error(`Invalid #EXT-X-KEY: ${method} requires a URI attribute.`); } const keyFormat = attributes.get('keyformat') ?? 'identity'; if (keyFormat === 'identity') { throw new Error( 'For SAMPLE-AES and SAMPLE-AES-CTR encryption, the \'identity\' KEYFORMAT is not' + ' supported. If you think this format should be supported, please raise an issue.', ); } let psshBox: PsshBox | null = null; if (BASE64_DATA_URI_REGEX.test(uri)) { const commaIndex = uri.indexOf(','); const bytes = base64ToBytes(uri.slice(commaIndex + 1)); if ( bytes.length >= 8 && bytes[4] === 0x70 && bytes[5] === 0x73 && bytes[6] === 0x73 && bytes[7] === 0x68 ) { const size = toDataView(bytes).getUint32(0); psshBox = parsePsshBoxContents(bytes.subarray(8, Math.min(size, bytes.length))); } } currentKey = { method, psshBox, }; } else { throw new Error( `Unsupported encryption method '${method}'. If you think this method should be supported,` + ` please raise an issue.`, ); } } else if (line.startsWith(TAG_MEDIA_SEQUENCE)) { const value = line.slice(TAG_MEDIA_SEQUENCE.length); const number = Number(value); if (!Number.isInteger(number) || number < 0) { throw new Error(`Invalid EXT-X-MEDIA-SEQUENCE value '${value}'.`); } setNextSequenceNumber(number); } else if (line.startsWith(TAG_BYTERANGE)) { const parsed = parseByteRange(line.slice(TAG_BYTERANGE.length)); if (parsed.offset === null) { if (lastByteRangeEnd === null) { throw new Error( 'Invalid M3U8 file; #EXT-X-BYTERANGE without offset requires a previous byte range.', ); } parsed.offset = lastByteRangeEnd; } nextByteRange = parsed as { length: number; offset: number }; lastByteRangeEnd = parsed.offset + parsed.length; } else if (line.startsWith(TAG_PROGRAM_DATE_TIME)) { if (prevLastSegment) { // No need to spend effort parsing dates if we're gonna discard it anyway. Also would be wrong to do // the segment shifting! continue; } const dateTime = line.slice(TAG_PROGRAM_DATE_TIME.length); const dateTimeMs = Date.parse(dateTime); if (!Number.isFinite(dateTimeMs)) { continue; } const dateTimeSeconds = dateTimeMs / 1000; if (lastProgramDateTimeSeconds === dateTimeSeconds) { continue; } if (lastProgramDateTimeSeconds === null && this.segments.length > 0) { // "If the first EXT-X-PROGRAM-DATE-TIME tag in a Playlist appears after // one or more Media Segment URIs, the client SHOULD extrapolate // backward from that tag (using EXTINF durations and/or media // timestamps) to associate dates with those segments." const lastSegment = last(this.segments)!; const lastSegmentEnd = lastSegment.timestamp + lastSegment.duration; const offset = dateTimeSeconds - lastSegmentEnd; for (const segment of this.segments) { segment.unixEpochTimestamp = segment.timestamp + offset; if (offsetTimestampsByDateTime) { segment.timestamp = segment.unixEpochTimestamp; } } } lastProgramDateTimeSeconds = dateTimeSeconds; accumulatedUnixTime = dateTimeSeconds; if (offsetTimestampsByDateTime) { accumulatedTime = dateTimeSeconds; // Snap the accumulated time into Unix space } } else if (line === TAG_DISCONTINUITY) { currentFirstSegment = null; // Note: the init segment is not reset; the #EXT-X-MAP statement simply lasts until the next // #EXT-X-MAP statement. } else if (line.startsWith(TAG_TARGETDURATION)) { const value = line.slice(TAG_TARGETDURATION.length); const duration = Number(value); if (!Number.isFinite(duration) || duration < 0) { throw new Error(`Invalid EXT-X-TARGETDURATION value '${value}'.`); } this.refreshInterval = duration; targetDuration = duration; } else if (line === TAG_ENDLIST) { this.streamHasEnded = true; break; // No need to keep reading after this } else if (line.startsWith(TAG_PLAYLIST_TYPE)) { const type = line.slice(TAG_PLAYLIST_TYPE.length); if (type.toLowerCase() === 'vod') { // A VOD playlist cannot be updated per spec so we can be sure the stream has ended this.streamHasEnded = true; } } } if (!headerRead) { throw new Error('Invalid M3U8 file; no #EXTM3U header.'); } } async getFirstSegment() { if (this.segments.length === 0) { await this.runUpdateSegments(); } return this.segments[0] ?? null; } async getSegmentAt(timestamp: number, options: SegmentRetrievalOptions) { if (this.segments.length === 0) { await this.runUpdateSegments(); } // If we're skipping the live wait BUT there's no wait time, we're actually not lazy for the first iteration let isLazy = !!options.skipLiveWait && this.getRemainingWaitTimeMs() > 0; while (true) { const index = binarySearchLessOrEqual(this.segments, timestamp, x => x.timestamp); if (index === -1) { return null; } if (index < this.segments.length - 1 || this.streamHasEnded || isLazy) { return this.segments[index]!; } const segment = this.segments[index]!; if (timestamp < segment.timestamp + segment.duration) { return segment; } await this.runUpdateSegments(); if (options.skipLiveWait) { isLazy = true; // Definitely lazy in the next iteration } } } async getNextSegment(segment: Segment, options: SegmentRetrievalOptions) { const index = this.segments.indexOf(segment as HlsSegment); assert(index !== -1); const nextIndex = index + 1; // If we're skipping the live wait BUT there's no wait time, we're actually not lazy for the first iteration let isLazy = !!options.skipLiveWait && this.getRemainingWaitTimeMs() > 0; while (true) { if (nextIndex < this.segments.length) { return this.segments[nextIndex]!; } if (this.streamHasEnded || isLazy) { return null; } await this.runUpdateSegments(); if (options.skipLiveWait) { isLazy = true; // Definitely lazy in the next iteration } } } async getPreviousSegment(segment: Segment) { const index = this.segments.indexOf(segment as HlsSegment); assert(index !== -1); return this.segments[index - 1] ?? null; } getInputForSegment(segment: Segment): Input { const hlsSegment = segment as HlsSegment; const cacheEntry = this.inputCache.find(x => x.segment === hlsSegment); if (cacheEntry) { cacheEntry.age = this.nextInputCacheAge++; return cacheEntry.input; } let initInput: Input | null = null; if (hlsSegment.initSegment || hlsSegment.firstSegment) { initInput = this.getInputForSegment((hlsSegment.initSegment ?? hlsSegment.firstSegment)!); } const formatOptions: InputFormatOptions = { ...this.input._formatOptions, isobmff: { ...this.input._formatOptions.isobmff, // Intercept calls to resolveKeyId to inject our psshBox knowledge into it resolveKeyId: this.input._formatOptions.isobmff?.resolveKeyId && ((options) => { if ( !hlsSegment.encryption || !( hlsSegment.encryption.method === 'SAMPLE-AES' || hlsSegment.encryption.method === 'SAMPLE-AES-CTR' ) || !hlsSegment.encryption.psshBox ) { return this.input._formatOptions.isobmff!.resolveKeyId!(options); } let psshBoxes = options.psshBoxes; const { psshBox } = hlsSegment.encryption; if ( (psshBox.keyIds === null || psshBox.keyIds.includes(options.keyId)) && !psshBoxes.some(x => psshBoxesAreEqual(x, psshBox)) ) { psshBoxes = [...psshBoxes, psshBox]; } return this.input._formatOptions.isobmff!.resolveKeyId!({ ...options, psshBoxes }); }), }, }; const input = new Input({ source: new CustomPathedSource( hlsSegment.location.path, async (request) => { assert(request.isRoot); // Shouldn't fail since we don't allow recursive HLS const proxiedRequest: SourceRequest = { ...request, isRoot: false, }; let ref: SourceRef; const needsSlice = hlsSegment.location.offset > 0 || hlsSegment.location.length !== null; if ( !hlsSegment.encryption || hlsSegment.encryption.method === 'SAMPLE-AES' || hlsSegment.encryption.method === 'SAMPLE-AES-CTR' ) { ref = await this.input._getSourceCached(proxiedRequest); if (needsSlice) { const slice = ref.source.slice( hlsSegment.location.offset, hlsSegment.location.length ?? undefined, ); const sliceRef = slice.ref(); ref.free(); ref = sliceRef; } } else if (hlsSegment.encryption.method === 'AES-128') { const encryption = hlsSegment.encryption; assert(encryption.iv); let ciphertextRef = await this.input._getSourceCached(proxiedRequest); if (needsSlice) { // Slice before decrypting const slice = ciphertextRef.source.slice( hlsSegment.location.offset, hlsSegment.location.length ?? undefined, ); const sliceRef = slice.ref(); ciphertextRef.free(); ciphertextRef = sliceRef; } const ciphertextReader = new Reader(ciphertextRef.source); const stream = createAes128CbcDecryptStream(ciphertextReader, async () => { using keyRef = await this.input._getSourceCached( { path: encryption.keyUri, isRoot: false }, ENCRYPTION_KEY_CACHE_GROUP, ); const keyReader = new Reader(keyRef.source); const keySlice = await keyReader.requestSlice(0, AES_128_BLOCK_SIZE); if (!keySlice) { throw new Error('Invalid AES-128 key; expected at least 16 bytes of data.'); } const key = readBytes(keySlice, AES_128_BLOCK_SIZE); return { key, iv: encryption.iv! }; }, () => { ciphertextRef.free(); }); ref = new ReadableStreamSource(stream).ref(); } else { assert(false); } return ref; }, ), // Do not allow recursive HLS. Cool on paper, but allows for nasty infinite-depth request trees. formats: this.input._formats.filter(x => !(x instanceof HlsInputFormat)), initInput: initInput ?? undefined, formatOptions, }); input._onFormatDetermined = (format) => { if ( (hlsSegment.encryption?.method === 'SAMPLE-AES' || hlsSegment.encryption?.method === 'SAMPLE-AES-CTR') && !format._isIsobmff ) { // These methods can also be used for formats such as MPEG-TS // eslint-disable-next-line @stylistic/max-len // (see https://developer.apple.com/library/archive/documentation/AudioVideo/Conceptual/HLS_Sample_Encryption/Encryption/Encryption.html) // but we don't support them there yet, so instead of silently decrypting nothing, we throw an error. throw new Error( 'The SAMPLE-AES and SAMPLE-AES-CTR encryption methods are currently only supported for' + ' ISOBMFF files.', ); } }; this.inputCache.push({ segment: hlsSegment, input, age: this.nextInputCacheAge++, }); const MAX_INPUT_CACHE_SIZE = 4; if (this.inputCache.length > MAX_INPUT_CACHE_SIZE) { const minAgeIndex = arrayArgmin(this.inputCache, x => x.age); assert(minAgeIndex !== -1); this.inputCache.splice(minAgeIndex, 1); // DON'T dispose here; the Input might still be used! The source disposal will happen with GC logic } return input; } async getLiveRefreshInterval() { if (this.getRemainingWaitTimeMs() === 0) { await this.runUpdateSegments(); } return this.streamHasEnded ? null : this.refreshInterval; } } ===== src/hls/hls-muxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { MediaCodec, validateAudioChunkMetadata, validateVideoChunkMetadata } from '../codec'; import { Logging } from '../logging'; import { EncodedAudioPacketSource, EncodedVideoPacketSource } from '../media-source'; import { arrayArgmax, assert, AsyncMutex, findLastIndex, joinPaths, textEncoder, toArray, UNDETERMINED_LANGUAGE, } from '../misc'; import { Muxer } from '../muxer'; import { Output, OutputAudioTrack, OutputSubtitleTrack, OutputTrack, OutputVideoTrack, TrackType, } from '../output'; import { HlsOutputFormat, HlsOutputFormatOptions, HlsOutputPlaylistInfo, HlsOutputSegmentInfo, OutputFormat, } from '../output-format'; import { Writer } from '../writer'; import { EncodedPacket } from '../packet'; import { SubtitleCue, SubtitleMetadata } from '../subtitles'; import { NullTarget, PathedTarget, Target, TargetRequest } from '../target'; import { HLS_MIME_TYPE } from './hls-misc'; type HlsTrackData = { track: OutputTrack; packets: EncodedPacket[]; playlist: Playlist; // We must store it on the TrackData, reading it directly from the track leads to async race conditions! closed: boolean; info: { type: 'video'; decoderConfig: VideoDecoderConfig; } | { type: 'audio'; decoderConfig: AudioDecoderConfig; }; }; type HlsVideoTrackData = HlsTrackData & { info: { type: 'video' } }; type HlsAudioTrackData = HlsTrackData & { info: { type: 'audio' } }; type PlaylistSegment = { path: string; duration: number; timestamp: number; byteSize: number; byteOffset: number | null; info: HlsOutputSegmentInfo | null; }; type Playlist = { id: number; path: string; tracks: OutputTrack[]; segmentFormat: OutputFormat; currentSegmentStartTimestamp: number | null; currentSegmentStartTimestampIsFixed: boolean; nextSegmentId: number; initSegment: PlaylistSegment | null; writtenSegments: PlaylistSegment[]; peakBitrate: number | null; averageBitrate: number | null; mediaSequence: number; done: boolean; singleFile: { target: Target; path: string; nextOffset: number; info: HlsOutputSegmentInfo; } | null; // For HLS, having a single mutex is too coarse. Every playlist is basically independent and therefore we can have // a per-playlist mutex instead of a per-muxer one. This means two packets from different playlists coming in don't // block each other. mutex: AsyncMutex; }; type PlaylistDeclaration = { playlist: Playlist; groupId: string | null; noUri: boolean; references: PlaylistDeclaration[]; }; export class HlsMuxer extends Muxer { format: HlsOutputFormat; getPlaylistPath: NonNullable; getSegmentPath: NonNullable; getInitPath: NonNullable; targetSegmentDuration: number; trackDatas: HlsTrackData[] = []; singleFilePerPlaylist: boolean; isLive: boolean; maxLiveSegmentCount: number; isRelativeToUnixEpoch = false; globalTargetDuration: number; numWrittenMasterPlaylists = 0; playlists: Playlist[] = []; playlistDeclarations: PlaylistDeclaration[] = []; constructor(output: Output, format: HlsOutputFormat) { if (!(output._target instanceof PathedTarget)) { throw new TypeError('HLS outputs require `OutputOptions.target` to be a PathedTarget.'); } super(output); this.format = format; this.targetSegmentDuration = format._options.targetDuration ?? 2; this.singleFilePerPlaylist = format._options.singleFilePerPlaylist ?? false; this.isLive = format._options.live ?? false; this.maxLiveSegmentCount = format._options.maxLiveSegmentCount ?? Infinity; this.globalTargetDuration = this.targetSegmentDuration; this.getPlaylistPath = format._options.getPlaylistPath ?? (({ n }) => `playlist-${n}.m3u8`); this.getSegmentPath = format._options.getSegmentPath ?? (info => info.isSingleFile ? `segments-${info.playlist.n}${info.format.fileExtension}` : `segment-${info.playlist.n}-${info.n}${info.format.fileExtension}`); this.getInitPath = format._options.getInitPath ?? (playlist => `init-${playlist.n}${playlist.segmentFormat.fileExtension}`); } async start(): Promise { const release = await this.mutex.acquire(); const someRelative = this.output._tracks.some(t => t.metadata.isRelativeToUnixEpoch); const someNotRelative = this.output._tracks.some(t => !t.metadata.isRelativeToUnixEpoch); if (someRelative && someNotRelative) { throw new Error( 'All tracks must agree on `relativeToUnixEpoch`: some tracks are relative to the Unix epoch and some' + ' are not.', ); } this.isRelativeToUnixEpoch = someRelative; // Upon starting, we now need to assign the tracks to separate playlists. This assignment will make use of the // track pairability information provided by the user as well as other metadata specified on the tracks. The // resulting master playlist should preserve track pairability; meaning that all tracks that are pairable // remain pairable, and no two tracks become pairable that are meant to be mutually exclusive. // The algorithm determines "groups" by enumerating all pairable tracks for each track, and then materializes // each group either as #EXT-X-MEDIA tags or top-level #EXT-X-STREAM-INF tags. The algorithm is biased towards // video being the top-level grouping, since that's the standard practice. const groupAssignment = new Map(); const groups: { name: string; key: string; tracks: OutputTrack[]; needsEmit: boolean; firstNoUri: boolean; }[] = []; let hasVideo = false; let illegalPairingDetected = false; let keyPacketsOnlyPairingWarned = false; // First, let's build the "sibling" groups induced by track pairability for (const track of this.output._tracks) { if (track.type === 'video') { hasVideo = true; } const pairableGroups = new Map(); for (const otherTrack of this.output._tracks) { if (track === otherTrack) { continue; } if (!track.canBePairedWith(otherTrack)) { continue; } if (track.type === otherTrack.type) { if (!illegalPairingDetected) { Logging._warn( `Illegal pairing of two ${track.type} tracks detected, which is not possible in HLS;` + ` treating them as unpaired.`, ); illegalPairingDetected = true; } continue; } // Key-packets-only tracks can neither pair with nor be paired with other tracks if ( (track.isVideoTrack() && track.metadata.hasOnlyKeyPackets) || (otherTrack.isVideoTrack() && otherTrack.metadata.hasOnlyKeyPackets) ) { if (!keyPacketsOnlyPairingWarned) { Logging._warn( `A key-packets-only video track is pairable with another track, which is not` + ` possible in HLS; treating them as unpaired.`, ); keyPacketsOnlyPairingWarned = true; } continue; } let groupTracks = pairableGroups.get(otherTrack.source._codec); if (!groupTracks) { pairableGroups.set(otherTrack.source._codec, groupTracks = []); } groupTracks.push(otherTrack); } for (const [, pairableTracks] of pairableGroups) { const key = pairableTracks.map(x => x.id).join('-'); const group = groups.find(x => x.key === key); if (!group) { groups.push({ name: pairableTracks[0]!.type + '-' + (groups.length + 1), key, tracks: pairableTracks, needsEmit: false, firstNoUri: false, }); } let assignedGroups = groupAssignment.get(track); if (!assignedGroups) { groupAssignment.set(track, assignedGroups = []); } assignedGroups.push(key); } } const mainType: TrackType = hasVideo ? 'video' : 'audio'; const variantStreams: { tracks: OutputTrack[]; linkedGroup: typeof groups[number] | null; }[] = []; const unpairedVideoTracks: OutputTrack[] = []; const unpairedAudioTracks: OutputTrack[] = []; // Now, create the top-level variant streams for (const track of this.output._tracks) { const assignedGroupKeys = groupAssignment.get(track); if (assignedGroupKeys) { assert(assignedGroupKeys.length > 0); if (track.type !== mainType) { continue; } for (const key of assignedGroupKeys) { const group = groups.find(x => x.key === key); assert(group); if (assignedGroupKeys.length === 1 && group.tracks.length === 1) { const otherGroupKeys = groupAssignment.get(group.tracks[0]!); assert(otherGroupKeys !== undefined); if (otherGroupKeys.length === 1) { const otherGroup = groups.find(x => x.key === otherGroupKeys[0]!)!; if (otherGroup.tracks.length === 1) { assert(otherGroup.tracks[0] === track); variantStreams.push({ tracks: [track, group.tracks[0]!], linkedGroup: null, }); continue; } } } variantStreams.push({ tracks: [track], linkedGroup: group, }); group.needsEmit = true; } } else { if (track.type === 'video') { unpairedVideoTracks.push(track); } else if (track.type === 'audio') { unpairedAudioTracks.push(track); } } } const getMetadataKeyForTrack = ({ metadata }: OutputTrack) => { let key = ''; key += `${metadata.languageCode ?? UNDETERMINED_LANGUAGE}-`; key += `${metadata.name ?? ''}-`; key += `${metadata.disposition?.default ?? true}-`; key += `${metadata.disposition?.primary ?? false}-`; key += `${metadata.disposition?.forced ?? false}-`; return key; }; // Video tracks that can't be paired with any other track always live on the top-level, the question is just if // they need to be separated into #EXT-X-MEDIA tags or not if (unpairedVideoTracks.length > 0) { const uniqueMetadata = new Set(unpairedVideoTracks.map(getMetadataKeyForTrack)); if (uniqueMetadata.size > 1) { // They differ in metadata, emit as group const group: typeof groups[number] = { key: unpairedVideoTracks.map(x => x.id).join('-'), name: 'video-' + (groups.length + 1), tracks: unpairedVideoTracks, needsEmit: true, firstNoUri: true, }; groups.push(group); variantStreams.push({ tracks: [unpairedVideoTracks[0]!], linkedGroup: group, }); } else { for (const track of unpairedVideoTracks) { variantStreams.push({ tracks: [track], linkedGroup: null, }); } } } // Audio tracks that can't be paired with any other track always live on the top-level, the question is just if // they need to be separated into #EXT-X-MEDIA tags or not if (unpairedAudioTracks.length > 0) { const uniqueMetadata = new Set(unpairedAudioTracks.map(getMetadataKeyForTrack)); if (uniqueMetadata.size > 1) { // They differ in metadata, emit as group const group: typeof groups[number] = { key: unpairedAudioTracks.map(x => x.id).join('-'), name: 'audio-' + (groups.length + 1), tracks: unpairedAudioTracks, needsEmit: true, firstNoUri: true, }; groups.push(group); variantStreams.push({ tracks: [unpairedAudioTracks[0]!], linkedGroup: group, }); } else { for (const track of unpairedAudioTracks) { variantStreams.push({ tracks: [track], linkedGroup: null, }); } } } const deduceSegmentFormat = (tracks: OutputTrack[]) => { const codecs: MediaCodec[] = []; let videoCount = 0; let audioCount = 0; let requiresRotationMetadata = false; let candidate: OutputFormat | null = null; let candidateScore = -Infinity; for (const track of tracks) { if (track.isVideoTrack()) { videoCount++; requiresRotationMetadata ||= (track.metadata.rotation ?? 0) !== 0; } else if (track.isAudioTrack()) { audioCount++; } codecs.push(track.source._codec); } for (const format of toArray(this.format._options.segmentFormat)) { const supportedCodecs = format.getSupportedCodecs(); const trackCounts = format.getSupportedTrackCounts(); if (codecs.some(codec => !supportedCodecs.includes(codec))) { continue; } if (videoCount < trackCounts.video.min || videoCount > trackCounts.video.max) { continue; } if (audioCount < trackCounts.audio.min || audioCount > trackCounts.audio.max) { continue; } let score = 0; if (requiresRotationMetadata && format.supportsVideoRotationMetadata) { score++; } if (score > candidateScore) { candidate = format; candidateScore = score; } } // We must find a format. If no format is found, that means we incorrectly gated track creation and // assignment at an earlier step. assert(candidate); return candidate; }; const registerPlaylist = async (tracks: OutputTrack[]) => { if (tracks.some(track => this.playlists.some(playlist => playlist.tracks.includes(track)))) { throw new Error('Internal error: track is already registered in a playlist.'); // Should be unreachable } const format = deduceSegmentFormat(tracks); const id = this.playlists.length + 1; const path = await this.getPlaylistPath({ n: id, tracks, segmentFormat: format, }); validatePlaylistPath(path); const playlist: Playlist = { id: this.playlists.length + 1, path, tracks, segmentFormat: format, currentSegmentStartTimestamp: null, currentSegmentStartTimestampIsFixed: false, nextSegmentId: 1, initSegment: null, writtenSegments: [], peakBitrate: null, averageBitrate: null, mediaSequence: 0, done: false, singleFile: null, mutex: new AsyncMutex(), }; this.playlists.push(playlist); return playlist; }; // Now, finally let's create all declarations. Each declaration maps to one #EXT-X-MEDIA or #EXT-X-STREAM-INF // tag in the final master playlist. for (const group of groups) { if (!group.needsEmit) { continue; } for (let i = 0; i < group.tracks.length; i++) { const track = group.tracks[i]!; let playlist = this.playlists.find(x => x.tracks[0]!.id === track.id); playlist ??= await registerPlaylist([track]); this.playlistDeclarations.push({ playlist, groupId: group.name, noUri: group.firstNoUri && i === 0, references: [], }); } } for (const variant of variantStreams) { // Since tracks can only be assigned to one playlist, the first track's ID acts as a "playlist key" let playlist = this.playlists.find(x => x.tracks[0]!.id === variant.tracks[0]!.id); playlist ??= await registerPlaylist(variant.tracks); this.playlistDeclarations.push({ playlist, groupId: null, noUri: false, references: variant.linkedGroup ? this.playlistDeclarations.filter(x => x.groupId === variant.linkedGroup!.name) : [], }); } release(); } async getMimeType(): Promise { return HLS_MIME_TYPE; } private allTracksAreKnown(playlist: Playlist) { for (const track of playlist.tracks) { if (!track.source._closed && !this.trackDatas.some(x => x.track === track)) { return false; // We haven't seen a sample from this open track yet } } return true; } // eslint-disable-next-line @typescript-eslint/no-misused-promises override async onTrackClose(track: OutputTrack) { const trackData = this.trackDatas.find(x => x.track === track); if (trackData) { trackData.closed = true; } const playlist = this.playlists.find(x => x.tracks.includes(track)); assert(playlist); // If there isn't one then the assignment algo failed innit const release = await playlist.mutex.acquire(); try { await this.advancePlaylist(playlist); } finally { release(); } } getVideoTrackData(track: OutputVideoTrack, meta?: EncodedVideoChunkMetadata) { let trackData = this.trackDatas.find(x => x.track === track) as HlsVideoTrackData; if (trackData) { return trackData; } validateVideoChunkMetadata(meta); assert(meta); assert(meta?.decoderConfig); const playlists = this.playlists.filter(x => x.tracks.includes(track)); assert(playlists.length === 1); trackData = { track, packets: [], playlist: playlists[0]!, closed: false, info: { type: 'video', decoderConfig: meta.decoderConfig, }, }; this.trackDatas.push(trackData); return trackData; } getAudioTrackData(track: OutputAudioTrack, meta?: EncodedAudioChunkMetadata) { let trackData = this.trackDatas.find(x => x.track === track) as HlsAudioTrackData; if (trackData) { return trackData; } validateAudioChunkMetadata(meta); assert(meta); assert(meta?.decoderConfig); const playlists = this.playlists.filter(x => x.tracks.includes(track)); assert(playlists.length === 1); trackData = { track, packets: [], playlist: playlists[0]!, closed: false, info: { type: 'audio', decoderConfig: meta.decoderConfig, }, }; this.trackDatas.push(trackData); return trackData; } async addEncodedVideoPacket( track: OutputVideoTrack, packet: EncodedPacket, meta?: EncodedVideoChunkMetadata, ) { const trackData = this.getVideoTrackData(track, meta); const playlist = trackData.playlist; const release = await playlist.mutex.acquire(); try { this.validateTimestamp(track, packet.timestamp, packet.type === 'key'); trackData.packets.push(packet); if (playlist.currentSegmentStartTimestamp === null) { playlist.currentSegmentStartTimestamp = packet.timestamp; } else if (!playlist.currentSegmentStartTimestampIsFixed) { playlist.currentSegmentStartTimestamp = Math.min( playlist.currentSegmentStartTimestamp, packet.timestamp, ); } await this.advancePlaylist(playlist); } finally { release(); } } async addEncodedAudioPacket( track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata, ) { const trackData = this.getAudioTrackData(track, meta); const playlist = trackData.playlist; const release = await playlist.mutex.acquire(); try { this.validateTimestamp(track, packet.timestamp, packet.type === 'key'); trackData.packets.push(packet); if (playlist.currentSegmentStartTimestamp === null) { playlist.currentSegmentStartTimestamp = packet.timestamp; } else if (!playlist.currentSegmentStartTimestampIsFixed) { playlist.currentSegmentStartTimestamp = Math.min( playlist.currentSegmentStartTimestamp, packet.timestamp, ); } await this.advancePlaylist(playlist); } finally { release(); } } async addSubtitleCue( // eslint-disable-next-line @typescript-eslint/no-unused-vars track: OutputSubtitleTrack, // eslint-disable-next-line @typescript-eslint/no-unused-vars cue: SubtitleCue, // eslint-disable-next-line @typescript-eslint/no-unused-vars meta?: SubtitleMetadata, ) { throw new Error('Unreachable.'); } async advancePlaylist(playlist: Playlist) { assert(!playlist.done); if (!this.allTracksAreKnown(playlist)) { return; } if (playlist.currentSegmentStartTimestamp === null) { // All tracks are known but we never received any data - all tracks must be closed already await this.onPlaylistDone(playlist); return; } const trackDatas = this.trackDatas.filter(x => playlist.tracks.includes(x.track)); const videoTrack = trackDatas.find(x => x.info.type === 'video') as HlsVideoTrackData | undefined; const audioTrack = trackDatas.find(x => x.info.type === 'audio') as HlsAudioTrackData | undefined; // Loop in case we can finalize multiple segments while (true) { // This here is the core segmentation logic. The segmentation logic figures out which packets are to be // written into the next segment, and if we can write a segment at all. If tracks are still open and have // not provided sufficient media data, no segment will be written. The packets will be added to the segment // to maximize its duration AND keep it from exceeding the target duration. This condition is extended with // a key frame rule for video, meaning the algorithm must guarantee that every segment with video data // begins with a video key frame. // // The logic is quite complex but is solved in a straight-forward way: all possible permutations of the // problem are checked in a nested if-else structure, making sure all cases behave correctly. This was the // easiest, least error-prone way I found to express this behavior. const currentSegmentEndTimestamp = playlist.currentSegmentStartTimestamp + this.targetSegmentDuration; // These store the index (exclusive) until when packets can be added to the next segment let videoEndIndex = 0; let audioEndIndex = 0; if (videoTrack && (!videoTrack.closed || videoTrack.packets.length > 0)) { // A video track is active (and maybe an audio track too) const allBelow = videoTrack.packets.every(x => x.timestamp < currentSegmentEndTimestamp); let bestKeyPacket: EncodedPacket | null = null; let bestKeyPacketIndex: number | null = null; if (allBelow) { if (!videoTrack.closed) { // Not enough data yet return; } } else { // Find the best key packet timestamp for (let i = 0; i < videoTrack.packets.length; i++) { const packet = videoTrack.packets[i]!; if (bestKeyPacket !== null && packet.timestamp > currentSegmentEndTimestamp) { break; } if (i > 0 && packet.type === 'key') { bestKeyPacket = packet; bestKeyPacketIndex = i; } } } if (bestKeyPacketIndex !== null) { videoEndIndex = bestKeyPacketIndex; if (audioTrack) { // The audio track must go at least until the video key frame const index = audioTrack.packets.findIndex(x => x.timestamp >= bestKeyPacket!.timestamp); if (index !== -1) { audioEndIndex = index; } else { if (audioTrack.closed) { audioEndIndex = audioTrack.packets.length; } else { return; } } } } else { if (!videoTrack.closed) { return; } // Include the entire rest of the video (since there's no key frame to split it on) videoEndIndex = videoTrack.packets.length; const maxIndex = arrayArgmax(videoTrack.packets, x => x.timestamp); const maxPacket = videoTrack.packets[maxIndex]; assert(maxPacket); if (audioTrack) { if (maxPacket.timestamp < currentSegmentEndTimestamp) { // The audio must go until at least the start of the next segment const index = audioTrack.packets.findIndex(x => x.timestamp >= currentSegmentEndTimestamp); if (index !== -1) { audioEndIndex = index; } else { if (audioTrack.closed) { audioEndIndex = audioTrack.packets.length; } else { return; } } } else { // The audio must go beyond the last video packet const index = audioTrack.packets.findIndex(x => x.timestamp > maxPacket.timestamp); if (index !== -1) { audioEndIndex = index; } else { if (audioTrack.closed) { audioEndIndex = audioTrack.packets.length; } else { return; } } } } } } else if (audioTrack && (!audioTrack.closed || audioTrack.packets.length > 0)) { // There's only an audio track active const allBelow = audioTrack.packets.every(x => x.timestamp < currentSegmentEndTimestamp); if (allBelow) { if (audioTrack.closed) { // We can write all packets since they're all below audioEndIndex = audioTrack.packets.length; } else { // We don't know enough packets yet return; } } else { // Aim to make the segment at most as long as desired const index = findLastIndex(audioTrack.packets, x => x.timestamp <= currentSegmentEndTimestamp); audioEndIndex = Math.max(index, 1); // Always include at least the first packet } } if (videoEndIndex === 0 && audioEndIndex === 0) { // No more segments to write - if all tracks are closed, this playlist is done const allClosed = trackDatas.every(x => x.closed); if (allClosed) { await this.onPlaylistDone(playlist); } return; } // We can finalize a new segment! let segmentInfo: HlsOutputSegmentInfo | null = null; let relativeSegmentPath: string; let fullSegmentPath: string; assert(this.output._target instanceof PathedTarget); const pathedTarget = this.output._target; if (this.singleFilePerPlaylist) { if (playlist.singleFile === null) { // INTENTIONALLY shadow the outside `segmentInfo` because we don't want to set it. // In single-file mode, onSegment is called once in onPlaylistDone instead of per-segment, // so the outer `segmentInfo` intentionally stays null in this case. const segmentInfo: HlsOutputSegmentInfo = { n: playlist.nextSegmentId, format: playlist.segmentFormat, isSingleFile: true, playlist: toPlaylistInfo(playlist), }; relativeSegmentPath = await this.getSegmentPath(segmentInfo); validateSegmentPath(relativeSegmentPath); fullSegmentPath = joinPaths( joinPaths(pathedTarget.rootPath, playlist.path), relativeSegmentPath, ); const target = await this.output._getTarget({ path: fullSegmentPath, isRoot: false, mimeType: playlist.segmentFormat.mimeType, }); target._start(); playlist.singleFile = { target, path: relativeSegmentPath, nextOffset: 0, info: segmentInfo, }; } else { relativeSegmentPath = playlist.singleFile.path; fullSegmentPath = joinPaths( joinPaths(pathedTarget.rootPath, playlist.path), relativeSegmentPath, ); } } else { segmentInfo = { n: playlist.nextSegmentId, format: playlist.segmentFormat, isSingleFile: false, playlist: toPlaylistInfo(playlist), }; relativeSegmentPath = await this.getSegmentPath(segmentInfo); validateSegmentPath(relativeSegmentPath); fullSegmentPath = joinPaths(joinPaths(pathedTarget.rootPath, playlist.path), relativeSegmentPath); playlist.nextSegmentId++; } let segmentSize = 0; let outputTarget: Target | null = null; const output = new Output({ format: playlist.segmentFormat, target: new PathedTarget( fullSegmentPath, async (request: TargetRequest) => { const proxiedRequest: TargetRequest = { ...request, isRoot: false, }; if (request.isRoot) { if (playlist.singleFile) { const slice = playlist.singleFile.target.slice(playlist.singleFile.nextOffset); slice.on('write', ({ end }) => segmentSize = Math.max(segmentSize, end)); return slice; } else { const target = await this.output._getTarget(proxiedRequest); outputTarget = target; target.on('write', ({ end }) => segmentSize = Math.max(segmentSize, end)); return target; } } return this.output._getTarget(proxiedRequest); }, ), initTarget: async () => { if (playlist.initSegment) { // We already have an init segment from a previous segment return new NullTarget(); } if (playlist.singleFile) { playlist.initSegment = { path: playlist.singleFile.path, duration: 0, timestamp: 0, byteSize: 0, byteOffset: 0, info: null, }; const slice = playlist.singleFile.target.slice(playlist.singleFile.nextOffset); slice.on('write', ({ end }) => { playlist.initSegment!.byteSize = Math.max(playlist.initSegment!.byteSize, end); }); slice.on('finalized', () => { playlist.singleFile!.nextOffset = playlist.initSegment!.byteSize; }); return slice; } else { const playlistInfo = toPlaylistInfo(playlist); const initPath = await this.getInitPath(playlistInfo); validateInitPath(initPath); playlist.initSegment = { path: initPath, duration: 0, timestamp: 0, byteSize: 0, byteOffset: null, info: null, }; const fullInitPath = joinPaths( joinPaths(pathedTarget.rootPath, playlist.path), initPath, ); const target = await this.output._getTarget({ path: fullInitPath, isRoot: false, mimeType: playlist.segmentFormat.mimeType, }); target.on('write', ({ end }) => { playlist.initSegment!.byteSize = Math.max(playlist.initSegment!.byteSize, end); }); target.on('finalized', () => { this.format._options.onInit?.(target, playlistInfo); }); return target; } }, }); let maxEndTimestamp = -Infinity; try { let videoSource: EncodedVideoPacketSource | null = null; let audioSource: EncodedAudioPacketSource | null = null; if (videoTrack) { // Always add the track, no matter if it has packets or not (maintains underlying IDs) videoSource = new EncodedVideoPacketSource((videoTrack.track as OutputVideoTrack).source._codec); output.addVideoTrack(videoSource, videoTrack.track.metadata); } if (audioTrack) { // Always add the track, no matter if it has packets or not (maintains underlying IDs) audioSource = new EncodedAudioPacketSource((audioTrack.track as OutputAudioTrack).source._codec); output.addAudioTrack(audioSource, audioTrack.track.metadata); } await output.start(); // Add all of the packets if (videoTrack) { assert(videoSource); const meta = { decoderConfig: videoTrack.info.decoderConfig }; for (let i = 0; i < videoEndIndex; i++) { const packet = videoTrack.packets[i]!; await videoSource.add(packet, meta); maxEndTimestamp = Math.max(maxEndTimestamp, packet.timestamp + packet.duration); } } if (audioTrack) { assert(audioSource); const meta = { decoderConfig: audioTrack.info.decoderConfig }; for (let i = 0; i < audioEndIndex; i++) { const packet = audioTrack.packets[i]!; await audioSource.add(packet, meta); maxEndTimestamp = Math.max(maxEndTimestamp, packet.timestamp + packet.duration); } } await output.finalize(); } catch (e) { await output.cancel(); throw e; } if (segmentInfo) { assert(outputTarget); this.format._options.onSegment?.(outputTarget, segmentInfo); } if (videoEndIndex > 0) { assert(videoTrack); videoTrack.packets.splice(0, videoEndIndex); } if (audioEndIndex > 0) { assert(audioTrack); audioTrack.packets.splice(0, audioEndIndex); } let minNextTimestamp = Infinity; if (videoTrack && videoTrack.packets.length > 0) { minNextTimestamp = videoTrack.packets[0]!.timestamp; } if (audioTrack && audioTrack.packets.length > 0) { minNextTimestamp = Math.min(minNextTimestamp, audioTrack.packets[0]!.timestamp); } const nextSegmentStartTimestamp = minNextTimestamp < Infinity ? minNextTimestamp : maxEndTimestamp; // Happens for the last segment for example assert(Number.isFinite(nextSegmentStartTimestamp)); const segmentDuration = nextSegmentStartTimestamp - playlist.currentSegmentStartTimestamp; assert(segmentDuration >= 0); playlist.writtenSegments.push({ path: relativeSegmentPath, duration: segmentDuration, timestamp: playlist.currentSegmentStartTimestamp, byteSize: segmentSize, byteOffset: playlist.singleFile ? playlist.singleFile.nextOffset : null, info: segmentInfo ?? null, }); this.globalTargetDuration = Math.max(this.globalTargetDuration, segmentDuration); playlist.currentSegmentStartTimestamp = nextSegmentStartTimestamp; playlist.currentSegmentStartTimestampIsFixed = true; // After the first segment, the timestamp is now fixed if (playlist.singleFile) { playlist.singleFile.nextOffset += segmentSize; } if (this.isLive) { while (playlist.writtenSegments.length > this.maxLiveSegmentCount) { const popped = playlist.writtenSegments.shift()!; playlist.mediaSequence++; if (!this.singleFilePerPlaylist) { assert(popped.info); this.format._options.onSegmentPopped?.(popped.path, popped.info); } } await this.writePlaylist(playlist); await this.tryWriteMasterPlaylist(); } } } private async onPlaylistDone(playlist: Playlist) { assert(!playlist.done); playlist.done = true; if (playlist.singleFile) { await playlist.singleFile.target._flush(); await playlist.singleFile.target._finalize(); this.format._options.onSegment?.(playlist.singleFile.target, playlist.singleFile.info); } await this.writePlaylist(playlist); if (this.isLive && playlist.writtenSegments.length === 0) { await this.tryWriteMasterPlaylist(); } } private updatePlaylistBitrates(playlist: Playlist) { const segments = playlist.writtenSegments; let peakBitrate = 0; let totalBits = 0; let totalDuration = 0; // Per spec, peak bitrate is the largest bit rate of any contiguous set of segments whose total duration is // between 0.5 and 1.5 times the target duration for (let i = 0; i < segments.length; i++) { totalDuration += segments[i]!.duration; let windowBytes = 0; let windowDuration = 0; for (let j = i; j < segments.length; j++) { windowBytes += segments[j]!.byteSize; windowDuration += segments[j]!.duration; if ( windowDuration >= 0.5 * this.globalTargetDuration && windowDuration <= 1.5 * this.globalTargetDuration ) { peakBitrate = Math.max(peakBitrate, 8 * windowBytes / windowDuration); } if (windowDuration > 1.5 * this.globalTargetDuration) { break; } } } // Fallback: if no contiguous set falls within the range, use per-segment max if (peakBitrate === 0) { for (const segment of segments) { const segmentDuration = segment.duration || 1; // To catch 0-duration segments which can happen peakBitrate = Math.max(peakBitrate, 8 * segment.byteSize / segmentDuration); } } for (const segment of segments) { totalBits += 8 * segment.byteSize; } playlist.peakBitrate = peakBitrate; playlist.averageBitrate = totalBits / (totalDuration || 1); } private async writePlaylist(playlist: Playlist) { assert(this.output._target instanceof PathedTarget); const pathedTarget = this.output._target; this.updatePlaylistBitrates(playlist); let hasByteOffsets = false; for (const segment of playlist.writtenSegments) { hasByteOffsets ||= segment.byteOffset !== null; } const isKeyPacketsOnly = playlist.tracks[0]!.isVideoTrack() && playlist.tracks[0].metadata.hasOnlyKeyPackets; let version = 3; if (isKeyPacketsOnly || hasByteOffsets) { version = 4; } if (playlist.initSegment) { version = 5; } if (playlist.initSegment && !isKeyPacketsOnly) { // "if it contains the EXT-X-MAP tag in a Media Playlist that does not contain EXT-X-I-FRAMES-ONLY" version = 6; } // In live mode, target duration is not allowed to change, so we use the nominal value const targetDuration = this.isLive ? this.targetSegmentDuration : this.globalTargetDuration; const playlistPath = joinPaths(pathedTarget.rootPath, playlist.path); const playlistText = '#EXTM3U\n' + `#EXT-X-VERSION:${version}\n` + (!this.isLive ? '#EXT-X-PLAYLIST-TYPE:VOD\n' : '') + `#EXT-X-TARGETDURATION:${Math.ceil(targetDuration)}\n` // Must be a "decimal-integer" + (Number.isFinite(this.maxLiveSegmentCount) ? `#EXT-X-MEDIA-SEQUENCE:${playlist.mediaSequence}\n` : '') + '#EXT-X-INDEPENDENT-SEGMENTS\n' + (isKeyPacketsOnly ? '#EXT-X-I-FRAMES-ONLY\n' : '') + (playlist.initSegment ? (`#EXT-X-MAP:URI="${playlist.initSegment.path}"` + (playlist.initSegment.byteOffset !== null ? `,BYTERANGE="${playlist.initSegment.byteSize}@${playlist.initSegment.byteOffset}"` : '') + '\n') : '') + '\n' + (playlist.writtenSegments .map(segment => ( `#EXTINF:${+segment.duration.toFixed(12)},\n` // Trailing comma mandated by spec + (this.isRelativeToUnixEpoch ? `#EXT-X-PROGRAM-DATE-TIME:${new Date(1000 * segment.timestamp).toISOString()}\n` : '') + (segment.byteOffset !== null ? `#EXT-X-BYTERANGE:${segment.byteSize}@${segment.byteOffset}\n` : '') + `${segment.path}\n` )) .join('')) + (playlist.done ? (playlist.writtenSegments.length > 0 ? '\n' : '') + '#EXT-X-ENDLIST\n' : ''); this.format._options.onPlaylist?.(playlistText, toPlaylistInfo(playlist)); const target = await this.output._getTarget({ path: playlistPath, isRoot: false, mimeType: HLS_MIME_TYPE, }); const writer = new Writer(target, true); writer.start(); writer.write(textEncoder.encode(playlistText)); await writer.flush(); await writer.finalize(); } private async writeMasterPlaylist() { assert(this.output._target instanceof PathedTarget); const pathedTarget = this.output._target; let masterPlaylistText = '#EXTM3U\n'; let firstVariantWritten = false; let lastGroupId: string | null = null; let groupIdTrackCount = 0; let hasHadDefaultTrackInGroup = false; for (const decl of this.playlistDeclarations) { if (decl.groupId === null) { const isKeyPacketsOnly = decl.playlist.tracks[0]!.isVideoTrack() && decl.playlist.tracks[0].metadata.hasOnlyKeyPackets; const codecs: string[] = []; for (const track of decl.playlist.tracks) { const trackData = this.trackDatas.find(x => x.track === track); const codecString = trackData?.info.decoderConfig.codec ?? track.source._codec; codecs.push(codecString); } let peakDeclBitrate = 0; let maxRefAverageBitrate = 0; if (decl.references.length > 0) { const firstRef = decl.references[0]!; const firstTrack = firstRef.playlist.tracks[0]!; const trackData = this.trackDatas.find(x => x.track === firstTrack); const codecString = trackData?.info.decoderConfig.codec ?? firstTrack.source._codec; codecs.push(codecString); for (const ref of decl.references) { assert(ref.playlist.peakBitrate !== null); peakDeclBitrate = Math.max(peakDeclBitrate, ref.playlist.peakBitrate); maxRefAverageBitrate = Math.max(maxRefAverageBitrate, ref.playlist.averageBitrate ?? 0); } } assert(decl.playlist.peakBitrate !== null); const totalPeakBitrate = decl.playlist.peakBitrate + peakDeclBitrate; const totalAverageBitrate = (decl.playlist.averageBitrate ?? 0) + maxRefAverageBitrate; if (!firstVariantWritten) { masterPlaylistText += '\n'; firstVariantWritten = true; } if (isKeyPacketsOnly) { masterPlaylistText += `#EXT-X-I-FRAME-STREAM-INF:`; } else { masterPlaylistText += `#EXT-X-STREAM-INF:`; } masterPlaylistText += `BANDWIDTH=${Math.ceil(totalPeakBitrate)}`; if (totalAverageBitrate > 0) { masterPlaylistText += `,AVERAGE-BANDWIDTH=${Math.ceil(totalAverageBitrate)}`; } masterPlaylistText += `,CODECS="${codecs.join(',')}"`; const videoTrack = decl.playlist.tracks.find(x => x.isVideoTrack()); if (videoTrack?.isVideoTrack()) { const trackData = this.trackDatas.find(x => x.track === videoTrack) as HlsVideoTrackData | undefined; const decoderConfig = trackData?.info.decoderConfig; if (decoderConfig) { let width = decoderConfig.displayAspectWidth ?? decoderConfig.codedWidth; let height = decoderConfig.displayAspectHeight ?? decoderConfig.codedHeight; if (width !== undefined && height !== undefined) { if ( videoTrack.metadata.rotation !== undefined && videoTrack.metadata.rotation % 180 === 90 ) { [width, height] = [height, width]; } masterPlaylistText += `,RESOLUTION=${width}x${height}`; } } // FRAME-RATE is not defined for EXT-X-I-FRAME-STREAM-INF if (!isKeyPacketsOnly && videoTrack.metadata.frameRate !== undefined) { // Spec requires that frame rate be rounded to 3 decimal places masterPlaylistText += `,FRAME-RATE=${+videoTrack.metadata.frameRate.toFixed(3)}`; } } if (!isKeyPacketsOnly) { const groupIdForType = new Map(); for (const ref of decl.references) { assert(ref.groupId !== null); const type = ref.playlist.tracks[0]!.type; groupIdForType.set(type, ref.groupId); } for (const [type, id] of groupIdForType) { masterPlaylistText += `,${type.toUpperCase()}="${id}"`; } } if (isKeyPacketsOnly) { // EXT-X-I-FRAME-STREAM-INF is standalone with a URI attribute masterPlaylistText += `,URI="${decl.playlist.path}"`; masterPlaylistText += '\n'; } else { masterPlaylistText += '\n'; masterPlaylistText += `${decl.playlist.path}\n`; } } else { assert(decl.playlist.tracks.length === 1); const track = decl.playlist.tracks[0]!; const type = track.type; let name = track.metadata.name ?? null; const languageCode = track.metadata.languageCode; const disposition = track.metadata.disposition; if (lastGroupId === null || decl.groupId !== lastGroupId) { groupIdTrackCount = 0; masterPlaylistText += '\n'; hasHadDefaultTrackInGroup = false; } lastGroupId = decl.groupId; groupIdTrackCount++; masterPlaylistText += `#EXT-X-MEDIA:TYPE=${type.toUpperCase()},GROUP-ID="${decl.groupId}"`; if (name !== null && /[\n\r"]/.test(name)) { Logging._warn( 'Dropping track name since it includes a line feed, carriage return, or double quote' + ' character, which are not allowed in HLS playlist attributes.', ); name = null; } // Name is required, so we have to set it to SOMETHING name ??= `${languageCode ?? decl.groupId}-${groupIdTrackCount}`; masterPlaylistText += `,NAME="${name}"`; if (languageCode !== undefined) { masterPlaylistText += `,LANGUAGE="${languageCode}"`; } const dispositionPrimary = disposition?.primary ?? false; const dispositionDefault = disposition?.default ?? true; const dispositionForced = disposition?.forced ?? false; if (dispositionPrimary && !hasHadDefaultTrackInGroup) { // HLS's "DEFAULT" behaves like our "primary" masterPlaylistText += ',DEFAULT=YES'; hasHadDefaultTrackInGroup = true; // Only one DEFAULT label per group allowed } if (dispositionPrimary || dispositionDefault) { masterPlaylistText += ',AUTOSELECT=YES'; } if (dispositionForced) { masterPlaylistText += ',FORCED=YES'; } if (type === 'audio') { const trackData = this.trackDatas.find(x => x.track === track) as HlsAudioTrackData | undefined; const decoderConfig = trackData?.info.decoderConfig; if (decoderConfig) { masterPlaylistText += `,CHANNELS="${decoderConfig.numberOfChannels}"`; } } if (!decl.noUri) { masterPlaylistText += `,URI="${decl.playlist.path}"`; } masterPlaylistText += '\n'; } } this.format._options.onMaster?.(masterPlaylistText); const release = await this.mutex.acquire(); try { let writer: Writer; if (this.numWrittenMasterPlaylists === 0) { // For the first master playlist write, we use the normal root writer getter, so that the target // returned by Output.target emits valid write events. writer = await this.output._getRootWriter(true); } else { // For subsequent master playlist writes, we *must* obtain a different target in order to overwrite // the file. const target = await this.output._getTarget({ path: pathedTarget.rootPath, isRoot: true, mimeType: HLS_MIME_TYPE, }); writer = new Writer(target, true); writer.start(); } writer.write(textEncoder.encode(masterPlaylistText)); await writer.flush(); await writer.finalize(); this.numWrittenMasterPlaylists++; } finally { release(); } } private async tryWriteMasterPlaylist() { assert(this.isLive); // The master playlist is written once all playlists have either produced at least one segment or are done for (const playlist of this.playlists) { if (playlist.writtenSegments.length === 0 && !playlist.done) { return; } } await this.writeMasterPlaylist(); } async finalize() { const releases = await Promise.all(this.playlists.map(p => p.mutex.acquire())); releases.forEach(release => release()); for (const trackData of this.trackDatas) { trackData.closed = true; } await Promise.all(this.playlists.map(playlist => ( playlist.done ? Promise.resolve() : this.advancePlaylist(playlist) ))); if (!this.isLive) { await this.writeMasterPlaylist(); } } } const validatePlaylistPath = (path: string) => { if (typeof path !== 'string') { throw new TypeError('options.getPlaylistPath must return or resolve to a string'); } if (/[\n\r"]/.test(path)) { throw new TypeError( 'Playlist paths cannot contain line feed, carriage return, or double quote characters.', ); } }; const validateSegmentPath = (path: string) => { if (typeof path !== 'string') { throw new TypeError('options.getSegmentPath must return or resolve to a string'); } if (/[\n\r"]/.test(path)) { throw new TypeError( 'Segment paths cannot contain line feed or carriage return characters.', ); } }; const validateInitPath = (path: string) => { if (typeof path !== 'string') { throw new TypeError('options.getInitPath must return or resolve to a string'); } if (/[\n\r"]/.test(path)) { throw new TypeError( 'Init paths cannot contain line feed, carriage return, or double quote characters.', ); } }; const toPlaylistInfo = (playlist: Playlist): HlsOutputPlaylistInfo => { return { n: playlist.id, tracks: playlist.tracks, segmentFormat: playlist.segmentFormat, }; }; ===== src/hls/hls-demuxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AUDIO_CODECS, AudioCodec, inferCodecFromCodecString, MediaCodec, VIDEO_CODECS, VideoCodec } from '../codec'; import { Demuxer, DurationMetadataRequestOptions } from '../demuxer'; import { Input } from '../input'; import { InputAudioTrackBacking, InputTrackBacking, InputVideoTrackBacking, } from '../input-track'; import { PacketRetrievalOptions } from '../media-sink'; import { DEFAULT_TRACK_DISPOSITION, MetadataTags, TrackDisposition } from '../metadata'; import { TrackType } from '../output'; import { assert, joinPaths, MaybePromise, Rotation, UNDETERMINED_LANGUAGE } from '../misc'; import { EncodedPacket } from '../packet'; import { readAllLines } from '../reader'; import { AttributeList, canIgnoreLine, HLS_MIME_TYPE, TAG_EXTINF, TAG_I_FRAME_STREAM_INF, TAG_I_FRAMES_ONLY, TAG_MEDIA, TAG_STREAM_INF, } from './hls-misc'; import { HlsSegmentedInput } from './hls-segmented-input'; import { SegmentedInputTrackDeclaration } from '../segmented-input'; import { PathedSource } from '../source'; type InternalTrack = { id: number; demuxer: HlsDemuxer; backingTrack: InputTrackBacking | null; default: boolean; autoselect: boolean; languageCode: string; lineNumber: number; fullPath: string; fullCodecString: string; pairingMask: bigint; peakBitrate: number | null; averageBitrate: number | null; name: string | null; hasOnlyKeyPackets: boolean; info: { type: 'video'; width: number | null; height: number | null; } | { type: 'audio'; numberOfChannels: number | null; }; }; type InternalVideoTrack = InternalTrack & { info: { type: 'video' } }; type InternalAudioTrack = InternalTrack & { info: { type: 'audio' } }; export class HlsDemuxer extends Demuxer { metadataPromise: Promise | null = null; trackBackings: InputTrackBacking[] | null = null; internalTracks: InternalTrack[] | null = null; segmentedInputs: HlsSegmentedInput[] = []; hasMasterPlaylist = true; constructor(input: Input) { super(input); } readMetadata() { return this.metadataPromise ??= (async () => { assert(this.input._rootSource instanceof PathedSource); const slice = await this.input._reader.requestEntireFile(); assert(slice); const lines = readAllLines(slice, slice.length, { ignore: canIgnoreLine }); // Important: get the root path AFTER reading data to get the final root path, possibly affected by // redirects. Any follow requests should be related to the redirected path, not the original one. const { rootPath } = this.input._rootSource; const variantStreams: { fullPath: string; attributes: AttributeList; lineNumber: number; hasOnlyKeyPackets: boolean; }[] = []; const mediaTags: { fullPath: string | null; attributes: AttributeList; lineNumber: number; }[] = []; // Let's first iterate through the entire file, collecting all variant streams and media tags for (let i = 1; i < lines.length; i++) { const line = lines[i]!; if (line.startsWith(TAG_STREAM_INF)) { const streamInfLineNumber = i; const playlistPath = lines[++i]; if (playlistPath === undefined) { throw new Error('Incorrect M3U8 file; a line must follow the #EXT-X-STREAM-INF tag.'); } const fullPath = joinPaths(rootPath, playlistPath); const attributes = new AttributeList(line.slice(TAG_STREAM_INF.length)); const bandwidth = attributes.getAsNumber('bandwidth'); if (bandwidth === null) { throw new Error( 'Invalid M3U8 file; #EXT-X-STREAM-INF tag requires a BANDWIDTH attribute with a valid' + ' numerical value.', ); } variantStreams.push({ fullPath, attributes, lineNumber: streamInfLineNumber, hasOnlyKeyPackets: false, }); } else if (line.startsWith(TAG_I_FRAME_STREAM_INF)) { const attributes = new AttributeList(line.slice(TAG_I_FRAME_STREAM_INF.length)); const playlistPath = attributes.get('uri'); if (playlistPath === null) { throw new Error( 'Invalid M3U8 file; #EXT-X-I-FRAME-STREAM-INF tag requires a URI attribute.', ); } const bandwidth = attributes.getAsNumber('bandwidth'); if (bandwidth === null) { throw new Error( 'Invalid M3U8 file; #EXT-X-I-FRAME-STREAM-INF tag requires a BANDWIDTH attribute with a' + ' valid numerical value.', ); } const fullPath = joinPaths(rootPath, playlistPath); variantStreams.push({ fullPath, attributes, lineNumber: i, hasOnlyKeyPackets: true, }); } else if (line.startsWith(TAG_MEDIA)) { const attributes = new AttributeList(line.slice(TAG_MEDIA.length)); const type = attributes.get('type'); if (type === null) { throw new Error( 'Invalid M3U8 file; #EXT-X-MEDIA tag requires a TYPE attribute.', ); } const groupId = attributes.get('group-id'); if (groupId === null) { throw new Error( 'Invalid M3U8 file; #EXT-X-MEDIA tag requires a GROUP-ID attribute.', ); } let fullPath: string | null = null; const uri = attributes.get('uri'); if (uri !== null) { fullPath = joinPaths(rootPath, uri); } mediaTags.push({ fullPath, attributes, lineNumber: i }); } else if (line === TAG_I_FRAMES_ONLY) { // iFramesOnlyTagFound = true; } else if (line.startsWith(TAG_EXTINF)) { // This is a media playlist, not a master playlist const segmentedInput = new HlsSegmentedInput(this, rootPath, null, lines); this.segmentedInputs = [segmentedInput]; this.hasMasterPlaylist = false; this.trackBackings = await segmentedInput.getTrackBackings(); return; } } const videoGroupIds = [...new Set( mediaTags .filter(tag => tag.attributes.get('type')!.toLowerCase() === 'video') .map(tag => tag.attributes.get('group-id')!)), ]; const audioGroupIds = [...new Set( mediaTags .filter(tag => tag.attributes.get('type')!.toLowerCase() === 'audio') .map(tag => tag.attributes.get('group-id')!)), ]; // Now, let's process & resolve all variant streams in parallel, mapping each of them to tracks. const internalTracksByVariant = await Promise.all(variantStreams.map(async (variantStream, i) => { const result: InternalTrack[] = []; const codecsList = variantStream.attributes.get('codecs'); let codecStrings: string[]; if (codecsList) { codecStrings = codecsList.split(',').map(x => x.trim()); } else { // No codecs were specified, we need to read the underlying media data const segmentedInput = this.getSegmentedInputForPath(variantStream.fullPath); const trackBackings = await segmentedInput.getTrackBackings(); const tracksWithCodec = await Promise.all( trackBackings.map(async t => ({ track: t, codec: await t.getCodec() })), ); codecStrings = await Promise.all( tracksWithCodec .filter(x => x.codec !== null) .map(x => x.track.getDecoderConfig().then(x => x!.codec)), ); } const videoGroupId = variantStream.attributes.get('video'); const audioGroupId = variantStream.attributes.get('audio'); const containsVideoCodecs = codecStrings.some(x => VIDEO_CODECS.includes(inferCodecFromCodecString(x) as VideoCodec), ); const containsAudioCodecs = codecStrings.some(x => AUDIO_CODECS.includes(inferCodecFromCodecString(x) as AudioCodec), ); if (videoGroupId !== null && !containsVideoCodecs) { // A video group is linked but no video codec is listed, sigh. Let's resolve the video codec. if (!videoGroupIds.includes(videoGroupId)) { throw new Error( `Invalid M3U8 file; variant stream references video group "${videoGroupId}" which` + ` is not defined in any #EXT-X-MEDIA tags.`, ); } // We only need to look at the first matching tag, since all tags are required to have the same // codec anyway const matchingVideoMediaTag = mediaTags.find((mediaTag) => { const groupId = mediaTag.attributes.get('group-id')!; const type = mediaTag.attributes.get('type')!; return groupId === videoGroupId && type.toLowerCase() === 'video'; }); outer: if (matchingVideoMediaTag) { const uri = matchingVideoMediaTag.attributes.get('uri'); if (uri === null) { break outer; } const fullPath = joinPaths(rootPath, uri); const segmentedInput = this.getSegmentedInputForPath(fullPath); const trackBackings = await segmentedInput.getTrackBackings(); const videoTrack = trackBackings.find(x => x.getType() === 'video'); if (!videoTrack || (await videoTrack.getCodec()) === null) { break outer; } const additionalCodecString = await videoTrack.getDecoderConfig().then(x => x?.codec ?? null); assert(additionalCodecString !== null); codecStrings.push(additionalCodecString); } } if (audioGroupId !== null && !containsAudioCodecs) { // An audio group is linked but no audio codec is listed, sigh. Let's resolve the audio codec. if (!audioGroupIds.includes(audioGroupId)) { throw new Error( `Invalid M3U8 file; variant stream references audio group "${audioGroupId}" which` + ` is not defined in any #EXT-X-MEDIA tags.`, ); } // We only need to look at the first matching tag, since all tags are required to have the same // codec anyway const matchingAudioMediaTag = mediaTags.find((tag) => { const groupId = tag.attributes.get('group-id')!; const type = tag.attributes.get('type')!; return groupId === audioGroupId && type.toLowerCase() === 'audio'; }); outer: if (matchingAudioMediaTag) { const uri = matchingAudioMediaTag.attributes.get('uri'); if (uri === null) { break outer; } const fullPath = joinPaths(rootPath, uri); const segmentedInput = this.getSegmentedInputForPath(fullPath); const trackBackings = await segmentedInput.getTrackBackings(); const audioTrack = trackBackings.find(x => x.getType() === 'audio'); if (!audioTrack || (await audioTrack.getCodec()) === null) { break outer; } const additionalCodecString = await audioTrack.getDecoderConfig().then(x => x?.codec ?? null); assert(additionalCodecString !== null); codecStrings.push(additionalCodecString); } } // Unique that shit codecStrings = [...new Set(codecStrings)]; let videoCodecString: string | null = null; let audioCodecString: string | null = null; const bandwidth = variantStream.attributes.getAsNumber('bandwidth'); assert(bandwidth !== null); const averageBandwidth = variantStream.attributes.getAsNumber('average-bandwidth'); const name = variantStream.attributes.get('name'); // Now, finally, loop over each codec string for the variant and resolve each one to one or more tracks. for (const codecString of codecStrings) { const inferredCodec = inferCodecFromCodecString(codecString); if (inferredCodec === null) { continue; } if (VIDEO_CODECS.includes(inferredCodec as VideoCodec)) { if (videoCodecString !== null) { throw new Error( 'Unsupported M3U8 file; multiple video codecs found in the CODECS attribute of a' + ' variant stream.', ); } videoCodecString = codecString; const videoGroupId = variantStream.attributes.get('video'); if (videoGroupId === null) { const resolution = variantStream.attributes.get('resolution'); let width: number | null = null; let height: number | null = null; if (resolution) { const match = resolution.match(/^(\d+)x(\d+)$/); if (match) { width = Number(match[1]); height = Number(match[2]); } } result.push({ id: -1, demuxer: this, backingTrack: null, default: true, autoselect: true, languageCode: UNDETERMINED_LANGUAGE, lineNumber: variantStream.lineNumber, fullPath: variantStream.fullPath, fullCodecString: videoCodecString, pairingMask: 1n << BigInt(i), peakBitrate: bandwidth, averageBitrate: averageBandwidth, name, hasOnlyKeyPackets: variantStream.hasOnlyKeyPackets, info: { type: 'video', width, height, }, }); } else { if (!videoGroupIds.includes(videoGroupId)) { throw new Error( `Invalid M3U8 file; variant stream references video group "${videoGroupId}"` + ` which is not defined in any #EXT-X-MEDIA tags.`, ); } for (const mediaTag of mediaTags) { const groupId = mediaTag.attributes.get('group-id')!; const type = mediaTag.attributes.get('type')!; if (groupId !== videoGroupId || type.toLowerCase() !== 'video') { continue; } const resolution = mediaTag.attributes.get('resolution') ?? variantStream.attributes.get('resolution'); let width: number | null = null; let height: number | null = null; if (resolution) { const match = resolution.match(/^(\d+)x(\d+)$/); if (match) { width = Number(match[1]); height = Number(match[2]); } } result.push({ id: -1, demuxer: this, backingTrack: null, default: getMediaTagDefault(mediaTag.attributes), // Autoselect is inferred to be true if the default is true autoselect: getMediaTagDefault(mediaTag.attributes) || getMediaTagAutoselect(mediaTag.attributes), languageCode: preprocessLanguageCode(mediaTag.attributes.get('language')), lineNumber: mediaTag.lineNumber, fullPath: mediaTag.fullPath ?? variantStream.fullPath, fullCodecString: videoCodecString, pairingMask: 1n << BigInt(i), peakBitrate: null, averageBitrate: null, name: mediaTag.attributes.get('name'), hasOnlyKeyPackets: variantStream.hasOnlyKeyPackets, info: { type: 'video', width, height, }, }); } } } else if (AUDIO_CODECS.includes(inferredCodec as AudioCodec)) { if (audioCodecString !== null) { throw new Error( 'Unsupported M3U8 file; multiple audio codecs found in the CODECS attribute of a' + ' variant stream.', ); } audioCodecString = codecString; const audioGroupId = variantStream.attributes.get('audio'); if (audioGroupId === null) { const channels = variantStream.attributes.get('channels'); const parsedChannels = channels !== null ? Number(channels.split('/')[0]!) : null; result.push({ id: -1, demuxer: this, backingTrack: null, default: true, autoselect: true, languageCode: UNDETERMINED_LANGUAGE, lineNumber: variantStream.lineNumber, fullPath: variantStream.fullPath, fullCodecString: audioCodecString, pairingMask: 1n << BigInt(i), peakBitrate: bandwidth, averageBitrate: averageBandwidth, name, hasOnlyKeyPackets: variantStream.hasOnlyKeyPackets, info: { type: 'audio', numberOfChannels: parsedChannels !== null && Number.isInteger(parsedChannels) && parsedChannels > 0 ? parsedChannels : null, }, }); } else { if (!audioGroupIds.includes(audioGroupId)) { throw new Error( `Invalid M3U8 file; variant stream references audio group "${audioGroupId}"` + ` which is not defined in any #EXT-X-MEDIA tags.`, ); } for (const mediaTag of mediaTags) { const groupId = mediaTag.attributes.get('group-id')!; const type = mediaTag.attributes.get('type')!; if (groupId !== audioGroupId || type.toLowerCase() !== 'audio') { continue; } const channels = mediaTag.attributes.get('channels') ?? variantStream.attributes.get('channels'); const parsedChannels = channels !== null ? Number(channels.split('/')[0]!) : null; result.push({ id: -1, demuxer: this, backingTrack: null, default: getMediaTagDefault(mediaTag.attributes), // Autoselect is inferred to be true if the default is true autoselect: getMediaTagDefault(mediaTag.attributes) || getMediaTagAutoselect(mediaTag.attributes), languageCode: preprocessLanguageCode(mediaTag.attributes.get('language')), lineNumber: mediaTag.lineNumber, fullPath: mediaTag.fullPath ?? variantStream.fullPath, fullCodecString: audioCodecString, pairingMask: 1n << BigInt(i), peakBitrate: null, averageBitrate: null, name: mediaTag.attributes.get('name'), hasOnlyKeyPackets: variantStream.hasOnlyKeyPackets, info: { type: 'audio', numberOfChannels: parsedChannels !== null && Number.isInteger(parsedChannels) && parsedChannels > 0 ? parsedChannels : null, }, }); } } } } return result; })); const internalTracks: InternalTrack[] = []; const addInternalTrack = (track: InternalTrack) => { const existingTrack = internalTracks.find(x => x.fullPath === track.fullPath && x.info.type === track.info.type, ); if (existingTrack) { existingTrack.pairingMask |= track.pairingMask; existingTrack.default ||= track.default; existingTrack.autoselect ||= track.autoselect; existingTrack.lineNumber = Math.min(existingTrack.lineNumber, track.lineNumber); if (track.peakBitrate !== null) { existingTrack.peakBitrate = Math.max( existingTrack.peakBitrate ?? -Infinity, track.peakBitrate, ); } if (track.averageBitrate !== null) { existingTrack.averageBitrate = Math.max( existingTrack.averageBitrate ?? -Infinity, track.averageBitrate, ); } if (existingTrack.languageCode === UNDETERMINED_LANGUAGE) { existingTrack.languageCode = track.languageCode; } } else { track.id = internalTracks.length + 1; internalTracks.push(track); } }; for (const variantInternalTracks of internalTracksByVariant) { for (const trackEntry of variantInternalTracks) { addInternalTrack(trackEntry); } } // Order tracks by how they appear in the file internalTracks.sort((a, b) => a.lineNumber - b.lineNumber); this.trackBackings = []; for (const internalTrack of internalTracks) { if (internalTrack.info.type === 'video') { this.trackBackings.push( new HlsInputVideoTrackBacking(internalTrack as InternalVideoTrack), ); } else { this.trackBackings.push( new HlsInputAudioTrackBacking(internalTrack as InternalAudioTrack), ); } } this.internalTracks = internalTracks; })(); } async getTrackBackings() { await this.readMetadata(); assert(this.trackBackings); return this.trackBackings; } getSegmentedInputForPath(path: string) { let segmentedInput = this.segmentedInputs.find(x => x.path === path); if (segmentedInput) { return segmentedInput; } let decls: SegmentedInputTrackDeclaration[] | null = null; if (this.internalTracks) { const tracks = this.internalTracks.filter(x => x.fullPath === path); decls = tracks.map(x => ({ id: x.id, type: x.info.type, })); } segmentedInput = new HlsSegmentedInput(this, path, decls, null); this.segmentedInputs.push(segmentedInput); return segmentedInput; } async getMetadataTags(): Promise { return {}; } async getMimeType(): Promise { return HLS_MIME_TYPE; } override dispose(): void { if (this.segmentedInputs) { for (const segInput of this.segmentedInputs) { segInput.dispose(); } this.segmentedInputs.length = 0; } } } abstract class HlsInputTrackBacking implements InputTrackBacking { hydrationPromise: Promise | null = null; constructor(public internalTrack: InternalTrack) {} abstract getType(): TrackType; abstract getDecoderConfig(): Promise; hydrate() { return this.hydrationPromise ??= (async () => { const segmentedInput = this.internalTrack.demuxer.getSegmentedInputForPath(this.internalTrack.fullPath); let trackBacking: InputTrackBacking | null = null; const trackBackings = await segmentedInput.getTrackBackings(); const matchingType = trackBackings.filter(x => x.getType() === this.getType()); if (matchingType.length === 1) { // Avoids reading fields on the track trackBacking = matchingType[0]!; } else { if (this instanceof HlsInputVideoTrackBacking) { for (const backing of matchingType) { if ((await backing.getCodec()) === this.getCodec()) { trackBacking = backing; break; } } } else { assert(this instanceof HlsInputAudioTrackBacking); for (const backing of matchingType) { if ((await backing.getCodec()) === this.getCodec()) { trackBacking = backing; break; } } } } if (!trackBacking) { throw new Error('Could not find matching track in underlying media data.'); } this.internalTrack.backingTrack = trackBacking; })(); } /** If the backing track is already present, delegate synchronously; otherwise, hydrate first. */ delegate(fn: () => MaybePromise): MaybePromise { if (this.internalTrack.backingTrack) { return fn(); } return this.hydrate().then(fn); } getCodec(): MediaCodec | null { throw new Error('Not implemented on base class.'); } getDisposition(): TrackDisposition { return { ...DEFAULT_TRACK_DISPOSITION, // Meanings are swapped in HLS: "Default" means that a track is the primary track. default: this.internalTrack.autoselect, primary: this.internalTrack.default, }; } getId(): number { return this.internalTrack.id; } getPairingMask(): bigint { return this.internalTrack.pairingMask; } getInternalCodecId(): string | number | Uint8Array | null { return null; } getLanguageCode(): string { return this.internalTrack.languageCode; } getName(): string | null { return this.internalTrack.name; } getNumber(): number { assert(this.internalTrack.demuxer.internalTracks); const trackType = this.internalTrack.info.type; let number = 0; for (const track of this.internalTrack.demuxer.internalTracks) { if (track.info.type === trackType) { number++; } if (track === this.internalTrack) { break; } } return number; } getTimeResolution(): MaybePromise { return this.delegate(() => this.internalTrack.backingTrack!.getTimeResolution()); } isRelativeToUnixEpoch(): MaybePromise { return this.delegate(() => this.internalTrack.backingTrack!.isRelativeToUnixEpoch()); } getUnixTimeForTimestamp(timestamp: number): MaybePromise { return this.delegate(() => this.internalTrack.backingTrack!.getUnixTimeForTimestamp(timestamp)); } getBitrate(): number | null { return this.internalTrack.peakBitrate; } getAverageBitrate(): number | null { return this.internalTrack.averageBitrate; } async getDurationFromMetadata(options: DurationMetadataRequestOptions): Promise { await this.hydrate(); return this.internalTrack.backingTrack!.getDurationFromMetadata(options); } async getLiveRefreshInterval(): Promise { await this.hydrate(); return this.internalTrack.backingTrack!.getLiveRefreshInterval(); } getHasOnlyKeyPackets() { return this.internalTrack.hasOnlyKeyPackets || null; } async getFirstPacket(options: PacketRetrievalOptions): Promise { await this.hydrate(); return this.internalTrack.backingTrack!.getFirstPacket(options); } async getPacket(timestamp: number, options: PacketRetrievalOptions): Promise { await this.hydrate(); return this.internalTrack.backingTrack!.getPacket(timestamp, options); } async getKeyPacket(timestamp: number, options: PacketRetrievalOptions): Promise { await this.hydrate(); return this.internalTrack.backingTrack!.getKeyPacket(timestamp, options); } async getNextPacket(packet: EncodedPacket, options: PacketRetrievalOptions): Promise { await this.hydrate(); return this.internalTrack.backingTrack!.getNextPacket(packet, options); } async getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions): Promise { await this.hydrate(); return this.internalTrack.backingTrack!.getNextKeyPacket(packet, options); } } class HlsInputVideoTrackBacking extends HlsInputTrackBacking implements InputVideoTrackBacking { override internalTrack!: InternalVideoTrack; constructor(internalTrack: InternalVideoTrack) { super(internalTrack); } get backingVideoTrack() { return this.internalTrack.backingTrack as InputVideoTrackBacking | null; } getType() { return 'video' as const; } override getCodec(): VideoCodec | null { const inferredCodec = inferCodecFromCodecString(this.internalTrack.fullCodecString); return inferredCodec as VideoCodec; } getCodedWidth(): MaybePromise { return this.delegate(() => this.backingVideoTrack!.getCodedWidth()); } getCodedHeight(): MaybePromise { return this.delegate(() => this.backingVideoTrack!.getCodedHeight()); } getSquarePixelWidth(): MaybePromise { return this.delegate(() => this.backingVideoTrack!.getSquarePixelWidth()); } getSquarePixelHeight(): MaybePromise { return this.delegate(() => this.backingVideoTrack!.getSquarePixelHeight()); } getMetadataDisplayWidth(): number | null { if (this.backingVideoTrack) { return null; } return this.internalTrack.info.width; } getMetadataDisplayHeight(): number | null { if (this.backingVideoTrack) { return null; } return this.internalTrack.info.height; } getRotation(): MaybePromise { return this.delegate(() => this.backingVideoTrack!.getRotation()); } async getColorSpace(): Promise { await this.hydrate(); return this.backingVideoTrack!.getColorSpace(); } async canBeTransparent(): Promise { await this.hydrate(); return this.backingVideoTrack!.canBeTransparent(); } getMetadataCodecParameterString(): string | null { if (this.backingVideoTrack) { return null; } return this.internalTrack.fullCodecString; } async getDecoderConfig(): Promise { await this.hydrate(); return this.backingVideoTrack!.getDecoderConfig(); } } class HlsInputAudioTrackBacking extends HlsInputTrackBacking implements InputAudioTrackBacking { override internalTrack!: InternalAudioTrack; constructor(internalTrack: InternalAudioTrack) { super(internalTrack); } get backingAudioTrack() { return this.internalTrack.backingTrack as InputAudioTrackBacking | null; } getType() { return 'audio' as const; } override getCodec(): AudioCodec | null { const inferredCodec = inferCodecFromCodecString(this.internalTrack.fullCodecString); return inferredCodec as AudioCodec; } getNumberOfChannels(): MaybePromise { if (this.internalTrack.info.numberOfChannels !== null) { return this.internalTrack.info.numberOfChannels; } return this.delegate(() => this.backingAudioTrack!.getNumberOfChannels()); } getSampleRate(): MaybePromise { return this.delegate(() => this.backingAudioTrack!.getSampleRate()); } getMetadataCodecParameterString(): string | null { if (this.backingAudioTrack) { return null; } return this.internalTrack.fullCodecString; } async getDecoderConfig(): Promise { await this.hydrate(); return this.backingAudioTrack!.getDecoderConfig(); } } const getMediaTagDefault = (attributes: AttributeList) => { const value = attributes.get('default'); if (value === null) { return false; } const normalized = value.toUpperCase(); if (normalized === 'YES') { return true; } if (normalized === 'NO') { return false; } throw new Error( `Invalid M3U8 file; #EXT-X-MEDIA DEFAULT attribute must be YES or NO, got "${value}".`, ); }; const getMediaTagAutoselect = (attributes: AttributeList) => { const value = attributes.get('autoselect'); if (value === null) { return false; } const normalized = value.toUpperCase(); if (normalized === 'YES') { return true; } if (normalized === 'NO') { return false; } throw new Error( `Invalid M3U8 file; #EXT-X-MEDIA AUTOSELECT attribute must be YES or NO, got "${value}".`, ); }; const preprocessLanguageCode = (code: string | null) => { if (code === null) { return UNDETERMINED_LANGUAGE; } const languageSubtag = code.split('-')[0]; if (!languageSubtag) { return UNDETERMINED_LANGUAGE; } // Technically invalid, for now: The language subtag might be a language code from ISO 639-1, // ISO 639-2, ISO 639-3, ISO 639-5 or some other thing (source: Wikipedia). But, `languageCode` is // documented as ISO 639-2. Changing the definition would be a breaking change. This will get // cleaned up in the future by defining languageCode to be BCP 47 instead. return languageSubtag; }; ===== src/hls/hls-misc.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ export const HLS_MIME_TYPE = 'application/vnd.apple.mpegurl'; export const TAG_STREAM_INF = '#EXT-X-STREAM-INF:'; export const TAG_I_FRAME_STREAM_INF = '#EXT-X-I-FRAME-STREAM-INF:'; export const TAG_MEDIA = '#EXT-X-MEDIA:'; export const TAG_EXTINF = '#EXTINF:'; export const TAG_MAP = '#EXT-X-MAP:'; export const TAG_KEY = '#EXT-X-KEY:'; export const TAG_MEDIA_SEQUENCE = '#EXT-X-MEDIA-SEQUENCE:'; export const TAG_BYTERANGE = '#EXT-X-BYTERANGE:'; export const TAG_PROGRAM_DATE_TIME = '#EXT-X-PROGRAM-DATE-TIME:'; export const TAG_DISCONTINUITY = '#EXT-X-DISCONTINUITY'; export const TAG_TARGETDURATION = '#EXT-X-TARGETDURATION:'; export const TAG_ENDLIST = '#EXT-X-ENDLIST'; export const TAG_PLAYLIST_TYPE = '#EXT-X-PLAYLIST-TYPE:'; export const TAG_I_FRAMES_ONLY = '#EXT-X-I-FRAMES-ONLY'; export const canIgnoreLine = (line: string) => line.length === 0 || (line.startsWith('#') && !line.startsWith('#EXT')); export class AttributeList { _attributes: Record = {}; constructor(str: string) { let key = ''; let value = ''; let inValue = false; let inQuotes = false; for (let i = 0; i < str.length; i++) { const char = str[i]!; if (char === '"') { inQuotes = !inQuotes; } else if (char === '=' && !inValue && !inQuotes) { inValue = true; } else if (char === ',' && !inQuotes) { if (key) { this._attributes[key.trim().toLowerCase()] = value; } key = ''; value = ''; inValue = false; } else if (inValue) { value += char; } else { key += char; } } if (key) { this._attributes[key.trim().toLowerCase()] = value; } } get(name: string) { return this._attributes[name.toLowerCase()] ?? null; } getAsNumber(name: string) { const value = this.get(name); if (value === null) { return null; } const num = Number(value); return Number.isFinite(num) ? num : null; } merge(other: AttributeList) { Object.assign(this._attributes, other._attributes); } } ===== src/mp3/mp3-muxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { toDataView } from '../misc'; import { metadataTagsAreEmpty } from '../metadata'; import { Muxer } from '../muxer'; import { Output, OutputAudioTrack } from '../output'; import { Mp3OutputFormat } from '../output-format'; import { EncodedPacket } from '../packet'; import { Writer } from '../writer'; import { getXingOffset, INFO, readMp3FrameHeader, XING } from '../../shared/mp3-misc'; import { Mp3Writer, XingFrameData } from './mp3-writer'; import { Id3V2Writer } from '../id3'; export class Mp3Muxer extends Muxer { private format: Mp3OutputFormat; private writer!: Writer; private mp3Writer!: Mp3Writer; private xingFrameData: XingFrameData | null = null; private frameCount = 0; private framePositions: number[] = []; private xingFramePos: number | null = null; constructor(output: Output, format: Mp3OutputFormat) { super(output); this.format = format; } async start() { const release = await this.mutex.acquire(); this.writer = await this.output._getRootWriter(this.format._options.xingHeader === false); this.mp3Writer = new Mp3Writer(this.writer); if (!metadataTagsAreEmpty(this.output._metadataTags)) { const id3Writer = new Id3V2Writer(this.writer); id3Writer.writeId3V2Tag(this.output._metadataTags); } release(); } async getMimeType() { return 'audio/mpeg'; } async addEncodedVideoPacket() { throw new Error('MP3 does not support video.'); } async addEncodedAudioPacket( track: OutputAudioTrack, packet: EncodedPacket, ) { const release = await this.mutex.acquire(); try { const writeXingHeader = this.format._options.xingHeader !== false; if (!this.xingFrameData && writeXingHeader) { const view = toDataView(packet.data); if (view.byteLength < 4) { throw new Error('Invalid MP3 header in sample.'); } const word = view.getUint32(0, false); const header = readMp3FrameHeader(word, null).header; if (!header) { throw new Error('Invalid MP3 header in sample.'); } const xingOffset = getXingOffset(header.mpegVersionId, header.channel); if (view.byteLength >= xingOffset + 4) { const word = view.getUint32(xingOffset, false); const isXing = word === XING || word === INFO; if (isXing) { // This is not a data frame, so let's completely ignore this sample return; } } this.xingFrameData = { mpegVersionId: header.mpegVersionId, layer: header.layer, frequencyIndex: header.frequencyIndex, sampleRate: header.sampleRate, channel: header.channel, modeExtension: header.modeExtension, copyright: header.copyright, original: header.original, emphasis: header.emphasis, frameCount: null, fileSize: null, toc: null, }; // Write a Xing frame because this muxer doesn't make any bitrate constraints, meaning we don't know if // this will be a constant or variable bitrate file. Therefore, always write the Xing frame. this.xingFramePos = this.writer.getPos(); this.mp3Writer.writeXingFrame(this.xingFrameData); this.frameCount++; } this.validateTimestamp(track, packet.timestamp, packet.type === 'key'); if (writeXingHeader) { this.framePositions.push(this.writer.getPos()); } this.writer.write(packet.data); this.frameCount++; await this.writer.flush(); } finally { release(); } } async addSubtitleCue() { throw new Error('MP3 does not support subtitles.'); } async finalize() { if (!this.xingFrameData || this.xingFramePos === null) { return; } const release = await this.mutex.acquire(); const endPos = this.writer.getPos(); const audioDataEndPos = endPos - this.xingFramePos; this.writer.seek(this.xingFramePos); const toc = new Uint8Array(100); for (let i = 0; i < 100; i++) { const index = Math.floor(this.framePositions.length * (i / 100)); const byteOffset = this.framePositions[index]! - this.xingFramePos; toc[i] = 256 * (byteOffset / audioDataEndPos); } this.xingFrameData.frameCount = this.frameCount; this.xingFrameData.fileSize = audioDataEndPos; this.xingFrameData.toc = toc; if (this.format._options.onXingFrame) { this.writer.startTrackingWrites(); } this.mp3Writer.writeXingFrame(this.xingFrameData); if (this.format._options.onXingFrame) { const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onXingFrame(data, start); } release(); } } ===== src/mp3/mp3-reader.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { MP3_FRAME_HEADER_SIZE, getMp3ChannelCount, Mp3FrameHeader, readMp3FrameHeader } from '../../shared/mp3-misc'; import { Reader, readU32Be } from '../reader'; export const readNextMp3FrameHeader = async ( reader: Reader, startPos: number, until: number | null, ref: Mp3FrameHeader | null = null, ): Promise<{ header: Mp3FrameHeader; startPos: number; } | null> => { const CHUNK_SIZE = 2 ** 16; // So we don't need to grab thousands of slices let currentPos = startPos; while (until === null || currentPos < until) { const maxLength = until !== null ? Math.min(CHUNK_SIZE, until - currentPos) : CHUNK_SIZE; let slice = reader.requestSliceRange(currentPos, MP3_FRAME_HEADER_SIZE, maxLength); if (slice instanceof Promise) slice = await slice; if (!slice || slice.length < MP3_FRAME_HEADER_SIZE) break; while (slice.remainingLength >= MP3_FRAME_HEADER_SIZE) { const posBeforeRead = slice.filePos; const word = readU32Be(slice); const remainingBytes = reader.fileSize !== null ? reader.fileSize - currentPos : null; const result = readMp3FrameHeader(word, remainingBytes); if ( result.header && (!ref || ( // This condition helps us recover malformed streams // https://stackoverflow.com/a/20884944 result.header.sampleRate === ref.sampleRate && result.header.mpegVersionId === ref.mpegVersionId && result.header.layer === ref.layer && getMp3ChannelCount(result.header.channel) === getMp3ChannelCount(ref.channel) )) ) { return { header: result.header, startPos: currentPos }; } slice.filePos = posBeforeRead + result.bytesAdvanced; currentPos = slice.filePos; } } return null; }; ===== src/mp3/mp3-writer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Writer } from '../writer'; import { computeMp3FrameSize, getXingOffset, KILOBIT_RATES, XING, XingFlags, } from '../../shared/mp3-misc'; export type XingFrameData = { mpegVersionId: number; layer: number; frequencyIndex: number; sampleRate: number; channel: number; modeExtension: number; copyright: number; original: number; emphasis: number; frameCount: number | null; fileSize: number | null; toc: Uint8Array | null; }; export class Mp3Writer { private helper = new Uint8Array(8); private helperView = new DataView(this.helper.buffer); constructor(private writer: Writer) {} writeU32(value: number) { this.helperView.setUint32(0, value, false); this.writer.write(this.helper.subarray(0, 4)); } writeXingFrame(data: XingFrameData) { const startPos = this.writer.getPos(); const firstByte = 0xff; const secondByte = 0xe0 | (data.mpegVersionId << 3) | (data.layer << 1); let lowSamplingFrequency: number; if (data.mpegVersionId & 2) { lowSamplingFrequency = (data.mpegVersionId & 1) ? 0 : 1; } else { lowSamplingFrequency = 1; } const padding = 0; const neededBytes = 155; let bitrateIndex = -1; const bitrateOffset = lowSamplingFrequency * 16 * 4 + data.layer * 16; // Let's find the lowest bitrate for which the frame size is sufficiently large to fit all the data for (let i = 0; i < 16; i++) { const kbr = KILOBIT_RATES[bitrateOffset + i]!; const size = computeMp3FrameSize(lowSamplingFrequency, data.layer, 1000 * kbr, data.sampleRate, padding); if (size >= neededBytes) { bitrateIndex = i; break; } } if (bitrateIndex === -1) { throw new Error('No suitable bitrate found.'); } const thirdByte = (bitrateIndex << 4) | (data.frequencyIndex << 2) | padding << 1; const fourthByte = (data.channel << 6) | (data.modeExtension << 4) | (data.copyright << 3) | (data.original << 2) | data.emphasis; this.helper[0] = firstByte; this.helper[1] = secondByte; this.helper[2] = thirdByte; this.helper[3] = fourthByte; this.writer.write(this.helper.subarray(0, 4)); const xingOffset = getXingOffset(data.mpegVersionId, data.channel); this.writer.seek(startPos + xingOffset); this.writeU32(XING); let flags = 0; if (data.frameCount !== null) { flags |= XingFlags.FrameCount; } if (data.fileSize !== null) { flags |= XingFlags.FileSize; } if (data.toc !== null) { flags |= XingFlags.Toc; } this.writeU32(flags); this.writeU32(data.frameCount ?? 0); this.writeU32(data.fileSize ?? 0); this.writer.write(data.toc ?? new Uint8Array(100)); const kilobitRate = KILOBIT_RATES[bitrateOffset + bitrateIndex]!; const frameSize = computeMp3FrameSize( lowSamplingFrequency, data.layer, 1000 * kilobitRate, data.sampleRate, padding, ); this.writer.seek(startPos + frameSize); } } ===== src/mp3/mp3-demuxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AudioCodec } from '../codec'; import { Demuxer } from '../demuxer'; import { Input } from '../input'; import { InputAudioTrackBacking } from '../input-track'; import { DEFAULT_TRACK_DISPOSITION, MetadataTags } from '../metadata'; import { PacketRetrievalOptions } from '../media-sink'; import { assert, AsyncMutex, binarySearchExact, binarySearchLessOrEqual, toDataView, UNDETERMINED_LANGUAGE, } from '../misc'; import { EncodedPacket, PLACEHOLDER_DATA } from '../packet'; import { Mp3FrameHeader, getXingOffset, INFO, XING, XingFlags, computeAverageMp3FrameSize, getMp3ChannelCount, } from '../../shared/mp3-misc'; import { ID3_V1_TAG_SIZE, ID3_V2_HEADER_SIZE, parseId3V1Tag, parseId3V2Tag, readId3V2Header, } from '../id3'; import { readNextMp3FrameHeader } from './mp3-reader'; import { readAscii, readBytes, Reader, readU32Be } from '../reader'; type Sample = { timestamp: number; duration: number; dataStart: number; dataSize: number; }; export class Mp3Demuxer extends Demuxer { reader: Reader; metadataPromise: Promise | null = null; firstFrameHeader: Mp3FrameHeader | null = null; firstFrameHeaderPos: number | null = null; loadedSamples: Sample[] = []; // All samples from the start of the file to lastLoadedPos metadataTags: MetadataTags | null = null; xingData: { frameCount: number | null; fileSize: number | null; } | null = null; trackBackings: Mp3AudioTrackBacking[] = []; readingMutex = new AsyncMutex(); lastSampleLoaded = false; lastLoadedPos = 0; nextTimestampInSamples = 0; constructor(input: Input) { super(input); this.reader = input._reader; } async readMetadata() { return this.metadataPromise ??= (async () => { // Keep loading until we find the first frame header while (!this.firstFrameHeader && !this.lastSampleLoaded) { await this.advanceReader(); } if (!this.firstFrameHeader) { throw new Error('No valid MP3 frame found.'); } this.trackBackings = [new Mp3AudioTrackBacking(this)]; })(); } async advanceReader() { if (this.lastLoadedPos === 0) { // Let's skip all ID3v2 tags at the start of the file while (true) { let slice = this.reader.requestSlice(this.lastLoadedPos, ID3_V2_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) { this.lastSampleLoaded = true; return; } const id3V2Header = readId3V2Header(slice); if (!id3V2Header) { break; } this.lastLoadedPos = slice.filePos + id3V2Header.size; } } const result = await readNextMp3FrameHeader( this.reader, this.lastLoadedPos, this.reader.fileSize, this.firstFrameHeader, ); if (!result) { this.lastSampleLoaded = true; return; } const header = result.header; this.lastLoadedPos = result.startPos + header.totalSize - 1; // -1 in case the frame is 1 byte too short const xingOffset = getXingOffset(header.mpegVersionId, header.channel); let slice = this.reader.requestSlice(result.startPos + xingOffset, 4); if (slice instanceof Promise) slice = await slice; if (slice) { const word = readU32Be(slice); const isXing = word === XING || word === INFO; if (isXing) { // There's no actual audio data in this frame, so let's skip it if (!this.xingData) { let xingDataSlice = this.reader.requestSlice(result.startPos + xingOffset + 4, 12); if (xingDataSlice instanceof Promise) xingDataSlice = await xingDataSlice; if (xingDataSlice) { const xingData = readBytes(xingDataSlice, 12); const view = toDataView(xingData); const flags = view.getUint32(0, false); this.xingData = { frameCount: (flags & XingFlags.FrameCount) ? view.getUint32(4, false) : null, fileSize: (flags & XingFlags.FileSize) ? view.getUint32(8, false) : null, }; } } return; } } if (!this.firstFrameHeader) { this.firstFrameHeader = header; this.firstFrameHeaderPos = result.startPos; } const sampleDuration = header.audioSamplesInFrame / this.firstFrameHeader.sampleRate; const sample: Sample = { timestamp: this.nextTimestampInSamples / this.firstFrameHeader.sampleRate, duration: sampleDuration, dataStart: result.startPos, dataSize: header.totalSize, }; this.loadedSamples.push(sample); this.nextTimestampInSamples += header.audioSamplesInFrame; return; } async getMimeType() { return 'audio/mpeg'; } async getTrackBackings() { await this.readMetadata(); return this.trackBackings; } async getMetadataTags() { const release = await this.readingMutex.acquire(); try { await this.readMetadata(); if (this.metadataTags) { return this.metadataTags; } this.metadataTags = {}; let currentPos = 0; let id3V2HeaderFound = false; while (true) { let headerSlice = this.reader.requestSlice(currentPos, ID3_V2_HEADER_SIZE); if (headerSlice instanceof Promise) headerSlice = await headerSlice; if (!headerSlice) break; const id3V2Header = readId3V2Header(headerSlice); if (!id3V2Header) { break; } id3V2HeaderFound = true; let contentSlice = this.reader.requestSlice(headerSlice.filePos, id3V2Header.size); if (contentSlice instanceof Promise) contentSlice = await contentSlice; if (!contentSlice) break; parseId3V2Tag(contentSlice, id3V2Header, this.metadataTags); currentPos = headerSlice.filePos + id3V2Header.size; } if (!id3V2HeaderFound && this.reader.fileSize !== null && this.reader.fileSize >= ID3_V1_TAG_SIZE) { // Try reading an ID3v1 tag at the end of the file let slice = this.reader.requestSlice(this.reader.fileSize - ID3_V1_TAG_SIZE, ID3_V1_TAG_SIZE); if (slice instanceof Promise) slice = await slice; assert(slice); const tag = readAscii(slice, 3); if (tag === 'TAG') { parseId3V1Tag(slice, this.metadataTags); } } return this.metadataTags; } finally { release(); } } } class Mp3AudioTrackBacking implements InputAudioTrackBacking { constructor(public demuxer: Mp3Demuxer) {} getType() { return 'audio' as const; } getId() { return 1; } getNumber() { return 1; } getTimeResolution() { assert(this.demuxer.firstFrameHeader); return this.demuxer.firstFrameHeader.sampleRate / this.demuxer.firstFrameHeader.audioSamplesInFrame; } isRelativeToUnixEpoch() { return false; } getUnixTimeForTimestamp() { return null; } getPairingMask() { return 1n; } getBitrate() { return null; } getAverageBitrate() { return null; } async getDurationFromMetadata() { const demuxer = this.demuxer; assert(demuxer.firstFrameHeader !== null); assert(demuxer.firstFrameHeaderPos !== null); if (demuxer.xingData) { if (demuxer.xingData.frameCount !== null) { return demuxer.xingData.frameCount * demuxer.firstFrameHeader.audioSamplesInFrame / demuxer.firstFrameHeader.sampleRate; } } else { // No Xing, assuming CBR if (demuxer.reader.fileSize !== null) { const averageFrameSize = computeAverageMp3FrameSize( demuxer.firstFrameHeader.lowSamplingFrequency, demuxer.firstFrameHeader.layer, demuxer.firstFrameHeader.bitrate, demuxer.firstFrameHeader.sampleRate, ); const frameCount = (demuxer.reader.fileSize - demuxer.firstFrameHeaderPos) / averageFrameSize; return Math.round(frameCount) * demuxer.firstFrameHeader.audioSamplesInFrame / demuxer.firstFrameHeader.sampleRate; } } return null; } async getLiveRefreshInterval() { return null; } getName() { return null; } getLanguageCode() { return UNDETERMINED_LANGUAGE; } getCodec(): AudioCodec { return 'mp3'; } getInternalCodecId() { return null; } getNumberOfChannels() { assert(this.demuxer.firstFrameHeader); return getMp3ChannelCount(this.demuxer.firstFrameHeader.channel); } getSampleRate() { assert(this.demuxer.firstFrameHeader); return this.demuxer.firstFrameHeader.sampleRate; } getDisposition() { return { ...DEFAULT_TRACK_DISPOSITION, }; } async getDecoderConfig(): Promise { assert(this.demuxer.firstFrameHeader); return { codec: 'mp3', numberOfChannels: getMp3ChannelCount(this.demuxer.firstFrameHeader.channel), sampleRate: this.demuxer.firstFrameHeader.sampleRate, }; } async getPacketAtIndex(sampleIndex: number, options: PacketRetrievalOptions) { if (sampleIndex === -1) { return null; } const rawSample = this.demuxer.loadedSamples[sampleIndex]; if (!rawSample) { return null; } let data: Uint8Array; if (options.metadataOnly) { data = PLACEHOLDER_DATA; } else { let slice = this.demuxer.reader.requestSlice(rawSample.dataStart, rawSample.dataSize); if (slice instanceof Promise) slice = await slice; if (!slice) { return null; // Data didn't fit into the rest of the file } data = readBytes(slice, rawSample.dataSize); } return new EncodedPacket( data, 'key', rawSample.timestamp, rawSample.duration, sampleIndex, rawSample.dataSize, ); } getFirstPacket(options: PacketRetrievalOptions) { return this.getPacketAtIndex(0, options); } async getNextPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { const release = await this.demuxer.readingMutex.acquire(); try { const sampleIndex = binarySearchExact( this.demuxer.loadedSamples, packet.timestamp, x => x.timestamp, ); if (sampleIndex === -1) { throw new Error('Packet was not created from this track.'); } const nextIndex = sampleIndex + 1; // Ensure the next sample exists while ( nextIndex >= this.demuxer.loadedSamples.length && !this.demuxer.lastSampleLoaded ) { await this.demuxer.advanceReader(); } return this.getPacketAtIndex(nextIndex, options); } finally { release(); } } async getPacket(timestamp: number, options: PacketRetrievalOptions) { const release = await this.demuxer.readingMutex.acquire(); try { while (true) { const index = binarySearchLessOrEqual( this.demuxer.loadedSamples, timestamp, x => x.timestamp, ); if (index === -1 && this.demuxer.loadedSamples.length > 0) { // We're before the first sample return null; } if (this.demuxer.lastSampleLoaded) { // All data is loaded, return what we found return this.getPacketAtIndex(index, options); } if (index >= 0 && index + 1 < this.demuxer.loadedSamples.length) { // The next packet also exists, we're done return this.getPacketAtIndex(index, options); } // Otherwise, keep loading data await this.demuxer.advanceReader(); } } finally { release(); } } getKeyPacket(timestamp: number, options: PacketRetrievalOptions) { return this.getPacket(timestamp, options); } getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { return this.getNextPacket(packet, options); } } ===== src/aes.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { assert, MaybePromise } from './misc'; import { readBytes, Reader } from './reader'; // Inspired in part by https://github.com/halloweeks/AES-128-CBC/blob/main/AES_128_CBC.h export const AES_128_BLOCK_SIZE = 16; const Te4 = new Uint32Array(256); const Td0 = new Uint32Array(256); const Td1 = new Uint32Array(256); const Td2 = new Uint32Array(256); const Td3 = new Uint32Array(256); const Td4 = new Uint32Array(256); const rcon = new Uint32Array(10); let tablesGenerated = false; // Generating the tables once is much more bundle size-efficient than shipping them in the bundle (entropy ftw) const generateAesTables = () => { const sbox = new Uint8Array(256); const log = new Uint8Array(256); const pow = new Uint8Array(256); // 1. Generate GF(2^8) log/exp tables // Primitive polynomial: x^8 + x^4 + x^3 + x + 1 (0x11B) for (let i = 0, p = 1; i < 256; i++) { pow[i] = p; log[p] = i; p = p ^ (p << 1) ^ (p & 0x80 ? 0x11B : 0); } // Helper: GF(2^8) multiplication const mul = (a: number, b: number) => (a && b) ? pow[(log[a]! + log[b]!) % 255]! : 0; // 2. Generate S-Box and Inverse S-Box sbox[0] = 0x63; // Special case for 0 // Loop for inverse (using log/exp) and Affine Transform for (let i = 1; i < 256; i++) { const x = pow[255 - log[i]!]!; // Multiplicative inverse let s = x ^ (x << 1) ^ (x << 2) ^ (x << 3) ^ (x << 4); s = (s >>> 8) ^ (s & 0xFF) ^ 0x63; // Affine transform sbox[i] = s; } // 3. Fill Tables for (let i = 0; i < 256; i++) { const s = sbox[i]!; // Forward S-Box value const is = sbox.indexOf(i); // Inverse S-Box value // Te4: Forward S-Box packed Te4[i] = (s << 24) | (s << 16) | (s << 8) | s; // Td4: Inverse S-Box packed Td4[i] = (is << 24) | (is << 16) | (is << 8) | is; // Td0-Td3: Inverse MixColumns applied to Inverse S-Box // Coefficients: 0x0E, 0x09, 0x0D, 0x0B (Order specific to Td0 structure) const b0 = mul(is, 0x0E); const b1 = mul(is, 0x09); const b2 = mul(is, 0x0D); const b3 = mul(is, 0x0B); const w = (b0 << 24) | (b1 << 16) | (b2 << 8) | b3; Td0[i] = w; Td1[i] = (w >>> 8) | (w << 24); // Rotate right 8 Td2[i] = (w >>> 16) | (w << 16); // Rotate right 16 Td3[i] = (w >>> 24) | (w << 8); // Rotate right 24 } // 4. Generate Rcon let r = 1; for (let i = 0; i < 10; i++) { rcon[i] = r << 24; r = (r << 1) ^ (r & 0x80 ? 0x11B : 0); } tablesGenerated = true; }; export type Aes128CbcContextInit = { key: Uint8Array; iv: Uint8Array; }; /** A context for doing AES-128-CBC operations. Better than the Web Crypto API since we can stream it. */ export class Aes128CbcContext { roundkey = new Uint32Array(44); iv = new Uint32Array(AES_128_BLOCK_SIZE / Uint32Array.BYTES_PER_ELEMENT); in = new Uint8Array(AES_128_BLOCK_SIZE); out = new Uint8Array(AES_128_BLOCK_SIZE); inView = new DataView(this.in.buffer); outView = new DataView(this.out.buffer); init({ key, iv }: Aes128CbcContextInit) { assert(key.byteLength === 16); assert(iv.byteLength === 16); if (!tablesGenerated) { generateAesTables(); } const keyView = new DataView(key.buffer, key.byteOffset, key.byteLength); const ivView = new DataView(iv.buffer, iv.byteOffset, iv.byteLength); this.roundkey[0] = keyView.getUint32(0, false); this.roundkey[1] = keyView.getUint32(4, false); this.roundkey[2] = keyView.getUint32(8, false); this.roundkey[3] = keyView.getUint32(12, false); this.iv[0] = ivView.getUint32(0, false); this.iv[1] = ivView.getUint32(4, false); this.iv[2] = ivView.getUint32(8, false); this.iv[3] = ivView.getUint32(12, false); for (let index = 4; index < 44; index += 4) { const temp = this.roundkey[index - 1]!; this.roundkey[index] = this.roundkey[index - 4]! ^ (Te4[(temp >>> 16) & 0xff]! & 0xff000000) ^ (Te4[(temp >>> 8) & 0xff]! & 0x00ff0000) ^ (Te4[(temp >>> 0) & 0xff]! & 0x0000ff00) ^ (Te4[(temp >>> 24) & 0xff]! & 0x000000ff) ^ rcon[(index / 4) - 1]!; this.roundkey[index + 1] = this.roundkey[index - 3]! ^ this.roundkey[index]!; this.roundkey[index + 2] = this.roundkey[index - 2]! ^ this.roundkey[index + 1]!; this.roundkey[index + 3] = this.roundkey[index - 1]! ^ this.roundkey[index + 2]!; } // Invert the order of the round keys for (let i = 0, j = 40; i < j; i += 4, j -= 4) { for (let k = 0; k < 4; k++) { const temp = this.roundkey[i + k]!; this.roundkey[i + k] = this.roundkey[j + k]!; this.roundkey[j + k] = temp; } } // Apply Inverse MixColumn transform to all round keys except first and last for (let index = 4; index < 40; index += 4) { for (let k = 0; k < 4; k++) { const rk = this.roundkey[index + k]!; this.roundkey[index + k] = Td0[Te4[(rk >>> 24) & 0xff]! & 0xff]! ^ Td1[Te4[(rk >>> 16) & 0xff]! & 0xff]! ^ Td2[Te4[(rk >>> 8) & 0xff]! & 0xff]! ^ Td3[Te4[(rk >>> 0) & 0xff]! & 0xff]!; } } } decrypt() { let s0 = this.inView.getUint32(0, false) ^ this.roundkey[0]!; let s1 = this.inView.getUint32(4, false) ^ this.roundkey[1]!; let s2 = this.inView.getUint32(8, false) ^ this.roundkey[2]!; let s3 = this.inView.getUint32(12, false) ^ this.roundkey[3]!; // Store input for CBC XOR later const temp0 = this.inView.getUint32(0, false); const temp1 = this.inView.getUint32(4, false); const temp2 = this.inView.getUint32(8, false); const temp3 = this.inView.getUint32(12, false); let t0, t1, t2, t3; // Rounds 1-9 for (let round = 1; round < 10; round++) { const offset = round * 4; t0 = Td0[s0 >>> 24]! ^ Td1[(s3 >>> 16) & 0xff]! ^ Td2[(s2 >>> 8) & 0xff]! ^ Td3[s1 & 0xff]! ^ this.roundkey[offset]!; t1 = Td0[s1 >>> 24]! ^ Td1[(s0 >>> 16) & 0xff]! ^ Td2[(s3 >>> 8) & 0xff]! ^ Td3[s2 & 0xff]! ^ this.roundkey[offset + 1]!; t2 = Td0[s2 >>> 24]! ^ Td1[(s1 >>> 16) & 0xff]! ^ Td2[(s0 >>> 8) & 0xff]! ^ Td3[s3 & 0xff]! ^ this.roundkey[offset + 2]!; t3 = Td0[s3 >>> 24]! ^ Td1[(s2 >>> 16) & 0xff]! ^ Td2[(s1 >>> 8) & 0xff]! ^ Td3[s0 & 0xff]! ^ this.roundkey[offset + 3]!; s0 = t0; s1 = t1; s2 = t2; s3 = t3; } // Final Round (10) const f0 = (Td4[(s0 >>> 24) & 0xff]! & 0xff000000) ^ (Td4[(s3 >>> 16) & 0xff]! & 0x00ff0000) ^ (Td4[(s2 >>> 8) & 0xff]! & 0x0000ff00) ^ (Td4[(s1 >>> 0) & 0xff]! & 0x000000ff) ^ this.roundkey[40]!; const f1 = (Td4[(s1 >>> 24) & 0xff]! & 0xff000000) ^ (Td4[(s0 >>> 16) & 0xff]! & 0x00ff0000) ^ (Td4[(s3 >>> 8) & 0xff]! & 0x0000ff00) ^ (Td4[(s2 >>> 0) & 0xff]! & 0x000000ff) ^ this.roundkey[41]!; const f2 = (Td4[(s2 >>> 24) & 0xff]! & 0xff000000) ^ (Td4[(s1 >>> 16) & 0xff]! & 0x00ff0000) ^ (Td4[(s0 >>> 8) & 0xff]! & 0x0000ff00) ^ (Td4[(s3 >>> 0) & 0xff]! & 0x000000ff) ^ this.roundkey[42]!; const f3 = (Td4[(s3 >>> 24) & 0xff]! & 0xff000000) ^ (Td4[(s2 >>> 16) & 0xff]! & 0x00ff0000) ^ (Td4[(s1 >>> 8) & 0xff]! & 0x0000ff00) ^ (Td4[(s0 >>> 0) & 0xff]! & 0x000000ff) ^ this.roundkey[43]!; // CBC XOR and output this.outView.setUint32(0, f0 ^ this.iv[0]!, false); this.outView.setUint32(4, f1 ^ this.iv[1]!, false); this.outView.setUint32(8, f2 ^ this.iv[2]!, false); this.outView.setUint32(12, f3 ^ this.iv[3]!, false); // Update IV for next block this.iv[0] = temp0; this.iv[1] = temp1; this.iv[2] = temp2; this.iv[3] = temp3; } } export const createAes128CbcDecryptStream = ( reader: Reader, getInit: () => MaybePromise, close: () => unknown, ) => { let initted = false; let pos = 0; const CHUNK_SIZE = 2 ** 16; const BLOCK_SIZE = 16; const aesContext = new Aes128CbcContext(); return new ReadableStream({ pull: async (controller) => { if (!initted) { aesContext.init(await getInit()); initted = true; } const requestedLength = CHUNK_SIZE + BLOCK_SIZE; let nextSlice = reader.requestSliceRange(pos, 0, requestedLength); if (nextSlice instanceof Promise) nextSlice = await nextSlice; if (!nextSlice || nextSlice.length === 0) { // Due to padding, this should never happen throw new Error('Invalid ciphertext.'); } const sliceLength = nextSlice.length; if (sliceLength % 16 !== 0) { throw new Error('Invalid ciphertext.'); } const bytesToRead = sliceLength === requestedLength ? sliceLength - BLOCK_SIZE // Don't read the last block : sliceLength; const input = readBytes(nextSlice, bytesToRead); const output = new Uint8Array(bytesToRead); for (let i = 0; i < bytesToRead; i += 16) { aesContext.in.set(input.subarray(i, i + 16)); aesContext.decrypt(); output.set(aesContext.out, i); } if (bytesToRead < sliceLength) { controller.enqueue(output); pos += bytesToRead; } else { // This is the last chunk const paddingLength = output[bytesToRead - 1]!; if (paddingLength === 0 || paddingLength > 16) { throw new Error('Invalid PKCS#7 padding. Incorrect key or corrupted data.'); } const trimmedOutput = output.subarray(0, bytesToRead - paddingLength); // PKCS#7 padding controller.enqueue(trimmedOutput); controller.close(); close(); } }, cancel: () => { close(); }, }); }; ===== src/input-format.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Demuxer } from './demuxer'; import { Input } from './input'; import { IsobmffDemuxer } from './isobmff/isobmff-demuxer'; import type { PsshBox } from './isobmff/isobmff-misc'; import { EBMLId, MAX_HEADER_SIZE, MIN_HEADER_SIZE, readAsciiString, readElementHeader, readElementSize, readUnsignedInt, readVarIntSize, } from './matroska/ebml'; import { MatroskaDemuxer } from './matroska/matroska-demuxer'; import { Mp3Demuxer } from './mp3/mp3-demuxer'; import { MP3_FRAME_HEADER_SIZE, getXingOffset, INFO, XING } from '../shared/mp3-misc'; import { ID3_V2_HEADER_SIZE, readId3V2Header } from './id3'; import { readNextMp3FrameHeader } from './mp3/mp3-reader'; import { OggDemuxer } from './ogg/ogg-demuxer'; import { WaveDemuxer } from './wave/wave-demuxer'; import { MAX_ADTS_FRAME_HEADER_SIZE, MIN_ADTS_FRAME_HEADER_SIZE, readAdtsFrameHeader } from './adts/adts-reader'; import { AdtsDemuxer } from './adts/adts-demuxer'; import { readAscii, readBytes, readU32Be } from './reader'; import { FlacDemuxer } from './flac/flac-demuxer'; import { MpegTsDemuxer } from './mpeg-ts/mpeg-ts-demuxer'; import { TS_PACKET_SIZE } from './mpeg-ts/mpeg-ts-misc'; import { HlsDemuxer } from './hls/hls-demuxer'; import { HLS_MIME_TYPE } from './hls/hls-misc'; import { PathedSource } from './source'; import { MaybePromise } from './misc'; /** * Base class representing an input media file format. * @group Input formats * @public */ export abstract class InputFormat { /** @internal */ abstract _canReadInput(input: Input): Promise; /** @internal */ abstract _createDemuxer(input: Input): Demuxer; /** Returns the name of the input format. */ abstract get name(): string; /** Returns the typical base MIME type of the input format. */ abstract get mimeType(): string; /** * Provided for tree-shakable checking. * @internal */ _isIsobmff = false; } /** * Format representing files compatible with the ISO base media file format (ISOBMFF), like MP4 or MOV files. * * This format can make use of {@link InputOptions.initInput}. When the file contents are fragmented but no track * initialization info is provided (no `moov` atom), then it must be provided via `initInput`. * * @group Input formats * @public */ export abstract class IsobmffInputFormat extends InputFormat { /** @internal */ protected async _getMajorBrand(input: Input) { let slice = input._reader.requestSlice(0, 12); if (slice instanceof Promise) slice = await slice; if (!slice) return null; slice.skip(4); const fourCc = readAscii(slice, 4); if ( fourCc !== 'ftyp' && fourCc !== 'styp' // Segment ) { return null; } return readAscii(slice, 4); } /** @internal */ _createDemuxer(input: Input) { return new IsobmffDemuxer(input); } /** @internal */ override _isIsobmff = true; } /** * MPEG-4 Part 14 (MP4) file format. * * Do not instantiate this class; use the {@link MP4} singleton instead. * * @group Input formats * @public */ export class Mp4InputFormat extends IsobmffInputFormat { /** @internal */ async _canReadInput(input: Input) { const majorBrand = await this._getMajorBrand(input); if (majorBrand !== null) { return majorBrand !== 'qt '; } let slice = input._reader.requestSlice(4, 4); if (slice instanceof Promise) slice = await slice; if (!slice) return false; const fourCc = readAscii(slice, 4); return fourCc === 'moof' || fourCc === 'sidx'; // Seen in HLS for example } get name() { return 'MP4'; } get mimeType() { return 'video/mp4'; } } /** * QuickTime File Format (QTFF), often called MOV. * * Do not instantiate this class; use the {@link QTFF} singleton instead. * * @group Input formats * @public */ export class QuickTimeInputFormat extends IsobmffInputFormat { /** @internal */ async _canReadInput(input: Input) { const majorBrand = await this._getMajorBrand(input); return majorBrand === 'qt '; } get name() { return 'QuickTime File Format'; } get mimeType() { return 'video/quicktime'; } } /** * Matroska file format. * * Do not instantiate this class; use the {@link MATROSKA} singleton instead. * * @group Input formats * @public */ export class MatroskaInputFormat extends InputFormat { /** @internal */ protected async isSupportedEBMLOfDocType(input: Input, desiredDocType: string) { let headerSlice = input._reader.requestSlice(0, MAX_HEADER_SIZE); if (headerSlice instanceof Promise) headerSlice = await headerSlice; if (!headerSlice) return false; const varIntSize = readVarIntSize(headerSlice); if (varIntSize === null) { return false; } if (varIntSize < 1 || varIntSize > 8) { return false; } const id = readUnsignedInt(headerSlice, varIntSize); if (id !== EBMLId.EBML) { return false; } const dataSize = readElementSize(headerSlice); if (typeof dataSize !== 'number') { return false; // Miss me with that shit } let dataSlice = input._reader.requestSlice(headerSlice.filePos, dataSize); if (dataSlice instanceof Promise) dataSlice = await dataSlice; if (!dataSlice) return false; const startPos = headerSlice.filePos; while (dataSlice.filePos <= startPos + dataSize - MIN_HEADER_SIZE) { const header = readElementHeader(dataSlice); if (!header) break; const { id, size } = header; const dataStartPos = dataSlice.filePos; if (size === undefined) return false; switch (id) { case EBMLId.EBMLVersion: { const ebmlVersion = readUnsignedInt(dataSlice, size); if (ebmlVersion !== 1) { return false; } }; break; case EBMLId.EBMLReadVersion: { const ebmlReadVersion = readUnsignedInt(dataSlice, size); if (ebmlReadVersion !== 1) { return false; } }; break; case EBMLId.DocType: { const docType = readAsciiString(dataSlice, size); if (docType !== desiredDocType) { return false; } }; break; case EBMLId.DocTypeVersion: { const docTypeVersion = readUnsignedInt(dataSlice, size); if (docTypeVersion > 4) { // Support up to Matroska v4 return false; } }; break; } dataSlice.filePos = dataStartPos + size; } return true; } /** @internal */ _canReadInput(input: Input) { return this.isSupportedEBMLOfDocType(input, 'matroska'); } /** @internal */ _createDemuxer(input: Input) { return new MatroskaDemuxer(input); } get name() { return 'Matroska'; } get mimeType() { return 'video/x-matroska'; } } /** * WebM file format, based on Matroska. * * Do not instantiate this class; use the {@link WEBM} singleton instead. * * @group Input formats * @public */ export class WebMInputFormat extends MatroskaInputFormat { /** @internal */ override _canReadInput(input: Input) { return this.isSupportedEBMLOfDocType(input, 'webm'); } override get name() { return 'WebM'; } override get mimeType() { return 'video/webm'; } } /** * MP3 file format. * * Do not instantiate this class; use the {@link MP3} singleton instead. * * @group Input formats * @public */ export class Mp3InputFormat extends InputFormat { /** @internal */ async _canReadInput(input: Input) { let currentPos = 0; while (true) { let slice = input._reader.requestSlice(currentPos, ID3_V2_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const id3V2Header = readId3V2Header(slice); if (!id3V2Header) { break; } currentPos = slice.filePos + id3V2Header.size; } const firstResult = await readNextMp3FrameHeader(input._reader, currentPos, currentPos + 4096); if (!firstResult) { return false; } const firstHeader = firstResult.header; const xingOffset = getXingOffset(firstHeader.mpegVersionId, firstHeader.channel); let slice = input._reader.requestSlice(firstResult.startPos + xingOffset, 4); if (slice instanceof Promise) slice = await slice; if (!slice) return false; const word = readU32Be(slice); const isXing = word === XING || word === INFO; if (isXing) { // Gotta be MP3 return true; } currentPos = firstResult.startPos + firstResult.header.totalSize; // Fine, we found one frame header, but we're still not entirely sure this is MP3. Let's check if we can find // another header right after it: const secondResult = await readNextMp3FrameHeader( input._reader, currentPos, currentPos + MP3_FRAME_HEADER_SIZE, ); if (!secondResult) { return false; } const secondHeader = secondResult.header; // In a well-formed MP3 file, we'd expect these two frames to share some similarities: if (firstHeader.channel !== secondHeader.channel || firstHeader.sampleRate !== secondHeader.sampleRate) { return false; } // We have found two matching consecutive MP3 frames, a strong indicator that this is an MP3 file return true; } /** @internal */ _createDemuxer(input: Input) { return new Mp3Demuxer(input); } get name() { return 'MP3'; } get mimeType() { return 'audio/mpeg'; } } /** * WAVE file format, based on RIFF. * * Do not instantiate this class; use the {@link WAVE} singleton instead. * * @group Input formats * @public */ export class WaveInputFormat extends InputFormat { /** @internal */ async _canReadInput(input: Input) { let slice = input._reader.requestSlice(0, 12); if (slice instanceof Promise) slice = await slice; if (!slice) return false; const riffType = readAscii(slice, 4); if (riffType !== 'RIFF' && riffType !== 'RIFX' && riffType !== 'RF64') { return false; } slice.skip(4); const format = readAscii(slice, 4); return format === 'WAVE'; } /** @internal */ _createDemuxer(input: Input) { return new WaveDemuxer(input); } get name() { return 'WAVE'; } get mimeType() { return 'audio/wav'; } } /** * Ogg file format. * * Do not instantiate this class; use the {@link OGG} singleton instead. * * @group Input formats * @public */ export class OggInputFormat extends InputFormat { /** @internal */ async _canReadInput(input: Input) { let slice = input._reader.requestSlice(0, 4); if (slice instanceof Promise) slice = await slice; if (!slice) return false; return readAscii(slice, 4) === 'OggS'; } /** @internal */ _createDemuxer(input: Input) { return new OggDemuxer(input); } get name() { return 'Ogg'; } get mimeType() { return 'application/ogg'; } } /** * FLAC file format. * * Do not instantiate this class; use the {@link FLAC} singleton instead. * * @group Input formats * @public */ export class FlacInputFormat extends InputFormat { /** @internal */ async _canReadInput(input: Input) { let currentPos = 0; // There might be ID3v2 headers at the start, skip 'em while (true) { let slice = input._reader.requestSlice(currentPos, ID3_V2_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const id3V2Header = readId3V2Header(slice); if (!id3V2Header) { break; } currentPos = slice.filePos + id3V2Header.size; } let slice = input._reader.requestSlice(currentPos, 4); if (slice instanceof Promise) slice = await slice; if (!slice) return false; return readAscii(slice, 4) === 'fLaC'; } get name() { return 'FLAC'; } get mimeType() { return 'audio/flac'; } /** @internal */ _createDemuxer(input: Input): Demuxer { return new FlacDemuxer(input); } } /** * ADTS file format. * * Do not instantiate this class; use the {@link ADTS} singleton instead. * * @group Input formats * @public */ export class AdtsInputFormat extends InputFormat { /** @internal */ async _canReadInput(input: Input) { let currentPos = 0; while (true) { let slice = input._reader.requestSlice(currentPos, ID3_V2_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const id3V2Header = readId3V2Header(slice); if (!id3V2Header) { break; } currentPos = slice.filePos + id3V2Header.size; } let slice = input._reader.requestSliceRange( currentPos, MIN_ADTS_FRAME_HEADER_SIZE, MAX_ADTS_FRAME_HEADER_SIZE, ); if (slice instanceof Promise) slice = await slice; if (!slice) return false; const firstHeader = readAdtsFrameHeader(slice); if (!firstHeader) { return false; } currentPos += firstHeader.frameLength; slice = input._reader.requestSliceRange( currentPos, MIN_ADTS_FRAME_HEADER_SIZE, MAX_ADTS_FRAME_HEADER_SIZE, ); if (slice instanceof Promise) slice = await slice; if (!slice) return false; const secondHeader = readAdtsFrameHeader(slice); if (!secondHeader) { return false; } return firstHeader.objectType === secondHeader.objectType && firstHeader.samplingFrequencyIndex === secondHeader.samplingFrequencyIndex && firstHeader.channelConfiguration === secondHeader.channelConfiguration; } /** @internal */ _createDemuxer(input: Input) { return new AdtsDemuxer(input); } get name() { return 'ADTS'; } get mimeType() { return 'audio/aac'; } } /** * MPEG Transport Stream (MPEG-TS) file format. * * This format can make use of {@link InputOptions.initInput} to initialize track information even when no * initialization information is provided for the track, for example because it has no key frames. In this case, tracks * are matched to each other based on their PID. * * Do not instantiate this class; use the {@link MPEG_TS} singleton instead. * * @group Input formats * @public */ export class MpegTsInputFormat extends InputFormat { /** @internal */ async _canReadInput(input: Input) { const lengthToCheck = TS_PACKET_SIZE + 16 + 1; let slice = input._reader.requestSlice(0, lengthToCheck); if (slice instanceof Promise) slice = await slice; if (!slice) return false; const bytes = readBytes(slice, lengthToCheck); if (bytes[0] === 0x47 && bytes[TS_PACKET_SIZE] === 0x47) { // Regular MPEG-TS return true; } else if (bytes[0] === 0x47 && bytes[TS_PACKET_SIZE + 16] === 0x47) { // MPEG-TS with Forward Error Correction return true; } else if (bytes[4] === 0x47 && bytes[4 + TS_PACKET_SIZE + 4] === 0x47) { // MPEG-2-TS (DVHS) return true; } return false; } /** @internal */ _createDemuxer(input: Input) { return new MpegTsDemuxer(input); } get name() { return 'MPEG Transport Stream'; } get mimeType() { return 'video/MP2T'; } } /** * Media described using the HTTP Live Streaming (HLS) protocol, with playlists in the M3U8 format. * * Do not instantiate this class; use the {@link HLS} singleton instead. * * @group Input formats * @public */ export class HlsInputFormat extends InputFormat { /** @internal */ async _canReadInput(input: Input) { let slice = input._reader.requestSlice(0, 7); if (slice instanceof Promise) slice = await slice; if (!slice) return false; const isM3u8 = readAscii(slice, 7) === '#EXTM3U'; if (!isM3u8) { return false; } if (!(input._rootSource instanceof PathedSource)) { throw new TypeError('HLS inputs require `InputOptions.source` to be a PathedSource or a ref to one.'); } input._rootSource._usedForHls = true; return true; } /** @internal */ _createDemuxer(input: Input) { return new HlsDemuxer(input); } get name() { return 'HTTP Live Streaming (HLS)'; } get mimeType() { return HLS_MIME_TYPE; } } /** * MP4 input format singleton. * @group Input formats * @public */ export const MP4 = /* #__PURE__ */ new Mp4InputFormat(); /** * QuickTime File Format input format singleton. * @group Input formats * @public */ export const QTFF = /* #__PURE__ */ new QuickTimeInputFormat(); /** * Matroska input format singleton. * @group Input formats * @public */ export const MATROSKA = /* #__PURE__ */ new MatroskaInputFormat(); /** * WebM input format singleton. * @group Input formats * @public */ export const WEBM = /* #__PURE__ */ new WebMInputFormat(); /** * MP3 input format singleton. * @group Input formats * @public */ export const MP3 = /* #__PURE__ */ new Mp3InputFormat(); /** * WAVE input format singleton. * @group Input formats * @public */ export const WAVE = /* #__PURE__ */ new WaveInputFormat(); /** * Ogg input format singleton. * @group Input formats * @public */ export const OGG = /* #__PURE__ */ new OggInputFormat(); /** * ADTS input format singleton. * @group Input formats * @public */ export const ADTS = /* #__PURE__ */ new AdtsInputFormat(); /** * FLAC input format singleton. * @group Input formats * @public */ export const FLAC = /* #__PURE__ */ new FlacInputFormat(); /** * MPEG-TS input format singleton. * @group Input formats * @public */ export const MPEG_TS = /* #__PURE__ */ new MpegTsInputFormat(); /** * HLS input format singleton. * @group Input formats * @public */ export const HLS = /* #__PURE__ */ new HlsInputFormat(); /** * List of all input format singletons. If you don't need to support all input formats, you should specify the * formats individually for better tree shaking. * @group Input formats * @public */ export const ALL_FORMATS: InputFormat[] = [HLS, MP4, QTFF, MATROSKA, WEBM, WAVE, OGG, FLAC, MP3, ADTS, MPEG_TS]; /** * List of input formats required for playback of typical HLS manifests. Includes HLS itself as well as the typical * segment formats: MPEG Transport Stream (.ts), MP4 (CMAF), ADTS (.aac) and MP3. * @group Input formats * @public */ export const HLS_FORMATS: InputFormat[] = [HLS, MP4, QTFF, MP3, ADTS, MPEG_TS]; /** * Additional per-format configuration. * @group Input formats * @public */ export type InputFormatOptions = { /** ISOBMFF-specific configuration. */ isobmff?: IsobmffInputFormatOptions; /** HLS-specific configuration. */ hls?: HlsInputFormatOptions; }; /** * Additional ISOBMFF input configuration. * @group Input formats * @public */ export type IsobmffInputFormatOptions = { /** * A callback that gets invoked for each key ID required for sample content decryption. The key ID is provided as a * 32-character lowercase hexadecimal string. * * Must return or resolve to a 32-character hexadecimal string or a 16-byte `Uint8Array`. */ resolveKeyId?: (options: { /** The key ID that is to be resolved to a key. This is a 32-character lowercase hexadecimal string. */ keyId: string; /** * Protection System Specific Header (pssh) boxes that apply to this key ID. Can be used to obtain a * description key from a DRM license server. */ psshBoxes: PsshBox[]; }) => MaybePromise; /** @internal */ _suppressPsshParsing?: boolean; }; /** * Additional HLS input configuration. * @group Input formats * @public */ export type HlsInputFormatOptions = { /** * Whether, in the presence of `#EXT-X-PROGRAM-DATE-TIME` tags, to offset track and packet timestamps to be relative * to the Unix epoch. * * Defaults to `true`, meaning packet timestamps map directly to wall-clock time. This guarantees AV sync across * multiple tracks, even with gaps present. * * When you don't want this mapping, you can set this value to `false`. In addition to timestamps not being Unix * timestamps anymore, any gaps in the playlist are also naturally removed. When `false`, you can still access the * wall-clock Unix timestamps via {@link InputTrack.getUnixTimeForTimestamp}. */ offsetTimestampsByDateTime?: boolean; }; export const validateInputFormatOptions = (options: InputFormatOptions, prefix: string) => { if (!options || typeof options !== 'object') { throw new TypeError(`${prefix}, when provided, must be an object.`); } if (options.isobmff !== undefined) { if (!options.isobmff || typeof options.isobmff !== 'object') { throw new TypeError(`${prefix}.isobmff, when provided, must be an object.`); } if (options.isobmff.resolveKeyId !== undefined && typeof options.isobmff.resolveKeyId !== 'function') { throw new TypeError(`${prefix}.isobmff.resolveKeyId, when provided, must be a function.`); } } if (options.hls !== undefined) { if (!options.hls || typeof options.hls !== 'object') { throw new TypeError(`${prefix}.hls, when provided, must be an object.`); } if ( options.hls.offsetTimestampsByDateTime !== undefined && typeof options.hls.offsetTimestampsByDateTime !== 'boolean' ) { throw new TypeError(`${prefix}.hls.offsetTimestampsByDateTime, when provided, must be a boolean.`); } } }; ===== src/tsconfig.json ===== { "extends": "../tsconfig.json", "compilerOptions": { "outDir": "../dist/modules", "rootDir": "..", "declaration": true, "declarationMap": true, "stripInternal": true, "noEmit": false, "composite": true }, "include": [ "**/*", "../shared/**/*" ] } ===== src/index.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ /// /// import { Logging } from './logging'; const MEDIABUNNY_LOADED_SYMBOL = Symbol.for('mediabunny loaded'); if ((globalThis as Record)[MEDIABUNNY_LOADED_SYMBOL]) { Logging._error( '[WARNING]\nMediabunny was loaded twice.' + ' This will likely cause Mediabunny not to work correctly.' + ' Check if multiple dependencies are importing different versions of Mediabunny,' + ' or if something is being bundled incorrectly.', ); } (globalThis as Record)[MEDIABUNNY_LOADED_SYMBOL] = true; export { Output, type OutputOptions, OutputTrack, OutputVideoTrack, OutputAudioTrack, OutputSubtitleTrack, OutputTrackGroup, type BaseTrackMetadata, type VideoTrackMetadata, type AudioTrackMetadata, type SubtitleTrackMetadata, type OutputEvents, } from './output'; export { OutputFormat, AdtsOutputFormat, type AdtsOutputFormatOptions, CmafOutputFormat, type CmafOutputFormatOptions, FlacOutputFormat, type FlacOutputFormatOptions, HlsOutputFormat, type HlsOutputFormatOptions, type HlsOutputPlaylistInfo, type HlsOutputSegmentInfo, IsobmffOutputFormat, type IsobmffOutputFormatOptions, MkvOutputFormat, type MkvOutputFormatOptions, MovOutputFormat, Mp3OutputFormat, type Mp3OutputFormatOptions, Mp4OutputFormat, MpegTsOutputFormat, type MpegTsOutputFormatOptions, OggOutputFormat, type OggOutputFormatOptions, WavOutputFormat, type WavOutputFormatOptions, WebMOutputFormat, type WebMOutputFormatOptions, type InclusiveIntegerRange, type TrackCountLimits, } from './output-format'; export { MediaSource, VideoSource, AudioSource, SubtitleSource, AudioBufferSource, AudioSampleSource, CanvasSource, EncodedAudioPacketSource, EncodedVideoPacketSource, MediaStreamAudioTrackSource, type MediaStreamAudioTrackSourceOptions, MediaStreamVideoTrackSource, type MediaStreamVideoTrackSourceOptions, TextSubtitleSource, VideoSampleSource, } from './media-source'; export { type MediaCodec, type VideoCodec, type AudioCodec, type SubtitleCodec, VIDEO_CODECS, AUDIO_CODECS, PCM_AUDIO_CODECS, NON_PCM_AUDIO_CODECS, SUBTITLE_CODECS, } from './codec'; export { canDecode, canDecodeVideo, canDecodeAudio, getDecodableCodecs, getDecodableVideoCodecs, getDecodableAudioCodecs, } from './decode'; export { type VideoEncodingConfig, type VideoEncodingAdditionalOptions, type VideoTransformOptions, type AudioEncodingConfig, type AudioEncodingAdditionalOptions, type AudioTransformOptions, canEncode, canEncodeVideo, canEncodeAudio, canEncodeSubtitles, getEncodableCodecs, getEncodableVideoCodecs, getEncodableAudioCodecs, getEncodableSubtitleCodecs, getFirstEncodableVideoCodec, getFirstEncodableAudioCodec, getFirstEncodableSubtitleCodec, Quality, QUALITY_VERY_LOW, QUALITY_LOW, QUALITY_MEDIUM, QUALITY_HIGH, QUALITY_VERY_HIGH, } from './encode'; export { Target, type TargetEvents, type TargetRequest, AppendOnlyStreamTarget, BufferTarget, type BufferTargetOptions, FilePathTarget, type FilePathTargetOptions, NullTarget, PathedTarget, RangedTarget, StreamTarget, type StreamTargetOptions, type StreamTargetChunk, } from './target'; export { type AnyIterable, ConcurrentRunner, type DeepReadonly, EventEmitter, type EventListenerOptions, type FilePath, type MaybePromise, } from './misc'; export { Logging, LogLevel, type LoggingEvents, } from './logging'; export { type PsshBox, } from './isobmff/isobmff-misc'; export { type Rational, type Rectangle, type Rotation, type SetOptional, type SetRequired, } from './misc'; export { type TrackType, ALL_TRACK_TYPES, } from './output'; export { Source, type SourceEvents, SourceRef, type SourceRequest, BlobSource, type BlobSourceOptions, BufferSource, CustomPathedSource, CustomSource, type CustomSourceOptions, FilePathSource, type FilePathSourceOptions, PathedSource, // eslint-disable-next-line @typescript-eslint/no-deprecated StreamSource, // eslint-disable-next-line @typescript-eslint/no-deprecated type StreamSourceOptions, RangedSource, ReadableStreamSource, type ReadableStreamSourceOptions, UrlSource, type UrlSourceOptions, } from './source'; export { InputFormat, type InputFormatOptions, AdtsInputFormat, FlacInputFormat, IsobmffInputFormat, type IsobmffInputFormatOptions, HlsInputFormat, type HlsInputFormatOptions, MatroskaInputFormat, Mp3InputFormat, Mp4InputFormat, MpegTsInputFormat, OggInputFormat, QuickTimeInputFormat, WaveInputFormat, WebMInputFormat, ALL_FORMATS, HLS_FORMATS, ADTS, FLAC, HLS, MATROSKA, MP3, MP4, MPEG_TS, OGG, QTFF, WAVE, WEBM, } from './input-format'; export { Input, type InputOptions, type InputEvents, InputDisposedError, UnsupportedInputFormatError, } from './input'; export { type DurationMetadataRequestOptions, } from './demuxer'; export { InputTrack, InputVideoTrack, InputAudioTrack, type InputTrackQuery, type PacketStats, asc, desc, prefer, } from './input-track'; export { EncodedPacket, type EncodedPacketSideData, type PacketType, } from './packet'; export { AudioSample, type AudioSampleInit, type AudioSampleCopyToOptions, AudioSampleResource, VideoSample, type VideoSampleInit, type VideoSamplePixelFormat, VideoSampleColorSpace, VideoSampleResource, type VideoSampleTransformOptions, type VideoSampleTransformationDescription, type CropRectangle, VIDEO_SAMPLE_PIXEL_FORMATS, type VideoDataPlane, registerVideoSampleTransformer, } from './sample'; export { AudioBufferSink, AudioSampleSink, BaseMediaSampleSink, CanvasSink, type CanvasSinkOptions, EncodedPacketSink, type PacketRetrievalOptions, VideoSampleSink, type VideoSinkDecoderOptions, type WrappedAudioBuffer, type WrappedCanvas, } from './media-sink'; export { Conversion, type ConversionOptions, type ConversionVideoOptions, type ConversionAudioOptions, ConversionCanceledError, type DiscardedTrack, } from './conversion'; export { CustomVideoDecoder, CustomVideoEncoder, CustomAudioDecoder, CustomAudioEncoder, registerDecoder, registerEncoder, } from './custom-coder'; export { type MetadataTags, type AttachedImage, RichImageData, AttachedFile, type TrackDisposition, } from './metadata'; // 🐡🦔 ===== src/resample.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { assert } from './misc'; import { AudioSample } from './sample'; /** * Utility class to handle audio resampling, handling both sample rate resampling as well as channel up/downmixing. * The advantage over doing this manually rather than using OfflineAudioContext to do it for us is the artifact-free * handling of putting multiple resampled audio samples back to back, which produces flaky results using * OfflineAudioContext. */ export class AudioResampler { sourceSampleRate: number | null = null; targetSampleRate: number; sourceNumberOfChannels: number | null = null; targetNumberOfChannels: number; startTime: number | null = null; onSample: (sample: AudioSample) => Promise; bufferSizeInFrames: number; bufferSizeInSamples: number; outputBuffer: Float32Array; /** Start frame of current buffer */ bufferStartFrame = 0; /** The highest index written to in the current buffer */ maxWrittenFrame: number | null = null; channelMixer!: (sourceData: Float32Array, sourceFrameIndex: number, targetChannelIndex: number) => number; tempSourceBuffer!: Float32Array; constructor(options: { targetSampleRate: number; targetNumberOfChannels: number; onSample: (sample: AudioSample) => Promise; }) { this.targetSampleRate = options.targetSampleRate; this.targetNumberOfChannels = options.targetNumberOfChannels; this.onSample = options.onSample; this.bufferSizeInFrames = Math.floor(this.targetSampleRate * 5.0); // 5 seconds this.bufferSizeInSamples = this.bufferSizeInFrames * this.targetNumberOfChannels; this.outputBuffer = new Float32Array(this.bufferSizeInSamples); } /** * Sets up the channel mixer to handle up/downmixing in the case where input and output channel counts don't match. */ doChannelMixerSetup(): void { assert(this.sourceNumberOfChannels !== null); const sourceNum = this.sourceNumberOfChannels; const targetNum = this.targetNumberOfChannels; // Logic taken from // https://developer.mozilla.org/en-US/docs/Web/API/Web_Audio_API/Basic_concepts_behind_Web_Audio_API // Most of the mapping functions are branchless. if (sourceNum === 1 && targetNum === 2) { // Mono to Stereo: M -> L, M -> R this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number) => { return sourceData[sourceFrameIndex * sourceNum]!; }; } else if (sourceNum === 1 && targetNum === 4) { // Mono to Quad: M -> L, M -> R, 0 -> SL, 0 -> SR this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number, targetChannelIndex: number) => { return sourceData[sourceFrameIndex * sourceNum]! * +(targetChannelIndex < 2); }; } else if (sourceNum === 1 && targetNum === 6) { // Mono to 5.1: 0 -> L, 0 -> R, M -> C, 0 -> LFE, 0 -> SL, 0 -> SR this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number, targetChannelIndex: number) => { return sourceData[sourceFrameIndex * sourceNum]! * +(targetChannelIndex === 2); }; } else if (sourceNum === 2 && targetNum === 1) { // Stereo to Mono: 0.5 * (L + R) this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number) => { const baseIdx = sourceFrameIndex * sourceNum; return 0.5 * (sourceData[baseIdx]! + sourceData[baseIdx + 1]!); }; } else if (sourceNum === 2 && targetNum === 4) { // Stereo to Quad: L -> L, R -> R, 0 -> SL, 0 -> SR this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number, targetChannelIndex: number) => { return sourceData[sourceFrameIndex * sourceNum + targetChannelIndex]! * +(targetChannelIndex < 2); }; } else if (sourceNum === 2 && targetNum === 6) { // Stereo to 5.1: L -> L, R -> R, 0 -> C, 0 -> LFE, 0 -> SL, 0 -> SR this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number, targetChannelIndex: number) => { return sourceData[sourceFrameIndex * sourceNum + targetChannelIndex]! * +(targetChannelIndex < 2); }; } else if (sourceNum === 4 && targetNum === 1) { // Quad to Mono: 0.25 * (L + R + SL + SR) this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number) => { const baseIdx = sourceFrameIndex * sourceNum; return 0.25 * ( sourceData[baseIdx]! + sourceData[baseIdx + 1]! + sourceData[baseIdx + 2]! + sourceData[baseIdx + 3]! ); }; } else if (sourceNum === 4 && targetNum === 2) { // Quad to Stereo: 0.5 * (L + SL), 0.5 * (R + SR) this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number, targetChannelIndex: number) => { const baseIdx = sourceFrameIndex * sourceNum; return 0.5 * ( sourceData[baseIdx + targetChannelIndex]! + sourceData[baseIdx + targetChannelIndex + 2]! ); }; } else if (sourceNum === 4 && targetNum === 6) { // Quad to 5.1: L -> L, R -> R, 0 -> C, 0 -> LFE, SL -> SL, SR -> SR this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number, targetChannelIndex: number) => { const baseIdx = sourceFrameIndex * sourceNum; // It's a bit harder to do this one branchlessly if (targetChannelIndex < 2) return sourceData[baseIdx + targetChannelIndex]!; // L, R if (targetChannelIndex === 2 || targetChannelIndex === 3) return 0; // C, LFE return sourceData[baseIdx + targetChannelIndex - 2]!; // SL, SR }; } else if (sourceNum === 6 && targetNum === 1) { // 5.1 to Mono: sqrt(1/2) * (L + R) + C + 0.5 * (SL + SR) this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number) => { const baseIdx = sourceFrameIndex * sourceNum; return Math.SQRT1_2 * (sourceData[baseIdx]! + sourceData[baseIdx + 1]!) + sourceData[baseIdx + 2]! + 0.5 * (sourceData[baseIdx + 4]! + sourceData[baseIdx + 5]!); }; } else if (sourceNum === 6 && targetNum === 2) { // 5.1 to Stereo: L + sqrt(1/2) * (C + SL), R + sqrt(1/2) * (C + SR) this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number, targetChannelIndex: number) => { const baseIdx = sourceFrameIndex * sourceNum; return sourceData[baseIdx + targetChannelIndex]! + Math.SQRT1_2 * (sourceData[baseIdx + 2]! + sourceData[baseIdx + targetChannelIndex + 4]!); }; } else if (sourceNum === 6 && targetNum === 4) { // 5.1 to Quad: L + sqrt(1/2) * C, R + sqrt(1/2) * C, SL, SR this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number, targetChannelIndex: number) => { const baseIdx = sourceFrameIndex * sourceNum; // It's a bit harder to do this one branchlessly if (targetChannelIndex < 2) { return sourceData[baseIdx + targetChannelIndex]! + Math.SQRT1_2 * sourceData[baseIdx + 2]!; } return sourceData[baseIdx + targetChannelIndex + 2]!; // SL, SR }; } else { // Discrete fallback: direct mapping with zero-fill or drop this.channelMixer = (sourceData: Float32Array, sourceFrameIndex: number, targetChannelIndex: number) => { return targetChannelIndex < sourceNum ? sourceData[sourceFrameIndex * sourceNum + targetChannelIndex]! : 0; }; } } ensureTempBufferSize(requiredSamples: number): void { let length = this.tempSourceBuffer.length; while (length < requiredSamples) { length *= 2; } if (length !== this.tempSourceBuffer.length) { const newBuffer = new Float32Array(length); newBuffer.set(this.tempSourceBuffer); this.tempSourceBuffer = newBuffer; } } async add(audioSample: AudioSample) { if (this.sourceSampleRate === null) { // This is the first sample, so let's init the missing data. Initting the sample rate from the decoded // sample is more reliable than using the file's metadata, because decoders are free to emit any sample rate // they see fit. this.sourceSampleRate = audioSample.sampleRate; this.sourceNumberOfChannels = audioSample.numberOfChannels; this.startTime = audioSample.timestamp; // Pre-allocate temporary buffer for source data this.tempSourceBuffer = new Float32Array(this.sourceSampleRate * this.sourceNumberOfChannels); this.doChannelMixerSetup(); } assert(this.startTime !== null); const requiredSamples = audioSample.numberOfFrames * audioSample.numberOfChannels; this.ensureTempBufferSize(requiredSamples); // Copy the audio data to the temp buffer const sourceDataSize = audioSample.allocationSize({ planeIndex: 0, format: 'f32' }); const sourceView = new Float32Array(this.tempSourceBuffer.buffer, 0, sourceDataSize / 4); audioSample.copyTo(sourceView, { planeIndex: 0, format: 'f32' }); const inputStartTime = audioSample.timestamp - this.startTime; const inputEndTime = inputStartTime + audioSample.duration; // Compute which output frames are affected by this sample const outputStartFrame = Math.floor(inputStartTime * this.targetSampleRate); const outputEndFrame = Math.ceil(inputEndTime * this.targetSampleRate); for (let outputFrame = outputStartFrame; outputFrame < outputEndFrame; outputFrame++) { if (outputFrame < this.bufferStartFrame) { continue; // Skip writes to the past } while (outputFrame >= this.bufferStartFrame + this.bufferSizeInFrames) { // The write is after the current buffer, so finalize it await this.finalizeCurrentBuffer(); this.bufferStartFrame += this.bufferSizeInFrames; } const bufferFrameIndex = outputFrame - this.bufferStartFrame; assert(bufferFrameIndex < this.bufferSizeInFrames); const outputTime = outputFrame / this.targetSampleRate; const inputTime = outputTime - inputStartTime; const sourcePosition = inputTime * this.sourceSampleRate; const sourceLowerFrame = Math.floor(sourcePosition); const sourceUpperFrame = Math.ceil(sourcePosition); const fraction = sourcePosition - sourceLowerFrame; // Process each output channel for (let targetChannel = 0; targetChannel < this.targetNumberOfChannels; targetChannel++) { let lowerSample = 0; let upperSample = 0; if (sourceLowerFrame >= 0 && sourceLowerFrame < audioSample.numberOfFrames) { lowerSample = this.channelMixer(sourceView, sourceLowerFrame, targetChannel); } if (sourceUpperFrame >= 0 && sourceUpperFrame < audioSample.numberOfFrames) { upperSample = this.channelMixer(sourceView, sourceUpperFrame, targetChannel); } // For resampling, we do naive linear interpolation to find the in-between sample. This produces // suboptimal results especially for downsampling (for which a low-pass filter would first need to be // applied), but AudioContext doesn't do this either, so, whatever, for now. const outputSample = lowerSample + fraction * (upperSample - lowerSample); // Write to output buffer (interleaved) const outputIndex = bufferFrameIndex * this.targetNumberOfChannels + targetChannel; this.outputBuffer[outputIndex]! += outputSample; // Add in case of overlapping samples } if (this.maxWrittenFrame === null) { this.maxWrittenFrame = bufferFrameIndex; } else { this.maxWrittenFrame = Math.max(this.maxWrittenFrame, bufferFrameIndex); } } } async finalizeCurrentBuffer() { if (this.maxWrittenFrame === null) { return; // Nothing to finalize } assert(this.startTime !== null); const samplesWritten = (this.maxWrittenFrame + 1) * this.targetNumberOfChannels; const outputData = new Float32Array(samplesWritten); outputData.set(this.outputBuffer.subarray(0, samplesWritten)); const audioSample = new AudioSample({ format: 'f32', sampleRate: this.targetSampleRate, numberOfChannels: this.targetNumberOfChannels, timestamp: this.startTime + this.bufferStartFrame / this.targetSampleRate, data: outputData, }); await this.onSample(audioSample); this.outputBuffer.fill(0); this.maxWrittenFrame = null; } finalize() { return this.finalizeCurrentBuffer(); } } ===== src/pcm.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ // https://github.com/dystopiancode/pcm-g711/blob/master/pcm-g711/g711.c export const toUlaw = (s16: number) => { const MULAW_MAX = 0x1FFF; const MULAW_BIAS = 33; let number = s16; let mask = 0x1000; let sign = 0; let position = 12; let lsb = 0; if (number < 0) { number = -number; sign = 0x80; } number += MULAW_BIAS; if (number > MULAW_MAX) { number = MULAW_MAX; } while ((number & mask) !== mask && position >= 5) { mask >>= 1; position--; } lsb = (number >> (position - 4)) & 0x0f; return ~(sign | ((position - 5) << 4) | lsb) & 0xFF; }; export const fromUlaw = (u8: number) => { const MULAW_BIAS = 33; let sign = 0; let position = 0; let number = ~u8; if (number & 0x80) { number &= ~(1 << 7); sign = -1; } position = ((number & 0xF0) >> 4) + 5; const decoded = ((1 << position) | ((number & 0x0F) << (position - 4)) | (1 << (position - 5))) - MULAW_BIAS; return (sign === 0) ? decoded : -decoded; }; export const toAlaw = (s16: number) => { const ALAW_MAX = 0xFFF; let mask = 0x800; let sign = 0; let position = 11; let lsb = 0; let number = s16; if (number < 0) { number = -number; sign = 0x80; } if (number > ALAW_MAX) { number = ALAW_MAX; } while ((number & mask) !== mask && position >= 5) { mask >>= 1; position--; } lsb = (number >> ((position === 4) ? 1 : (position - 4))) & 0x0f; return (sign | ((position - 4) << 4) | lsb) ^ 0x55; }; export const fromAlaw = (u8: number) => { let sign = 0x00; let position = 0; let number = u8 ^ 0x55; if (number & 0x80) { number &= ~(1 << 7); sign = -1; } position = ((number & 0xF0) >> 4) + 4; let decoded = 0; if (position !== 4) { decoded = ((1 << position) | ((number & 0x0F) << (position - 4)) | (1 << (position - 5))); } else { decoded = (number << 1) | 1; } return (sign === 0) ? decoded : -decoded; }; ===== src/wave/riff-writer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Writer } from '../writer'; export class RiffWriter { private helper = new Uint8Array(8); private helperView = new DataView(this.helper.buffer); constructor(private writer: Writer) {} writeU16(value: number) { this.helperView.setUint16(0, value, true); this.writer.write(this.helper.subarray(0, 2)); } writeU32(value: number) { this.helperView.setUint32(0, value, true); this.writer.write(this.helper.subarray(0, 4)); } writeU64(value: number) { this.helperView.setUint32(0, value, true); this.helperView.setUint32(4, Math.floor(value / 2 ** 32), true); this.writer.write(this.helper); } writeAscii(text: string) { this.writer.write(new TextEncoder().encode(text)); } } ===== src/wave/wave-demuxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AudioCodec } from '../codec'; import { Demuxer } from '../demuxer'; import { Input } from '../input'; import { InputAudioTrackBacking } from '../input-track'; import { PacketRetrievalOptions } from '../media-sink'; import { DEFAULT_TRACK_DISPOSITION, MetadataTags } from '../metadata'; import { assert, UNDETERMINED_LANGUAGE } from '../misc'; import { EncodedPacket, PLACEHOLDER_DATA } from '../packet'; import { readAscii, readBytes, Reader, readU16, readU32, readU64 } from '../reader'; import { ID3_V2_HEADER_SIZE, parseId3V2Tag, readId3V2Header } from '../id3'; export enum WaveFormat { PCM = 0x0001, IEEE_FLOAT = 0x0003, ALAW = 0x0006, MULAW = 0x0007, EXTENSIBLE = 0xFFFE, } export class WaveDemuxer extends Demuxer { reader: Reader; metadataPromise: Promise | null = null; dataStart = -1; dataSize = -1; audioInfo: { format: number; numberOfChannels: number; sampleRate: number; sampleSizeInBytes: number; blockSizeInBytes: number; } | null = null; trackBackings: WaveAudioTrackBacking[] = []; lastKnownPacketIndex = 0; metadataTags: MetadataTags = {}; constructor(input: Input) { super(input); this.reader = input._reader; } async readMetadata() { return this.metadataPromise ??= (async () => { let slice = this.reader.requestSlice(0, 12); if (slice instanceof Promise) slice = await slice; assert(slice); const riffType = readAscii(slice, 4); const littleEndian = riffType !== 'RIFX'; const isRf64 = riffType === 'RF64'; const outerChunkSize = readU32(slice, littleEndian); let totalFileSize = isRf64 ? this.reader.fileSize : Math.min(outerChunkSize + 8, this.reader.fileSize ?? Infinity); const format = readAscii(slice, 4); if (format !== 'WAVE') { throw new Error('Invalid WAVE file - wrong format'); } let chunksRead = 0; let dataChunkSize: number | null = null; let currentPos = slice.filePos; while (totalFileSize === null || currentPos < totalFileSize) { let slice = this.reader.requestSlice(currentPos, 8); if (slice instanceof Promise) slice = await slice; if (!slice) break; const chunkId = readAscii(slice, 4); const chunkSize = readU32(slice, littleEndian); const startPos = slice.filePos; if (isRf64 && chunksRead === 0 && chunkId !== 'ds64') { throw new Error('Invalid RF64 file: First chunk must be "ds64".'); } if (chunkId === 'fmt ') { await this.parseFmtChunk(startPos, chunkSize, littleEndian); } else if (chunkId === 'data') { dataChunkSize ??= chunkSize; this.dataStart = slice.filePos; this.dataSize = Math.min(dataChunkSize, (totalFileSize ?? Infinity) - this.dataStart); if (this.reader.fileSize === null) { break; // Stop once we hit the data chunk } } else if (chunkId === 'ds64') { // File and data chunk sizes are defined in here instead let ds64Slice = this.reader.requestSlice(startPos, chunkSize); if (ds64Slice instanceof Promise) ds64Slice = await ds64Slice; if (!ds64Slice) break; const riffChunkSize = readU64(ds64Slice, littleEndian); dataChunkSize = readU64(ds64Slice, littleEndian); totalFileSize = Math.min(riffChunkSize + 8, this.reader.fileSize ?? Infinity); } else if (chunkId === 'LIST') { await this.parseListChunk(startPos, chunkSize, littleEndian); } else if (chunkId === 'ID3 ' || chunkId === 'id3 ') { await this.parseId3Chunk(startPos, chunkSize); } currentPos = startPos + chunkSize + (chunkSize & 1); // Handle padding chunksRead++; } if (!this.audioInfo) { throw new Error('Invalid WAVE file - missing "fmt " chunk'); } if (this.dataStart === -1) { throw new Error('Invalid WAVE file - missing "data" chunk'); } const blockSize = this.audioInfo.blockSizeInBytes; this.dataSize = Math.floor(this.dataSize / blockSize) * blockSize; this.trackBackings.push(new WaveAudioTrackBacking(this)); })(); } private async parseFmtChunk(startPos: number, size: number, littleEndian: boolean) { let slice = this.reader.requestSlice(startPos, size); if (slice instanceof Promise) slice = await slice; if (!slice) return; // File too short let formatTag = readU16(slice, littleEndian); const numChannels = readU16(slice, littleEndian); const sampleRate = readU32(slice, littleEndian); slice.skip(4); // Bytes per second const blockAlign = readU16(slice, littleEndian); let bitsPerSample: number; if (size === 14) { // Plain WAVEFORMAT bitsPerSample = 8; } else { bitsPerSample = readU16(slice, littleEndian); } // Handle WAVEFORMATEXTENSIBLE if (size >= 18 && formatTag !== 0x0165) { const cbSize = readU16(slice, littleEndian); const remainingSize = size - 18; const extensionSize = Math.min(remainingSize, cbSize); if (extensionSize >= 22 && formatTag === WaveFormat.EXTENSIBLE) { // Parse WAVEFORMATEXTENSIBLE slice.skip(2 + 4); const subFormat = readBytes(slice, 16); // Get actual format from subFormat GUID formatTag = subFormat[0]! | (subFormat[1]! << 8); } } if (formatTag === WaveFormat.MULAW || formatTag === WaveFormat.ALAW) { bitsPerSample = 8; } this.audioInfo = { format: formatTag, numberOfChannels: numChannels, sampleRate, sampleSizeInBytes: Math.ceil(bitsPerSample / 8), blockSizeInBytes: blockAlign, }; } private async parseListChunk(startPos: number, size: number, littleEndian: boolean) { let slice = this.reader.requestSlice(startPos, size); if (slice instanceof Promise) slice = await slice; if (!slice) return; // File too short const infoType = readAscii(slice, 4); if (infoType !== 'INFO' && infoType !== 'INF0') { // exiftool.org claims INF0 can happen return; // Not an INFO chunk } let currentPos = slice.filePos; while (currentPos <= startPos + size - 8) { slice.filePos = currentPos; const chunkName = readAscii(slice, 4); const chunkSize = readU32(slice, littleEndian); const bytes = readBytes(slice, chunkSize); let stringLength = 0; for (let i = 0; i < bytes.length; i++) { if (bytes[i] === 0) { break; } stringLength++; } const value = String.fromCharCode(...bytes.subarray(0, stringLength)); this.metadataTags.raw ??= {}; this.metadataTags.raw[chunkName] = value; switch (chunkName) { case 'INAM': case 'TITL': { this.metadataTags.title ??= value; }; break; case 'TIT3': { this.metadataTags.description ??= value; }; break; case 'IART': { this.metadataTags.artist ??= value; }; break; case 'IPRD': { this.metadataTags.album ??= value; }; break; case 'IPRT': case 'ITRK': case 'TRCK': { const parts = value.split('/'); const trackNum = Number.parseInt(parts[0]!, 10); const tracksTotal = parts[1] && Number.parseInt(parts[1], 10); if (Number.isInteger(trackNum) && trackNum > 0) { this.metadataTags.trackNumber ??= trackNum; } if (tracksTotal && Number.isInteger(tracksTotal) && tracksTotal > 0) { this.metadataTags.tracksTotal ??= tracksTotal; } }; break; case 'ICRD': case 'IDIT': { const date = new Date(value); if (!Number.isNaN(date.getTime())) { this.metadataTags.date ??= date; } }; break; case 'YEAR': { const year = Number.parseInt(value, 10); if (Number.isInteger(year) && year > 0) { this.metadataTags.date ??= new Date(year, 0, 1); } }; break; case 'IGNR': case 'GENR': { this.metadataTags.genre ??= value; }; break; case 'ICMT': case 'CMNT': case 'COMM': { this.metadataTags.comment ??= value; }; break; } currentPos += 8 + chunkSize + (chunkSize & 1); // Handle padding } } private async parseId3Chunk(startPos: number, size: number) { // Parse ID3 tag embedded in WAV file (non-default, but used a lot in practice anyway) let slice = this.reader.requestSlice(startPos, size); if (slice instanceof Promise) slice = await slice; if (!slice) return; // File too short const id3V2Header = readId3V2Header(slice); if (id3V2Header) { // Clamp to the available data in case the ID3 header claims more than the WAV chunk provides // https://github.com/Vanilagy/mediabunny/issues/300 const availableSize = size - ID3_V2_HEADER_SIZE; id3V2Header.size = Math.min(id3V2Header.size, availableSize); if (id3V2Header.size > 0) { const contentSlice = slice.slice(startPos + ID3_V2_HEADER_SIZE, id3V2Header.size); parseId3V2Tag(contentSlice, id3V2Header, this.metadataTags); } } } getCodec(): AudioCodec | null { assert(this.audioInfo); if (this.audioInfo.format === WaveFormat.MULAW) { return 'ulaw'; } if (this.audioInfo.format === WaveFormat.ALAW) { return 'alaw'; } if (this.audioInfo.format === WaveFormat.PCM) { // All formats are little-endian if (this.audioInfo.sampleSizeInBytes === 1) { return 'pcm-u8'; } else if (this.audioInfo.sampleSizeInBytes === 2) { return 'pcm-s16'; } else if (this.audioInfo.sampleSizeInBytes === 3) { return 'pcm-s24'; } else if (this.audioInfo.sampleSizeInBytes === 4) { return 'pcm-s32'; } } if (this.audioInfo.format === WaveFormat.IEEE_FLOAT) { if (this.audioInfo.sampleSizeInBytes === 4) { return 'pcm-f32'; } } return null; } async getMimeType() { return 'audio/wav'; } async getTrackBackings() { await this.readMetadata(); return this.trackBackings; } async getMetadataTags() { await this.readMetadata(); return this.metadataTags; } } const PACKET_SIZE_IN_FRAMES = 2048; class WaveAudioTrackBacking implements InputAudioTrackBacking { constructor(public demuxer: WaveDemuxer) {} getType() { return 'audio' as const; } getId() { return 1; } getNumber() { return 1; } getCodec() { return this.demuxer.getCodec(); } getInternalCodecId() { assert(this.demuxer.audioInfo); return this.demuxer.audioInfo.format; } async getDecoderConfig(): Promise { const codec = this.demuxer.getCodec(); if (!codec) { return null; } assert(this.demuxer.audioInfo); return { codec, numberOfChannels: this.demuxer.audioInfo.numberOfChannels, sampleRate: this.demuxer.audioInfo.sampleRate, }; } getNumberOfChannels() { assert(this.demuxer.audioInfo); return this.demuxer.audioInfo.numberOfChannels; } getSampleRate() { assert(this.demuxer.audioInfo); return this.demuxer.audioInfo.sampleRate; } getTimeResolution() { assert(this.demuxer.audioInfo); return this.demuxer.audioInfo.sampleRate; } isRelativeToUnixEpoch() { return false; } getUnixTimeForTimestamp() { return null; } getPairingMask() { return 1n; } getBitrate() { return null; } getAverageBitrate() { return null; } async getDurationFromMetadata() { assert(this.demuxer.dataSize !== -1); return this.demuxer.dataSize / this.demuxer.audioInfo!.blockSizeInBytes / this.demuxer.audioInfo!.sampleRate; } async getLiveRefreshInterval() { return null; } getName() { return null; } getLanguageCode() { return UNDETERMINED_LANGUAGE; } getDisposition() { return { ...DEFAULT_TRACK_DISPOSITION, }; } private async getPacketAtIndex( packetIndex: number, options: PacketRetrievalOptions, ): Promise { assert(packetIndex >= 0); assert(this.demuxer.audioInfo); const startOffset = packetIndex * PACKET_SIZE_IN_FRAMES * this.demuxer.audioInfo.blockSizeInBytes; if (startOffset >= this.demuxer.dataSize) { return null; } const sizeInBytes = Math.min( PACKET_SIZE_IN_FRAMES * this.demuxer.audioInfo.blockSizeInBytes, this.demuxer.dataSize - startOffset, ); if (this.demuxer.reader.fileSize === null) { // If the file size is unknown, we weren't able to cap the dataSize in the init logic and we instead have to // rely on the headers telling us how large the file is. But, these might be wrong, so let's check if the // requested slice actually exists. let slice = this.demuxer.reader.requestSlice(this.demuxer.dataStart + startOffset, sizeInBytes); if (slice instanceof Promise) slice = await slice; if (!slice) { return null; } } let data: Uint8Array; if (options.metadataOnly) { data = PLACEHOLDER_DATA; } else { let slice = this.demuxer.reader.requestSlice(this.demuxer.dataStart + startOffset, sizeInBytes); if (slice instanceof Promise) slice = await slice; assert(slice); data = readBytes(slice, sizeInBytes); } const timestamp = packetIndex * PACKET_SIZE_IN_FRAMES / this.demuxer.audioInfo.sampleRate; const duration = sizeInBytes / this.demuxer.audioInfo.blockSizeInBytes / this.demuxer.audioInfo.sampleRate; this.demuxer.lastKnownPacketIndex = Math.max( packetIndex, this.demuxer.lastKnownPacketIndex, ); return new EncodedPacket( data, 'key', timestamp, duration, packetIndex, sizeInBytes, ); } getFirstPacket(options: PacketRetrievalOptions) { return this.getPacketAtIndex(0, options); } async getPacket(timestamp: number, options: PacketRetrievalOptions) { assert(this.demuxer.audioInfo); const packetIndex = Math.floor(Math.min( timestamp * this.demuxer.audioInfo.sampleRate / PACKET_SIZE_IN_FRAMES, (this.demuxer.dataSize - 1) / (PACKET_SIZE_IN_FRAMES * this.demuxer.audioInfo.blockSizeInBytes), )); if (packetIndex < 0) { return null; } const packet = await this.getPacketAtIndex(packetIndex, options); if (packet) { return packet; } if (packetIndex === 0) { return null; // Empty data chunk } assert(this.demuxer.reader.fileSize === null); // The file is shorter than we thought, meaning the packet we were looking for doesn't exist. So, let's find // the last packet by doing a sequential scan, instead. let currentPacket = await this.getPacketAtIndex(this.demuxer.lastKnownPacketIndex, options); while (currentPacket) { const nextPacket = await this.getNextPacket(currentPacket, options); if (!nextPacket) { break; } currentPacket = nextPacket; } return currentPacket; } getNextPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { assert(this.demuxer.audioInfo); const packetIndex = Math.round(packet.timestamp * this.demuxer.audioInfo.sampleRate / PACKET_SIZE_IN_FRAMES); return this.getPacketAtIndex(packetIndex + 1, options); } getKeyPacket(timestamp: number, options: PacketRetrievalOptions) { return this.getPacket(timestamp, options); } getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { return this.getNextPacket(packet, options); } } ===== src/wave/wave-muxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Muxer } from '../muxer'; import { Output, OutputAudioTrack } from '../output'; import { parsePcmCodec, PcmAudioCodec, validateAudioChunkMetadata } from '../codec'; import { WaveFormat } from './wave-demuxer'; import { RiffWriter } from './riff-writer'; import { Writer } from '../writer'; import { EncodedPacket } from '../packet'; import { WavOutputFormat } from '../output-format'; import { assert, assertNever, isIso88591Compatible, keyValueIterator } from '../misc'; import { MetadataTags, metadataTagsAreEmpty } from '../metadata'; import { Id3V2Writer } from '../id3'; import { Logging } from '../logging'; export class WaveMuxer extends Muxer { private format: WavOutputFormat; private isRf64: boolean; private writer!: Writer; private riffWriter!: RiffWriter; private headerWritten = false; private dataSize = 0; private sampleRate: number | null = null; private sampleCount = 0; private riffSizePos: number | null = null; private dataSizePos: number | null = null; private ds64RiffSizePos: number | null = null; private ds64DataSizePos: number | null = null; private ds64SampleCountPos: number | null = null; constructor(output: Output, format: WavOutputFormat) { super(output); this.format = format; this.isRf64 = !!format._options.large; } async start() { const release = await this.mutex.acquire(); this.writer = await this.output._getRootWriter(false); this.riffWriter = new RiffWriter(this.writer); // No writing needed here - we'll write the header with the first sample release(); } async getMimeType() { return 'audio/wav'; } async addEncodedVideoPacket() { throw new Error('WAVE does not support video.'); } async addEncodedAudioPacket( track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata, ) { const release = await this.mutex.acquire(); try { if (!this.headerWritten) { validateAudioChunkMetadata(meta); assert(meta); assert(meta.decoderConfig); this.writeHeader(track, meta.decoderConfig); this.sampleRate = meta.decoderConfig.sampleRate; this.headerWritten = true; } this.validateTimestamp(track, packet.timestamp, packet.type === 'key'); if (!this.isRf64 && this.writer.getPos() + packet.data.byteLength >= 2 ** 32) { throw new Error( 'Adding more audio data would exceed the maximum RIFF size of 4 GiB. To write larger files, use' + ' RF64 by setting `large: true` in the WavOutputFormatOptions.', ); } this.writer.write(packet.data); this.dataSize += packet.data.byteLength; this.sampleCount += Math.round(packet.duration * this.sampleRate!); await this.writer.flush(); } finally { release(); } } async addSubtitleCue() { throw new Error('WAVE does not support subtitles.'); } private writeHeader(track: OutputAudioTrack, config: AudioDecoderConfig) { if (this.format._options.onHeader) { this.writer.startTrackingWrites(); } let format: WaveFormat; const codec = track.source._codec; const pcmInfo = parsePcmCodec(codec as PcmAudioCodec); if (pcmInfo.dataType === 'ulaw') { format = WaveFormat.MULAW; } else if (pcmInfo.dataType === 'alaw') { format = WaveFormat.ALAW; } else if (pcmInfo.dataType === 'float') { format = WaveFormat.IEEE_FLOAT; } else { format = WaveFormat.PCM; } const channels = config.numberOfChannels; const sampleRate = config.sampleRate; const blockSize = pcmInfo.sampleSize * channels; // RIFF header this.riffWriter.writeAscii(this.isRf64 ? 'RF64' : 'RIFF'); if (this.isRf64) { this.riffWriter.writeU32(0xffffffff); // Not used in RF64 } else { this.riffSizePos = this.writer.getPos(); this.riffWriter.writeU32(0); // File size placeholder } this.riffWriter.writeAscii('WAVE'); if (this.isRf64) { this.riffWriter.writeAscii('ds64'); this.riffWriter.writeU32(28); // Chunk size this.ds64RiffSizePos = this.writer.getPos(); this.riffWriter.writeU64(0); // RIFF size placeholder this.ds64DataSizePos = this.writer.getPos(); this.riffWriter.writeU64(0); // Data size placeholder this.ds64SampleCountPos = this.writer.getPos(); this.riffWriter.writeU64(0); // Sample count placeholder this.riffWriter.writeU32(0); // Table length // Empty table } // fmt chunk this.riffWriter.writeAscii('fmt '); this.riffWriter.writeU32(16); // Chunk size this.riffWriter.writeU16(format); this.riffWriter.writeU16(channels); this.riffWriter.writeU32(sampleRate); this.riffWriter.writeU32(sampleRate * blockSize); // Bytes per second this.riffWriter.writeU16(blockSize); this.riffWriter.writeU16(8 * pcmInfo.sampleSize); // Metadata tags if (!metadataTagsAreEmpty(this.output._metadataTags)) { const metadataFormat = this.format._options.metadataFormat ?? 'info'; if (metadataFormat === 'info') { this.writeInfoChunk(this.output._metadataTags); } else if (metadataFormat === 'id3') { this.writeId3Chunk(this.output._metadataTags); } else { assertNever(metadataFormat); } } // data chunk this.riffWriter.writeAscii('data'); if (this.isRf64) { this.riffWriter.writeU32(0xffffffff); // Not used in RF64 } else { this.dataSizePos = this.writer.getPos(); this.riffWriter.writeU32(0); // Data size placeholder } if (this.format._options.onHeader) { const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onHeader(data, start); } } private writeInfoChunk(metadata: MetadataTags) { const startPos = this.writer.getPos(); this.riffWriter.writeAscii('LIST'); this.riffWriter.writeU32(0); // Size placeholder this.riffWriter.writeAscii('INFO'); const writtenTags = new Set(); const writeInfoTag = (tag: string, value: string) => { if (!isIso88591Compatible(value)) { // No Unicode supported here Logging._warn(`Didn't write tag '${tag}' because '${value}' is not ISO 8859-1-compatible.`); return; } const size = value.length + 1; // +1 for null terminator const bytes = new Uint8Array(size); for (let i = 0; i < value.length; i++) { bytes[i] = value.charCodeAt(i); } this.riffWriter.writeAscii(tag); this.riffWriter.writeU32(size); this.writer.write(bytes); // Add padding byte if size is odd if (size & 1) { this.writer.write(new Uint8Array(1)); } writtenTags.add(tag); }; for (const { key, value } of keyValueIterator(metadata)) { switch (key) { case 'title': { writeInfoTag('INAM', value); writtenTags.add('INAM'); }; break; case 'artist': { writeInfoTag('IART', value); writtenTags.add('IART'); }; break; case 'album': { writeInfoTag('IPRD', value); writtenTags.add('IPRD'); }; break; case 'trackNumber': { const string = metadata.tracksTotal !== undefined ? `${value}/${metadata.tracksTotal}` : value.toString(); writeInfoTag('ITRK', string); writtenTags.add('ITRK'); }; break; case 'genre': { writeInfoTag('IGNR', value); writtenTags.add('IGNR'); }; break; case 'date': { writeInfoTag('ICRD', value.toISOString().slice(0, 10)); writtenTags.add('ICRD'); }; break; case 'comment': { writeInfoTag('ICMT', value); writtenTags.add('ICMT'); }; break; case 'albumArtist': case 'discNumber': case 'tracksTotal': case 'discsTotal': case 'description': case 'lyrics': case 'images': { // Not supported in RIFF INFO }; break; case 'raw': { // Handled later }; break; default: assertNever(key); } } if (metadata.raw) { for (const key in metadata.raw) { const value = metadata.raw[key]; if (value == null || key.length !== 4 || writtenTags.has(key)) { continue; } if (typeof value === 'string') { writeInfoTag(key, value); } } } const endPos = this.writer.getPos(); const chunkSize = endPos - startPos - 8; this.writer.seek(startPos + 4); this.riffWriter.writeU32(chunkSize); this.writer.seek(endPos); // Add padding byte if chunk size is odd if (chunkSize & 1) { this.writer.write(new Uint8Array(1)); } } private writeId3Chunk(metadata: MetadataTags) { const startPos = this.writer.getPos(); // Write RIFF chunk header this.riffWriter.writeAscii('ID3 '); this.riffWriter.writeU32(0); // Size placeholder const id3Writer = new Id3V2Writer(this.writer); const id3TagSize = id3Writer.writeId3V2Tag(metadata); const endPos = this.writer.getPos(); // Update RIFF chunk size this.writer.seek(startPos + 4); this.riffWriter.writeU32(id3TagSize); this.writer.seek(endPos); // Add padding byte if chunk size is odd if (id3TagSize & 1) { this.writer.write(new Uint8Array(1)); } } async finalize() { const release = await this.mutex.acquire(); const endPos = this.writer.getPos(); if (this.isRf64) { // Write riff size assert(this.ds64RiffSizePos !== null); this.writer.seek(this.ds64RiffSizePos); this.riffWriter.writeU64(endPos - 8); // Write data size assert(this.ds64DataSizePos !== null); this.writer.seek(this.ds64DataSizePos); this.riffWriter.writeU64(this.dataSize); // Write sample count assert(this.ds64SampleCountPos !== null); this.writer.seek(this.ds64SampleCountPos); this.riffWriter.writeU64(this.sampleCount); } else { // Write file size assert(this.riffSizePos !== null); this.writer.seek(this.riffSizePos); this.riffWriter.writeU32(endPos - 8); // Write data chunk size assert(this.dataSizePos !== null); this.writer.seek(this.dataSizePos); this.riffWriter.writeU32(this.dataSize); } release(); } } ===== src/metadata.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { isRecordStringString } from './misc'; /** * Represents descriptive (non-technical) metadata about a media file, such as title, author, date, cover art, or other * attached files. Common tags are normalized by Mediabunny into a uniform format, while the `raw` field can be used to * directly read or write the underlying metadata tags (which differ by format). * * - For MP4/QuickTime files, the metadata refers to the data in `'moov'`-level `'udta'` and `'meta'` atoms. * - For WebM/Matroska files, the metadata refers to the Tags and Attachments elements whose target is 50 (MOVIE). * - For MP3 files, the metadata refers to the ID3v2 or ID3v1 tags. * - For Ogg files, there is no global metadata so instead, the metadata refers to the combined metadata of all tracks, * in Vorbis-style comment headers. * - For WAVE files, the metadata refers to the chunks within the RIFF INFO chunk. * - For ADTS files, the metadata refers to the ID3v2 tags. * - For FLAC files, the metadata lives in Vorbis style in the Vorbis comment block, or sometimes in ID3v2 tags at the * start of the file. * - For MPEG-TS files, metadata tags are currently not supported. * * @group Metadata tags * @public */ export type MetadataTags = { /** Title of the media (e.g. Gangnam Style, Titanic, etc.) */ title?: string; /** Short description or subtitle of the media. */ description?: string; /** Primary artist(s) or creator(s) of the work. */ artist?: string; /** Album, collection, or compilation the media belongs to. */ album?: string; /** Main credited artist for the album/collection as a whole. */ albumArtist?: string; /** Position of this track within its album or collection (1-based). */ trackNumber?: number; /** Total number of tracks in the album or collection. */ tracksTotal?: number; /** Disc index if the release spans multiple discs (1-based). */ discNumber?: number; /** Total number of discs in the release. */ discsTotal?: number; /** Genre or category describing the media's style or content (e.g. Metal, Horror, etc.) */ genre?: string; /** Release, recording or creation date of the media. */ date?: Date; /** Full text lyrics or transcript associated with the media. */ lyrics?: string; /** Freeform notes, remarks or commentary about the media. */ comment?: string; /** Embedded images such as cover art, booklet scans, artwork or preview frames. */ images?: AttachedImage[]; /** * The raw, underlying metadata tags. * * This field can be used for both reading and writing. When reading, it represents the original tags that were used * to derive the normalized fields, and any additional metadata that Mediabunny doesn't understand. When writing, it * can be used to set arbitrary metadata tags in the output file. * * The format of these tags differs per format: * - MP4/QuickTime: By default, the keys refer to the names of the individual atoms in the `'ilst'` atom inside the * `'meta'` atom, and the values are derived from the content of the `'data'` atom inside them. When a `'keys'` atom * is also used, then the keys reflect the keys specified there (such as `'com.apple.quicktime.version'`). * Additionally, any atoms within the `'udta'` atom are dumped into here, however with unknown internal format * (`Uint8Array`). * - WebM/Matroska: `SimpleTag` elements whose target is 50 (MOVIE), either containing string or `Uint8Array` * values. Additionally, all attached files (such as font files) are included here, where the key corresponds to * the FileUID and the value is an {@link AttachedFile}. * - MP3: The ID3v2 tags, or a single `'TAG'` key with the contents of the ID3v1 tag. The ID3v2 `'TXXX'` * user-defined text frames are exposed as a `Record`. * - ADTS: The ID3v2 tags, just like in MP3. * - Ogg: The key-value string pairs from the Vorbis-style comment header (see RFC 7845, Section 5.2). * Additionally, the `'vendor'` key refers to the vendor string within this header. * - WAVE: The individual metadata chunks within the RIFF INFO chunk. Values are always ISO 8859-1 strings. * - FLAC: The key-value string pairs from the vorbis metadata block (see RFC 9639, Section D.2.3). * Additionally, the `'vendor'` key refers to the vendor string within this header. If ID3v2 tags appear at the * start of the file, their content is stored just like for MP3. * - MPEG-TS: Not supported. */ raw?: Record | null>; }; /** * An embedded image such as cover art, booklet scan, artwork or preview frame. * * @group Metadata tags * @public */ export type AttachedImage = { /** The raw image data. */ data: Uint8Array; /** An RFC 6838 MIME type (e.g. image/jpeg, image/png, etc.) */ mimeType: string; /** The kind or purpose of the image. */ kind: 'coverFront' | 'coverBack' | 'unknown'; /** The name of the image file. */ name?: string; /** A description of the image. */ description?: string; }; /** * Image data with additional metadata. * * @group Metadata tags * @public */ export class RichImageData { /** Creates a new {@link RichImageData}. */ constructor( /** The raw image data. */ public data: Uint8Array, /** An RFC 6838 MIME type (e.g. image/jpeg, image/png, etc.) */ public mimeType: string, ) { if (!(data instanceof Uint8Array)) { throw new TypeError('data must be a Uint8Array.'); } if (typeof mimeType !== 'string') { throw new TypeError('mimeType must be a string.'); } } } /** * A file attached to a media file. * * @group Metadata tags * @public */ export class AttachedFile { /** Creates a new {@link AttachedFile}. */ constructor( /** The raw file data. */ public data: Uint8Array, /** An RFC 6838 MIME type (e.g. image/jpeg, image/png, font/ttf, etc.) */ public mimeType?: string, /** The name of the file. */ public name?: string, /** A description of the file. */ public description?: string, ) { if (!(data instanceof Uint8Array)) { throw new TypeError('data must be a Uint8Array.'); } if (mimeType !== undefined && typeof mimeType !== 'string') { throw new TypeError('mimeType, when provided, must be a string.'); } if (name !== undefined && typeof name !== 'string') { throw new TypeError('name, when provided, must be a string.'); } if (description !== undefined && typeof description !== 'string') { throw new TypeError('description, when provided, must be a string.'); } } }; export const validateMetadataTags = (tags: MetadataTags) => { if (!tags || typeof tags !== 'object') { throw new TypeError('tags must be an object.'); } if (tags.title !== undefined && typeof tags.title !== 'string') { throw new TypeError('tags.title, when provided, must be a string.'); } if (tags.description !== undefined && typeof tags.description !== 'string') { throw new TypeError('tags.description, when provided, must be a string.'); } if (tags.artist !== undefined && typeof tags.artist !== 'string') { throw new TypeError('tags.artist, when provided, must be a string.'); } if (tags.album !== undefined && typeof tags.album !== 'string') { throw new TypeError('tags.album, when provided, must be a string.'); } if (tags.albumArtist !== undefined && typeof tags.albumArtist !== 'string') { throw new TypeError('tags.albumArtist, when provided, must be a string.'); } if (tags.trackNumber !== undefined && (!Number.isInteger(tags.trackNumber) || tags.trackNumber <= 0)) { throw new TypeError('tags.trackNumber, when provided, must be a positive integer.'); } if ( tags.tracksTotal !== undefined && (!Number.isInteger(tags.tracksTotal) || tags.tracksTotal <= 0) ) { throw new TypeError('tags.tracksTotal, when provided, must be a positive integer.'); } if (tags.discNumber !== undefined && (!Number.isInteger(tags.discNumber) || tags.discNumber <= 0)) { throw new TypeError('tags.discNumber, when provided, must be a positive integer.'); } if ( tags.discsTotal !== undefined && (!Number.isInteger(tags.discsTotal) || tags.discsTotal <= 0) ) { throw new TypeError('tags.discsTotal, when provided, must be a positive integer.'); } if (tags.genre !== undefined && typeof tags.genre !== 'string') { throw new TypeError('tags.genre, when provided, must be a string.'); } if (tags.date !== undefined && (!(tags.date instanceof Date) || Number.isNaN(tags.date.getTime()))) { throw new TypeError('tags.date, when provided, must be a valid Date.'); } if (tags.lyrics !== undefined && typeof tags.lyrics !== 'string') { throw new TypeError('tags.lyrics, when provided, must be a string.'); } if (tags.images !== undefined) { if (!Array.isArray(tags.images)) { throw new TypeError('tags.images, when provided, must be an array.'); } for (const image of tags.images) { if (!image || typeof image !== 'object') { throw new TypeError('Each image in tags.images must be an object.'); } if (!(image.data instanceof Uint8Array)) { throw new TypeError('Each image.data must be a Uint8Array.'); } if (typeof image.mimeType !== 'string') { throw new TypeError('Each image.mimeType must be a string.'); } if (!['coverFront', 'coverBack', 'unknown'].includes(image.kind)) { throw new TypeError('Each image.kind must be \'coverFront\', \'coverBack\', or \'unknown\'.'); } } } if (tags.comment !== undefined && typeof tags.comment !== 'string') { throw new TypeError('tags.comment, when provided, must be a string.'); } if (tags.raw !== undefined) { if (!tags.raw || typeof tags.raw !== 'object') { throw new TypeError('tags.raw, when provided, must be an object.'); } for (const value of Object.values(tags.raw)) { if ( value !== null && typeof value !== 'string' && !(value instanceof Uint8Array) && !(value instanceof RichImageData) && !(value instanceof AttachedFile) && !isRecordStringString(value) ) { throw new TypeError( 'Each value in tags.raw must be a string, Uint8Array, RichImageData, AttachedFile, ' + 'Record, or null.', ); } } } }; export const metadataTagsAreEmpty = (tags: MetadataTags) => { return tags.title === undefined && tags.description === undefined && tags.artist === undefined && tags.album === undefined && tags.albumArtist === undefined && tags.trackNumber === undefined && tags.tracksTotal === undefined && tags.discNumber === undefined && tags.discsTotal === undefined && tags.genre === undefined && tags.date === undefined && tags.lyrics === undefined && (!tags.images || tags.images.length === 0) && tags.comment === undefined && (tags.raw === undefined || Object.keys(tags.raw).length === 0); }; /** * Specifies a track's disposition, i.e. information about its intended usage. * @public * @group Miscellaneous */ export type TrackDisposition = { /** * Indicates that this track is eligible for automatic selection by a player. Multiple tracks can be default tracks. */ default: boolean; /** Indicates that the track is the primary track among other tracks of its type. */ primary: boolean; /** * Indicates that players should always display this track by default, even if it goes against the user's default * preferences. For example, a subtitle track only containing translations of foreign-language audio. */ forced: boolean; /** Indicates that this track is in the content's original language. */ original: boolean; /** Indicates that this track contains commentary. */ commentary: boolean; /** Indicates that this track is intended for hearing-impaired users. */ hearingImpaired: boolean; /** Indicates that this track is intended for visually-impaired users. */ visuallyImpaired: boolean; }; export const DEFAULT_TRACK_DISPOSITION: TrackDisposition = { default: true, primary: true, forced: false, original: false, commentary: false, hearingImpaired: false, visuallyImpaired: false, }; export const validateTrackDisposition = (disposition: Partial) => { if (!disposition || typeof disposition !== 'object') { throw new TypeError('disposition must be an object.'); } if (disposition.default !== undefined && typeof disposition.default !== 'boolean') { throw new TypeError('disposition.default must be a boolean.'); } if (disposition.primary !== undefined && typeof disposition.primary !== 'boolean') { throw new TypeError('disposition.primary must be a boolean.'); } if (disposition.forced !== undefined && typeof disposition.forced !== 'boolean') { throw new TypeError('disposition.forced must be a boolean.'); } if (disposition.original !== undefined && typeof disposition.original !== 'boolean') { throw new TypeError('disposition.original must be a boolean.'); } if (disposition.commentary !== undefined && typeof disposition.commentary !== 'boolean') { throw new TypeError('disposition.commentary must be a boolean.'); } if (disposition.hearingImpaired !== undefined && typeof disposition.hearingImpaired !== 'boolean') { throw new TypeError('disposition.hearingImpaired must be a boolean.'); } if (disposition.visuallyImpaired !== undefined && typeof disposition.visuallyImpaired !== 'boolean') { throw new TypeError('disposition.visuallyImpaired must be a boolean.'); } }; ===== src/encode.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AUDIO_CODECS, AudioCodec, buildAudioCodecString, buildVideoCodecString, getAudioEncoderConfigExtension, getVideoEncoderConfigExtension, inferCodecFromCodecString, MediaCodec, PCM_AUDIO_CODECS, SUBTITLE_CODECS, SubtitleCodec, VIDEO_CODECS, VideoCodec, } from './codec'; import { customAudioEncoders, customVideoEncoders } from './custom-coder'; import { isFirefox, MaybePromise, Rotation } from './misc'; import { EncodedPacket } from './packet'; import { AudioSample, CropRectangle, validateCropRectangle, VideoSample, VideoSampleResource } from './sample'; export const canEncodeVideoMemo = new Map>(); export const canEncodeAudioMemo = new Map>(); /** * Configuration object that controls video encoding. Can be used to set codec, quality, and more. * @group Encoding * @public */ export type VideoEncodingConfig = { /** The video codec that should be used for encoding the video samples (frames). */ codec: VideoCodec; /** * The target bitrate for the encoded video, in bits per second. Alternatively, a subjective {@link Quality} can * be provided. */ bitrate: number | Quality; /** * The interval, in seconds, of how often frames are encoded as a key frame. The default is 2 seconds. Frequent key * frames improve seeking behavior but increase file size. When using multiple video tracks, you should give them * all the same key frame interval. */ keyFrameInterval?: number; /** * Video frames may change size over time. This field controls the behavior in case this happens. * * - `'deny'` (default) will throw an error, requiring all frames to have the exact same dimensions. * - `'passThrough'` will allow the change and directly pass the frame to the encoder. * - `'fill'` will stretch the image to fill the entire original box, potentially altering aspect ratio. * - `'contain'` will contain the entire image within the original box while preserving aspect ratio. This may lead * to letterboxing. * - `'cover'` will scale the image until the entire original box is filled, while preserving aspect ratio. * * The "original box" refers to the dimensions of the first encoded frame. */ sizeChangeBehavior?: 'deny' | 'passThrough' | 'fill' | 'contain' | 'cover'; /** * Optional transformations to apply to the video frames before they are passed to the encoder. */ transform?: VideoTransformOptions; /** Called for each successfully encoded packet. Both the packet and the encoding metadata are passed. */ onEncodedPacket?: (packet: EncodedPacket, meta: EncodedVideoChunkMetadata | undefined) => unknown; /** * Called when the internal [encoder config](https://www.w3.org/TR/webcodecs/#video-encoder-config), as used by the * WebCodecs API, is created. */ onEncoderConfig?: (config: VideoEncoderConfig) => unknown; /** Called right before a sample is passed to the encoder. */ onEncodedSample?: (sample: VideoSample) => unknown; } & VideoEncodingAdditionalOptions; /** * Options for transforming video frames before encoding. * @group Encoding * @public */ export type VideoTransformOptions = { /** * The width in pixels to resize the frames to. If height is not set, it will be deduced * automatically based on aspect ratio. */ width?: number; /** * The height in pixels to resize the frames to. If width is not set, it will be deduced * automatically based on aspect ratio. */ height?: number; /** * The fitting algorithm in case both width and height are set. * * - `'fill'` will stretch the image to fill the entire box, potentially altering aspect ratio. * - `'contain'` will contain the entire image within the box while preserving aspect ratio. This may lead to * letterboxing. * - `'cover'` will scale the image until the entire box is filled, while preserving aspect ratio. * * To avoid ambiguity, this field must not be set when `sizeChangeBehavior` is `'fill'`, `'contain'` or * `'deny'`, since `sizeChangeBehavior` already determines the fitting algorithm. */ fit?: 'fill' | 'contain' | 'cover'; /** * The clockwise rotation by which to rotate the frames. Rotation is applied before resizing. */ rotate?: Rotation; /** * Specifies the rectangular region of the frames to crop to. The crop region will automatically be * clamped to the dimensions of the frame. Cropping is performed after rotation but before resizing. */ crop?: CropRectangle; /** * Whether to discard or keep the transparency information of the video samples. The default is `'keep'`. */ alpha?: 'keep' | 'discard'; /** * The frame rate in hertz to normalize the video frame stream to. */ frameRate?: number; /** * Allows for custom user-defined processing of video frames, e.g. for applying overlays, color transformations, * or timestamp modifications. Will be called for each video frame after transformations and frame rate * corrections. * * Must return a {@link VideoSample}, a {@link VideoSampleResource} or a `CanvasImageSource`, an array of them, or * `null` for dropping the frame. When non-timestamped data is returned, the timestamp and duration from the input * sample will be used. */ process?: (sample: VideoSample) => MaybePromise< CanvasImageSource | VideoSample | VideoSampleResource | (CanvasImageSource | VideoSample | VideoSampleResource)[] | null >; /** * Forces every video frame through the transformation step even if no transformation properties are defined. * This can be used, for example, to bake rotation into the encoded video frames. */ force?: boolean; }; export const validateVideoEncodingConfig = (config: VideoEncodingConfig) => { if (!config || typeof config !== 'object') { throw new TypeError('Encoding config must be an object.'); } if (!VIDEO_CODECS.includes(config.codec)) { throw new TypeError(`Invalid video codec '${config.codec}'. Must be one of: ${VIDEO_CODECS.join(', ')}.`); } if (!(config.bitrate instanceof Quality) && (!Number.isInteger(config.bitrate) || config.bitrate <= 0)) { throw new TypeError('config.bitrate must be a positive integer or a quality.'); } if ( config.keyFrameInterval !== undefined && (!Number.isFinite(config.keyFrameInterval) || config.keyFrameInterval < 0) ) { throw new TypeError('config.keyFrameInterval, when provided, must be a non-negative number.'); } if ( config.sizeChangeBehavior !== undefined && !['deny', 'passThrough', 'fill', 'contain', 'cover'].includes(config.sizeChangeBehavior) ) { throw new TypeError( 'config.sizeChangeBehavior, when provided, must be \'deny\', \'passThrough\', \'fill\', \'contain\'' + ' or \'cover\'.', ); } if (config.transform !== undefined) { if (typeof config.transform !== 'object' || !config.transform) { throw new TypeError('config.transform, when provided, must be an object.'); } if ( config.transform.width !== undefined && (!Number.isInteger(config.transform.width) || config.transform.width <= 0) ) { throw new TypeError('config.transform.width, when provided, must be a positive integer.'); } if ( config.transform.height !== undefined && (!Number.isInteger(config.transform.height) || config.transform.height <= 0) ) { throw new TypeError('config.transform.height, when provided, must be a positive integer.'); } if (config.transform.fit !== undefined && !['fill', 'contain', 'cover'].includes(config.transform.fit)) { throw new TypeError('config.transform.fit, when provided, must be one of "fill", "contain", or "cover".'); } if ( config.transform.width !== undefined && config.transform.height !== undefined && config.transform.fit === undefined && !['fill', 'contain', 'cover'].includes(config.sizeChangeBehavior!) ) { throw new TypeError( 'When both config.transform.width and config.transform.height are provided, config.transform.fit' + ' must also be provided.', ); } if ( config.transform.fit !== undefined && ['fill', 'contain', 'cover'].includes(config.sizeChangeBehavior!) && config.transform.fit !== config.sizeChangeBehavior ) { throw new TypeError( 'config.transform.fit, when provided, cannot differ from config.sizeChangeBehavior when' + ' config.sizeChangeBehavior is \'fill\', \'contain\' or \'cover\', as sizeChangeBehavior already' + ' determines the fitting algorithm.', ); } if (config.transform.rotate !== undefined && ![0, 90, 180, 270].includes(config.transform.rotate)) { throw new TypeError('config.transform.rotate, when provided, must be 0, 90, 180 or 270.'); } if (config.transform.crop !== undefined) { validateCropRectangle(config.transform.crop, 'config.transform.'); } if (config.transform.process !== undefined && typeof config.transform.process !== 'function') { throw new TypeError('config.transform.process, when provided, must be a function.'); } if ( config.transform.frameRate !== undefined && (!Number.isFinite(config.transform.frameRate) || config.transform.frameRate <= 0) ) { throw new TypeError('config.transform.frameRate, when provided, must be a finite positive number.'); } if (config.transform.force !== undefined && typeof config.transform.force !== 'boolean') { throw new TypeError('config.transform.force, when provided, must be a boolean.'); } } if (config.onEncodedPacket !== undefined && typeof config.onEncodedPacket !== 'function') { throw new TypeError('config.onEncodedPacket, when provided, must be a function.'); } if (config.onEncoderConfig !== undefined && typeof config.onEncoderConfig !== 'function') { throw new TypeError('config.onEncoderConfig, when provided, must be a function.'); } if (config.onEncodedSample !== undefined && typeof config.onEncodedSample !== 'function') { throw new TypeError('config.onEncodedSample, when provided, must be a function.'); } validateVideoEncodingAdditionalOptions(config.codec, config); }; /** * Additional options that control video encoding. * @group Encoding * @public */ export type VideoEncodingAdditionalOptions = { /** * What to do with alpha data contained in the video samples. * * - `'discard'` (default): Only the samples' color data is kept; the video is opaque. * - `'keep'`: The samples' alpha data is also encoded as side data. Make sure to pair this mode with a container * format that supports transparency (such as WebM or Matroska). */ alpha?: 'discard' | 'keep'; /** Configures the bitrate mode; defaults to `'variable'`. */ bitrateMode?: 'constant' | 'variable'; /** * The latency mode used by the encoder; controls the performance-quality tradeoff. * * - `'quality'` (default): The encoder prioritizes quality over latency, and no frames can be dropped. * - `'realtime'`: The encoder prioritizes low latency over quality, and may drop frames if the encoder becomes * overloaded to keep up with real-time requirements. */ latencyMode?: 'quality' | 'realtime'; /** * The full codec string as specified in the Mediabunny Codec Registry. This string must match the codec * specified in `codec`. When not set, a fitting codec string will be constructed automatically by the library. */ fullCodecString?: string; /** * A hint that configures the hardware acceleration method of this codec. This is best left on `'no-preference'`, * the default. */ hardwareAcceleration?: 'no-preference' | 'prefer-hardware' | 'prefer-software'; /** * An encoding scalability mode identifier as defined by * [WebRTC-SVC](https://w3c.github.io/webrtc-svc/#scalabilitymodes*). */ scalabilityMode?: string; /** * An encoding video content hint as defined by * [mst-content-hint](https://w3c.github.io/mst-content-hint/#video-content-hints). */ contentHint?: string; }; export const validateVideoEncodingAdditionalOptions = (codec: VideoCodec, options: VideoEncodingAdditionalOptions) => { if (!options || typeof options !== 'object') { throw new TypeError('Encoding options must be an object.'); } if (options.alpha !== undefined && !['discard', 'keep'].includes(options.alpha)) { throw new TypeError('options.alpha, when provided, must be \'discard\' or \'keep\'.'); } if (options.bitrateMode !== undefined && !['constant', 'variable'].includes(options.bitrateMode)) { throw new TypeError('bitrateMode, when provided, must be \'constant\' or \'variable\'.'); } if (options.latencyMode !== undefined && !['quality', 'realtime'].includes(options.latencyMode)) { throw new TypeError('latencyMode, when provided, must be \'quality\' or \'realtime\'.'); } if (options.fullCodecString !== undefined && typeof options.fullCodecString !== 'string') { throw new TypeError('fullCodecString, when provided, must be a string.'); } if (options.fullCodecString !== undefined && inferCodecFromCodecString(options.fullCodecString) !== codec) { throw new TypeError( `fullCodecString, when provided, must be a string that matches the specified codec (${codec}).`, ); } if ( options.hardwareAcceleration !== undefined && !['no-preference', 'prefer-hardware', 'prefer-software'].includes(options.hardwareAcceleration) ) { throw new TypeError( 'hardwareAcceleration, when provided, must be \'no-preference\', \'prefer-hardware\' or' + ' \'prefer-software\'.', ); } if (options.scalabilityMode !== undefined && typeof options.scalabilityMode !== 'string') { throw new TypeError('scalabilityMode, when provided, must be a string.'); } if (options.contentHint !== undefined && typeof options.contentHint !== 'string') { throw new TypeError('contentHint, when provided, must be a string.'); } }; export const buildVideoEncoderConfig = (options: { codec: VideoCodec; width: number; height: number; bitrate: number | Quality; framerate: number | undefined; squarePixelWidth?: number; squarePixelHeight?: number; } & VideoEncodingAdditionalOptions): VideoEncoderConfig => { const resolvedBitrate = options.bitrate instanceof Quality ? options.bitrate._toVideoBitrate(options.codec, options.width, options.height) : options.bitrate; return { codec: options.fullCodecString ?? buildVideoCodecString( options.codec, options.width, options.height, resolvedBitrate, ), width: options.width, height: options.height, displayWidth: options.squarePixelWidth, displayHeight: options.squarePixelHeight, bitrate: resolvedBitrate, bitrateMode: options.bitrateMode, alpha: options.alpha ?? 'discard', framerate: options.framerate, latencyMode: options.latencyMode, hardwareAcceleration: options.hardwareAcceleration, scalabilityMode: options.scalabilityMode, contentHint: options.contentHint, ...getVideoEncoderConfigExtension(options.codec), }; }; /** * Configuration object that controls audio encoding. Can be used to set codec, quality, and more. * @group Encoding * @public */ export type AudioEncodingConfig = { /** The audio codec that should be used for encoding the audio samples. */ codec: AudioCodec; /** * The target bitrate for the encoded audio, in bits per second. Alternatively, a subjective {@link Quality} can * be provided. Required for compressed audio codecs, unused for PCM codecs. */ bitrate?: number | Quality; /** * Optional transformations to apply to the audio samples before they are passed to the encoder. */ transform?: AudioTransformOptions; /** Called for each successfully encoded packet. Both the packet and the encoding metadata are passed. */ onEncodedPacket?: (packet: EncodedPacket, meta: EncodedAudioChunkMetadata | undefined) => unknown; /** * Called when the internal [encoder config](https://www.w3.org/TR/webcodecs/#audio-encoder-config), as used by the * WebCodecs API, is created. */ onEncoderConfig?: (config: AudioEncoderConfig) => unknown; /** Called right before a sample is passed to the encoder. */ onEncodedSample?: (sample: AudioSample) => unknown; } & AudioEncodingAdditionalOptions; /** * Options for transforming audio samples before encoding. * @group Encoding * @public */ export type AudioTransformOptions = { /** The desired number of output channels to up/downmix to. */ numberOfChannels?: number; /** The desired output sample rate in hertz to resample to. */ sampleRate?: number; /** * The desired sample format (and therefore bit depth) of the audio samples before they are passed to the encoder. * Can be used to control bit depth with certain output codecs such as FLAC. */ sampleFormat?: 'u8' | 's16' | 's32' | 'f32'; /** * Allows for custom user-defined processing of audio samples, e.g. for applying audio effects or timestamp * modifications. Called for each audio sample after resampling and remixing. * * Must return an {@link AudioSample}, an array of them, or `null` for dropping the sample. */ process?: (sample: AudioSample) => MaybePromise< AudioSample | AudioSample[] | null >; }; export const validateAudioEncodingConfig = (config: AudioEncodingConfig) => { if (!config || typeof config !== 'object') { throw new TypeError('Encoding config must be an object.'); } if (!AUDIO_CODECS.includes(config.codec)) { throw new TypeError(`Invalid audio codec '${config.codec}'. Must be one of: ${AUDIO_CODECS.join(', ')}.`); } if ( config.bitrate === undefined && !((PCM_AUDIO_CODECS as readonly string[]).includes(config.codec) || config.codec === 'flac') ) { throw new TypeError('config.bitrate must be provided for compressed audio codecs.'); } if ( config.bitrate !== undefined && !(config.bitrate instanceof Quality) && (!Number.isInteger(config.bitrate) || config.bitrate <= 0) ) { throw new TypeError('config.bitrate, when provided, must be a positive integer or a quality.'); } if (config.transform !== undefined) { if (typeof config.transform !== 'object' || !config.transform) { throw new TypeError('config.transform, when provided, must be an object.'); } if ( config.transform.numberOfChannels !== undefined && (!Number.isInteger(config.transform.numberOfChannels) || config.transform.numberOfChannels <= 0) ) { throw new TypeError('config.transform.numberOfChannels, when provided, must be a positive integer.'); } if ( config.transform.sampleRate !== undefined && (!Number.isInteger(config.transform.sampleRate) || config.transform.sampleRate <= 0) ) { throw new TypeError('config.transform.sampleRate, when provided, must be a positive integer.'); } if ( config.transform.sampleFormat !== undefined && !['u8', 's16', 's32', 'f32'].includes(config.transform.sampleFormat) ) { throw new TypeError('config.transform.sampleFormat, when provided, must be one of: u8, s16, s32, f32.'); } if (config.transform.process !== undefined && typeof config.transform.process !== 'function') { throw new TypeError('config.transform.process, when provided, must be a function.'); } } if (config.onEncodedPacket !== undefined && typeof config.onEncodedPacket !== 'function') { throw new TypeError('config.onEncodedPacket, when provided, must be a function.'); } if (config.onEncoderConfig !== undefined && typeof config.onEncoderConfig !== 'function') { throw new TypeError('config.onEncoderConfig, when provided, must be a function.'); } if (config.onEncodedSample !== undefined && typeof config.onEncodedSample !== 'function') { throw new TypeError('config.onEncodedSample, when provided, must be a function.'); } validateAudioEncodingAdditionalOptions(config.codec, config); }; /** * Additional options that control audio encoding. * @group Encoding * @public */ export type AudioEncodingAdditionalOptions = { /** Configures the bitrate mode. */ bitrateMode?: 'constant' | 'variable'; /** * The full codec string as specified in the Mediabunny Codec Registry. This string must match the codec * specified in `codec`. When not set, a fitting codec string will be constructed automatically by the library. */ fullCodecString?: string; }; export const validateAudioEncodingAdditionalOptions = (codec: AudioCodec, options: AudioEncodingAdditionalOptions) => { if (!options || typeof options !== 'object') { throw new TypeError('Encoding options must be an object.'); } if (options.bitrateMode !== undefined && !['constant', 'variable'].includes(options.bitrateMode)) { throw new TypeError('bitrateMode, when provided, must be \'constant\' or \'variable\'.'); } if (options.fullCodecString !== undefined && typeof options.fullCodecString !== 'string') { throw new TypeError('fullCodecString, when provided, must be a string.'); } if (options.fullCodecString !== undefined && inferCodecFromCodecString(options.fullCodecString) !== codec) { throw new TypeError( `fullCodecString, when provided, must be a string that matches the specified codec (${codec}).`, ); } }; export const buildAudioEncoderConfig = (options: { codec: AudioCodec; numberOfChannels: number; sampleRate: number; bitrate?: number | Quality; } & AudioEncodingAdditionalOptions): AudioEncoderConfig => { const resolvedBitrate = options.bitrate instanceof Quality ? options.bitrate._toAudioBitrate(options.codec) : options.bitrate; return { codec: options.fullCodecString ?? buildAudioCodecString( options.codec, options.numberOfChannels, options.sampleRate, ), numberOfChannels: options.numberOfChannels, sampleRate: options.sampleRate, bitrate: resolvedBitrate, bitrateMode: options.bitrateMode, ...getAudioEncoderConfigExtension(options.codec), }; }; /** * Represents a subjective media quality level. * @group Encoding * @public */ export class Quality { /** @internal */ _factor: number; /** @internal */ constructor(factor: number) { this._factor = factor; } /** @internal */ _toVideoBitrate(codec: VideoCodec, width: number, height: number) { const pixels = width * height; const codecEfficiencyFactors = { avc: 1.0, // H.264/AVC (baseline) hevc: 0.6, // H.265/HEVC (~40% more efficient than AVC) vp9: 0.6, // Similar to HEVC av1: 0.4, // ~60% more efficient than AVC vp8: 1.2, // Slightly less efficient than AVC }; const referencePixels = 1920 * 1080; const referenceBitrate = 3000000; const scaleFactor = Math.pow(pixels / referencePixels, 0.95); // Slight non-linear scaling const baseBitrate = referenceBitrate * scaleFactor; const codecAdjustedBitrate = baseBitrate * codecEfficiencyFactors[codec]; const finalBitrate = codecAdjustedBitrate * this._factor; return Math.ceil(finalBitrate / 1000) * 1000; } /** @internal */ _toAudioBitrate(codec: AudioCodec) { if ((PCM_AUDIO_CODECS as readonly string[]).includes(codec) || codec === 'flac') { return undefined; } const baseRates = { aac: 128000, // 128kbps base for AAC opus: 64000, // 64kbps base for Opus mp3: 160000, // 160kbps base for MP3 vorbis: 64000, // 64kbps base for Vorbis ac3: 384000, // 384kbps base for AC-3 eac3: 192000, // 192kbps base for E-AC-3 }; const baseBitrate = baseRates[codec as keyof typeof baseRates]; if (!baseBitrate) { throw new Error(`Unhandled codec: ${codec}`); } let finalBitrate = baseBitrate * this._factor; if (codec === 'aac') { // AAC only works with specific bitrates, let's find the closest const validRates = [96000, 128000, 160000, 192000]; finalBitrate = validRates.reduce((prev, curr) => Math.abs(curr - finalBitrate) < Math.abs(prev - finalBitrate) ? curr : prev, ); } else if (codec === 'opus' || codec === 'vorbis') { finalBitrate = Math.max(6000, finalBitrate); } else if (codec === 'mp3') { const validRates = [ 8000, 16000, 24000, 32000, 40000, 48000, 64000, 80000, 96000, 112000, 128000, 160000, 192000, 224000, 256000, 320000, ]; finalBitrate = validRates.reduce((prev, curr) => Math.abs(curr - finalBitrate) < Math.abs(prev - finalBitrate) ? curr : prev, ); } return Math.round(finalBitrate / 1000) * 1000; } } /** * Represents a very low media quality. * @group Encoding * @public */ export const QUALITY_VERY_LOW = /* #__PURE__ */ new Quality(0.3); /** * Represents a low media quality. * @group Encoding * @public */ export const QUALITY_LOW = /* #__PURE__ */ new Quality(0.6); /** * Represents a medium media quality. * @group Encoding * @public */ export const QUALITY_MEDIUM = /* #__PURE__ */ new Quality(1); /** * Represents a high media quality. * @group Encoding * @public */ export const QUALITY_HIGH = /* #__PURE__ */ new Quality(2); /** * Represents a very high media quality. * @group Encoding * @public */ export const QUALITY_VERY_HIGH = /* #__PURE__ */ new Quality(4); /** * Checks if the browser is able to encode the given codec. * @group Encoding * @public */ export const canEncode = (codec: MediaCodec) => { if ((VIDEO_CODECS as readonly string[]).includes(codec)) { return canEncodeVideo(codec as VideoCodec); } else if ((AUDIO_CODECS as readonly string[]).includes(codec)) { return canEncodeAudio(codec as AudioCodec); } else if ((SUBTITLE_CODECS as readonly string[]).includes(codec)) { return canEncodeSubtitles(codec as SubtitleCodec); } throw new TypeError(`Unknown codec '${codec}'.`); }; /** * Checks if the browser is able to encode the given video codec with the given parameters. * @group Encoding * @public */ export const canEncodeVideo = async ( codec: VideoCodec, options: { width?: number; height?: number; bitrate?: number | Quality; } & VideoEncodingAdditionalOptions = {}, ) => { const { width = 1280, height = 720, bitrate = 1e6, ...restOptions } = options; if (!VIDEO_CODECS.includes(codec)) { return false; } if (!Number.isInteger(width) || width <= 0) { throw new TypeError('width must be a positive integer.'); } if (!Number.isInteger(height) || height <= 0) { throw new TypeError('height must be a positive integer.'); } if (!(bitrate instanceof Quality) && (!Number.isInteger(bitrate) || bitrate <= 0)) { throw new TypeError('bitrate must be a positive integer or a quality.'); } validateVideoEncodingAdditionalOptions(codec, restOptions); const encoderConfig = buildVideoEncoderConfig({ codec, width, height, bitrate, framerate: undefined, ...restOptions, alpha: 'discard', // Since we handle alpha ourselves }); const key = JSON.stringify(encoderConfig); const memoized = canEncodeVideoMemo.get(key); if (memoized) { return memoized; } const promise = (async () => { if (customVideoEncoders.some(x => x.supports(codec, encoderConfig))) { // There's a custom encoder return true; } if (typeof VideoEncoder === 'undefined') { return false; } const hasOddDimension = width % 2 === 1 || height % 2 === 1; if ( hasOddDimension && (codec === 'avc' || codec === 'hevc') ) { // Disallow odd dimensions for certain codecs return false; } const support = await VideoEncoder.isConfigSupported(encoderConfig); if (!support.supported) { return false; } if (isFirefox()) { // isConfigSupported on Firefox appears to unreliably indicate if encoding will actually succeed. Therefore, // we just try encoding a frame to see if it actually works. // https://github.com/Vanilagy/mediabunny/issues/222 // eslint-disable-next-line @typescript-eslint/no-misused-promises, no-async-promise-executor return new Promise(async (resolve) => { try { const encoder = new VideoEncoder({ output: () => {}, error: () => resolve(false), }); encoder.configure(encoderConfig); const frameData = new Uint8Array(width * height * 4); const frame = new VideoFrame(frameData, { format: 'RGBA', codedWidth: width, codedHeight: height, timestamp: 0, }); encoder.encode(frame); frame.close(); await encoder.flush(); resolve(true); } catch { resolve(false); } }); } return true; })(); canEncodeVideoMemo.set(key, promise); return promise; }; /** * Checks if the browser is able to encode the given audio codec with the given parameters. * @group Encoding * @public */ export const canEncodeAudio = async ( codec: AudioCodec, options: { numberOfChannels?: number; sampleRate?: number; bitrate?: number | Quality; } & AudioEncodingAdditionalOptions = {}, ) => { const { numberOfChannels = 2, sampleRate = 48000, bitrate = 128e3, ...restOptions } = options; if (!AUDIO_CODECS.includes(codec)) { return false; } if (!Number.isInteger(numberOfChannels) || numberOfChannels <= 0) { throw new TypeError('numberOfChannels must be a positive integer.'); } if (!Number.isInteger(sampleRate) || sampleRate <= 0) { throw new TypeError('sampleRate must be a positive integer.'); } if (!(bitrate instanceof Quality) && (!Number.isInteger(bitrate) || bitrate <= 0)) { throw new TypeError('bitrate must be a positive integer.'); } validateAudioEncodingAdditionalOptions(codec, restOptions); const encoderConfig = buildAudioEncoderConfig({ codec, numberOfChannels, sampleRate, bitrate, ...restOptions, }); const key = JSON.stringify(encoderConfig); const memoized = canEncodeAudioMemo.get(key); if (memoized) { return memoized; } const promise = (async () => { if (customAudioEncoders.some(x => x.supports(codec, encoderConfig))) { // There's a custom encoder return true; } if ((PCM_AUDIO_CODECS as readonly string[]).includes(codec)) { return true; // Because we encode these ourselves } if (typeof AudioEncoder === 'undefined') { return false; } const support = await AudioEncoder.isConfigSupported(encoderConfig); return support.supported === true; })(); canEncodeAudioMemo.set(key, promise); return promise; }; /** * Checks if the browser is able to encode the given subtitle codec. * @group Encoding * @public */ export const canEncodeSubtitles = async (codec: SubtitleCodec) => { if (!SUBTITLE_CODECS.includes(codec)) { return false; } return true; }; /** * Returns the list of all media codecs that can be encoded by the browser. * @group Encoding * @public */ export const getEncodableCodecs = async (): Promise => { const [videoCodecs, audioCodecs, subtitleCodecs] = await Promise.all([ getEncodableVideoCodecs(), getEncodableAudioCodecs(), getEncodableSubtitleCodecs(), ]); return [...videoCodecs, ...audioCodecs, ...subtitleCodecs]; }; /** * Returns the list of all video codecs that can be encoded by the browser. * @group Encoding * @public */ export const getEncodableVideoCodecs = async ( checkedCodecs: VideoCodec[] = VIDEO_CODECS as unknown as VideoCodec[], options?: { width?: number; height?: number; bitrate?: number | Quality; }, ): Promise => { const bools = await Promise.all(checkedCodecs.map(codec => canEncodeVideo(codec, options))); return checkedCodecs.filter((_, i) => bools[i]); }; /** * Returns the list of all audio codecs that can be encoded by the browser. * @group Encoding * @public */ export const getEncodableAudioCodecs = async ( checkedCodecs: AudioCodec[] = AUDIO_CODECS as unknown as AudioCodec[], options?: { numberOfChannels?: number; sampleRate?: number; bitrate?: number | Quality; }, ): Promise => { const bools = await Promise.all(checkedCodecs.map(codec => canEncodeAudio(codec, options))); return checkedCodecs.filter((_, i) => bools[i]); }; /** * Returns the list of all subtitle codecs that can be encoded by the browser. * @group Encoding * @public */ export const getEncodableSubtitleCodecs = async ( checkedCodecs: SubtitleCodec[] = SUBTITLE_CODECS as unknown as SubtitleCodec[], ): Promise => { const bools = await Promise.all(checkedCodecs.map(canEncodeSubtitles)); return checkedCodecs.filter((_, i) => bools[i]); }; /** * Returns the first video codec from the given list that can be encoded by the browser. * @group Encoding * @public */ export const getFirstEncodableVideoCodec = async ( checkedCodecs: VideoCodec[], options?: { width?: number; height?: number; bitrate?: number | Quality; }, ): Promise => { for (const codec of checkedCodecs) { if (await canEncodeVideo(codec, options)) { return codec; } } return null; }; /** * Returns the first audio codec from the given list that can be encoded by the browser. * @group Encoding * @public */ export const getFirstEncodableAudioCodec = async ( checkedCodecs: AudioCodec[], options?: { numberOfChannels?: number; sampleRate?: number; bitrate?: number | Quality; }, ): Promise => { for (const codec of checkedCodecs) { if (await canEncodeAudio(codec, options)) { return codec; } } return null; }; /** * Returns the first subtitle codec from the given list that can be encoded by the browser. * @group Encoding * @public */ export const getFirstEncodableSubtitleCodec = async ( checkedCodecs: SubtitleCodec[], ): Promise => { for (const codec of checkedCodecs) { if (await canEncodeSubtitles(codec)) { return codec; } } return null; }; ===== src/subtitles.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ export type SubtitleCue = { timestamp: number; // in seconds duration: number; // in seconds text: string; identifier?: string; settings?: string; notes?: string; }; export type SubtitleConfig = { description: string; }; export type SubtitleMetadata = { config?: SubtitleConfig; }; type SubtitleParserOptions = { codec: 'webvtt'; output: (cue: SubtitleCue, metadata: SubtitleMetadata) => unknown; }; const cueBlockHeaderRegex = /(?:(.+?)\n)?((?:\d{2}:)?\d{2}:\d{2}.\d{3})\s+-->\s+((?:\d{2}:)?\d{2}:\d{2}.\d{3})/g; const preambleStartRegex = /^WEBVTT(.|\n)*?\n{2}/; export const inlineTimestampRegex = /<(?:(\d{2}):)?(\d{2}):(\d{2}).(\d{3})>/g; export class SubtitleParser { private options: SubtitleParserOptions; private preambleText: string | null = null; private preambleEmitted = false; constructor(options: SubtitleParserOptions) { this.options = options; } parse(text: string) { text = text.replaceAll('\r\n', '\n').replaceAll('\r', '\n'); cueBlockHeaderRegex.lastIndex = 0; let match: RegExpMatchArray | null; if (!this.preambleText) { if (!preambleStartRegex.test(text)) { throw new Error('WebVTT preamble incorrect.'); } match = cueBlockHeaderRegex.exec(text); const preamble = text.slice(0, match?.index ?? text.length).trimEnd(); if (!preamble) { throw new Error('No WebVTT preamble provided.'); } this.preambleText = preamble; if (match) { text = text.slice(match.index); cueBlockHeaderRegex.lastIndex = 0; } } while ((match = cueBlockHeaderRegex.exec(text))) { const notes = text.slice(0, match.index); const cueIdentifier = match[1]; const matchEnd = match.index! + match[0].length; const bodyStart = text.indexOf('\n', matchEnd) + 1; const cueSettings = text.slice(matchEnd, bodyStart).trim(); let bodyEnd = text.indexOf('\n\n', matchEnd); if (bodyEnd === -1) bodyEnd = text.length; const startTime = parseSubtitleTimestamp(match[2]!); const endTime = parseSubtitleTimestamp(match[3]!); const duration = endTime - startTime; const body = text.slice(bodyStart, bodyEnd).trim(); text = text.slice(bodyEnd).trimStart(); cueBlockHeaderRegex.lastIndex = 0; const cue: SubtitleCue = { timestamp: startTime / 1000, duration: duration / 1000, text: body, identifier: cueIdentifier, settings: cueSettings, notes, }; const meta: SubtitleMetadata = {}; if (!this.preambleEmitted) { meta.config = { description: this.preambleText, }; this.preambleEmitted = true; } this.options.output(cue, meta); } } } const timestampRegex = /(?:(\d{2}):)?(\d{2}):(\d{2}).(\d{3})/; export const parseSubtitleTimestamp = (string: string) => { const match = timestampRegex.exec(string); if (!match) throw new Error('Expected match.'); return 60 * 60 * 1000 * Number(match[1] || '0') + 60 * 1000 * Number(match[2]) + 1000 * Number(match[3]) + Number(match[4]); }; export const formatSubtitleTimestamp = (timestamp: number) => { const hours = Math.floor(timestamp / (60 * 60 * 1000)); const minutes = Math.floor((timestamp % (60 * 60 * 1000)) / (60 * 1000)); const seconds = Math.floor((timestamp % (60 * 1000)) / 1000); const milliseconds = timestamp % 1000; return hours.toString().padStart(2, '0') + ':' + minutes.toString().padStart(2, '0') + ':' + seconds.toString().padStart(2, '0') + '.' + milliseconds.toString().padStart(3, '0'); }; ===== src/custom-coder.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { AudioCodec, VideoCodec } from './codec'; import { canDecodeAudioMemo, canDecodeVideoMemo } from './decode'; import { canEncodeAudioMemo, canEncodeVideoMemo } from './encode'; import { Logging } from './logging'; import { MaybePromise } from './misc'; import { EncodedPacket } from './packet'; import { AudioSample, VideoSample } from './sample'; /** * Base class for custom video decoders. To add your own custom video decoder, extend this class, implement the * abstract methods and static `supports` method, and register the decoder using {@link registerDecoder}. * @group Custom coders * @public */ export abstract class CustomVideoDecoder { /** The input video's codec. */ readonly codec!: VideoCodec; /** The input video's decoder config. */ readonly config!: VideoDecoderConfig; /** The callback to call when a decoded VideoSample is available. */ readonly onSample!: (sample: VideoSample) => unknown; /** Returns true if and only if the decoder can decode the given codec configuration. */ // eslint-disable-next-line @typescript-eslint/no-unused-vars static supports(codec: VideoCodec, config: VideoDecoderConfig): boolean { return false; } /** Called after decoder creation; can be used for custom initialization logic. */ abstract init(): MaybePromise; /** Decodes the provided encoded packet. */ abstract decode(packet: EncodedPacket): MaybePromise; /** Decodes all remaining packets and then resolves. */ abstract flush(): MaybePromise; /** Called when the decoder is no longer needed and its resources can be freed. */ abstract close(): MaybePromise; } /** * Base class for custom audio decoders. To add your own custom audio decoder, extend this class, implement the * abstract methods and static `supports` method, and register the decoder using {@link registerDecoder}. * @group Custom coders * @public */ export abstract class CustomAudioDecoder { /** The input audio's codec. */ readonly codec!: AudioCodec; /** The input audio's decoder config. */ readonly config!: AudioDecoderConfig; /** The callback to call when a decoded AudioSample is available. */ readonly onSample!: (sample: AudioSample) => unknown; /** Returns true if and only if the decoder can decode the given codec configuration. */ // eslint-disable-next-line @typescript-eslint/no-unused-vars static supports(codec: AudioCodec, config: AudioDecoderConfig): boolean { return false; } /** Called after decoder creation; can be used for custom initialization logic. */ abstract init(): MaybePromise; /** Decodes the provided encoded packet. */ abstract decode(packet: EncodedPacket): MaybePromise; /** Decodes all remaining packets and then resolves. */ abstract flush(): MaybePromise; /** Called when the decoder is no longer needed and its resources can be freed. */ abstract close(): MaybePromise; } /** * Base class for custom video encoders. To add your own custom video encoder, extend this class, implement the * abstract methods and static `supports` method, and register the encoder using {@link registerEncoder}. * @group Custom coders * @public */ export abstract class CustomVideoEncoder { /** The codec with which to encode the video. */ readonly codec!: VideoCodec; /** Config for the encoder. */ readonly config!: VideoEncoderConfig; /** The callback to call when an EncodedPacket is available. */ readonly onPacket!: (packet: EncodedPacket, meta?: EncodedVideoChunkMetadata) => unknown; /** Returns true if and only if the encoder can encode the given codec configuration. */ // eslint-disable-next-line @typescript-eslint/no-unused-vars static supports(codec: VideoCodec, config: VideoEncoderConfig): boolean { return false; } /** Called after encoder creation; can be used for custom initialization logic. */ abstract init(): MaybePromise; /** Encodes the provided video sample. */ abstract encode(videoSample: VideoSample, options: VideoEncoderEncodeOptions): MaybePromise; /** Encodes all remaining video samples and then resolves. */ abstract flush(): MaybePromise; /** Called when the encoder is no longer needed and its resources can be freed. */ abstract close(): MaybePromise; } /** * Base class for custom audio encoders. To add your own custom audio encoder, extend this class, implement the * abstract methods and static `supports` method, and register the encoder using {@link registerEncoder}. * @group Custom coders * @public */ export abstract class CustomAudioEncoder { /** The codec with which to encode the audio. */ readonly codec!: AudioCodec; /** Config for the encoder. */ readonly config!: AudioEncoderConfig; /** The callback to call when an EncodedPacket is available. */ readonly onPacket!: (packet: EncodedPacket, meta?: EncodedAudioChunkMetadata) => unknown; /** Returns true if and only if the encoder can encode the given codec configuration. */ // eslint-disable-next-line @typescript-eslint/no-unused-vars static supports(codec: AudioCodec, config: AudioEncoderConfig): boolean { return false; } /** Called after encoder creation; can be used for custom initialization logic. */ abstract init(): MaybePromise; /** Encodes the provided audio sample. */ abstract encode(audioSample: AudioSample): MaybePromise; /** Encodes all remaining audio samples and then resolves. */ abstract flush(): MaybePromise; /** Called when the encoder is no longer needed and its resources can be freed. */ abstract close(): MaybePromise; } export const customVideoDecoders: typeof CustomVideoDecoder[] = []; export const customAudioDecoders: typeof CustomAudioDecoder[] = []; export const customVideoEncoders: typeof CustomVideoEncoder[] = []; export const customAudioEncoders: typeof CustomAudioEncoder[] = []; /** * Registers a custom video or audio decoder. Registered decoders will automatically be used for decoding whenever * possible. * @group Custom coders * @public */ export const registerDecoder = (decoder: typeof CustomVideoDecoder | typeof CustomAudioDecoder) => { if (decoder.prototype instanceof CustomVideoDecoder) { const casted = decoder as typeof CustomVideoDecoder; if (customVideoDecoders.includes(casted)) { Logging._warn('Video decoder already registered.'); return; } customVideoDecoders.push(casted); canDecodeVideoMemo.clear(); } else if (decoder.prototype instanceof CustomAudioDecoder) { const casted = decoder as typeof CustomAudioDecoder; if (customAudioDecoders.includes(casted)) { Logging._warn('Audio decoder already registered.'); return; } customAudioDecoders.push(casted); canDecodeAudioMemo.clear(); } else { throw new TypeError('Decoder must be a CustomVideoDecoder or CustomAudioDecoder.'); } }; /** * Registers a custom video or audio encoder. Registered encoders will automatically be used for encoding whenever * possible. * @group Custom coders * @public */ export const registerEncoder = (encoder: typeof CustomVideoEncoder | typeof CustomAudioEncoder) => { if (encoder.prototype instanceof CustomVideoEncoder) { const casted = encoder as typeof CustomVideoEncoder; if (customVideoEncoders.includes(casted)) { Logging._warn('Video encoder already registered.'); return; } customVideoEncoders.push(casted); canEncodeVideoMemo.clear(); } else if (encoder.prototype instanceof CustomAudioEncoder) { const casted = encoder as typeof CustomAudioEncoder; if (customAudioEncoders.includes(casted)) { Logging._warn('Audio encoder already registered.'); return; } customAudioEncoders.push(casted); canEncodeAudioMemo.clear(); } else { throw new TypeError('Encoder must be a CustomVideoEncoder or CustomAudioEncoder.'); } }; ===== src/matroska/matroska-misc.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ export const buildMatroskaMimeType = (info: { isWebM: boolean; hasVideo: boolean; hasAudio: boolean; codecStrings: string[]; }) => { const base = info.hasVideo ? 'video/' : info.hasAudio ? 'audio/' : 'application/'; let string = base + (info.isWebM ? 'webm' : 'x-matroska'); if (info.codecStrings.length > 0) { const uniqueCodecMimeTypes = [...new Set(info.codecStrings.filter(Boolean))]; string += `; codecs="${uniqueCodecMimeTypes.join(', ')}"`; } return string; }; ===== src/matroska/matroska-muxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Bitstream } from '../../shared/bitstream'; import { COLOR_PRIMARIES_MAP, MATRIX_COEFFICIENTS_MAP, TRANSFER_CHARACTERISTICS_MAP, UNDETERMINED_LANGUAGE, assert, assertNever, colorSpaceIsComplete, imageMimeTypeToExtension, keyValueIterator, normalizeRotation, promiseWithResolvers, Rational, simplifyRational, textEncoder, toUint8Array, uint8ArraysAreEqual, writeBits, roundToDivisor, } from '../misc'; import { CODEC_STRING_MAP, EBML, EBMLElement, EBMLFloat32, EBMLFloat64, EBMLId, EBMLSignedInt, EBMLUnicodeString, EBMLWriter, } from './ebml'; import { buildMatroskaMimeType } from './matroska-misc'; import { MkvOutputFormat, WebMOutputFormat } from '../output-format'; import { Output, OutputAudioTrack, OutputSubtitleTrack, OutputTrack, OutputVideoTrack } from '../output'; import { SubtitleConfig, SubtitleCue, SubtitleMetadata, formatSubtitleTimestamp, inlineTimestampRegex, parseSubtitleTimestamp, } from '../subtitles'; import { aacChannelMap, aacFrequencyTable, buildAacAudioSpecificConfig } from '../../shared/aac-misc'; import { OPUS_SAMPLE_RATE, PCM_AUDIO_CODECS, PcmAudioCodec, SubtitleCodec, generateAv1CodecConfigurationFromCodecString, generateVp9CodecConfigurationFromCodecString, parsePcmCodec, validateAudioChunkMetadata, validateSubtitleMetadata, validateVideoChunkMetadata, } from '../codec'; import { MAX_ADTS_FRAME_HEADER_SIZE, MIN_ADTS_FRAME_HEADER_SIZE, readAdtsFrameHeader } from '../adts/adts-reader'; import { FileSlice } from '../reader'; import { Muxer } from '../muxer'; import { Writer } from '../writer'; import { EncodedPacket } from '../packet'; import { parseOpusIdentificationHeader } from '../codec-data'; import { AttachedFile } from '../metadata'; const MIN_CLUSTER_TIMESTAMP_MS = -(2 ** 15); const MAX_CLUSTER_TIMESTAMP_MS = 2 ** 15 - 1; const APP_NAME = 'Mediabunny'; const SEGMENT_SIZE_BYTES = 6; const CLUSTER_SIZE_BYTES = 5; type InternalMediaChunk = { data: Uint8Array; type: 'key' | 'delta'; timestamp: number; duration: number; additions: Uint8Array | null; }; type MatroskaTrackData = { chunkQueue: InternalMediaChunk[]; lastWrittenMsTimestamp: number | null; closed: boolean; } & ({ track: OutputVideoTrack; type: 'video'; info: { width: number; height: number; aspectRatio: Rational | null; decoderConfig: VideoDecoderConfig; alphaMode: boolean; }; } | { track: OutputAudioTrack; type: 'audio'; info: { numberOfChannels: number; sampleRate: number; decoderConfig: AudioDecoderConfig; /** * The "ADTS stripping" involves removing the ADTS header from each AAC packet. SOBMFF stores raw AAC data, not * ADTS-wrapped data. */ requiresAdtsStripping: boolean; }; } | { track: OutputSubtitleTrack; type: 'subtitle'; info: { config: SubtitleConfig; }; }); type MatroskaVideoTrackData = MatroskaTrackData & { type: 'video' }; type MatroskaAudioTrackData = MatroskaTrackData & { type: 'audio' }; type MatroskaSubtitleTrackData = MatroskaTrackData & { type: 'subtitle' }; const TRACK_TYPE_MAP: Record = { video: 1, audio: 2, subtitle: 17, }; export class MatroskaMuxer extends Muxer { private writer!: Writer; private ebmlWriter!: EBMLWriter; private format: WebMOutputFormat | MkvOutputFormat; private trackDatas: MatroskaTrackData[] = []; private allTracksKnown = promiseWithResolvers(); private segment: EBMLElement | null = null; private segmentInfo: EBMLElement | null = null; private seekHead: EBMLElement | null = null; private tracksElement: EBMLElement | null = null; private tagsElement: EBMLElement | null = null; private attachmentsElement: EBMLElement | null = null; private segmentDuration: EBMLElement | null = null; private cues: EBMLElement | null = null; private currentCluster: EBMLElement | null = null; private currentClusterStartMsTimestamp: number | null = null; private currentClusterMaxMsTimestamp: number | null = null; private trackDatasInCurrentCluster = new Map(); private startTimestamp = Infinity; private endTimestamp = -Infinity; constructor(output: Output, format: MkvOutputFormat) { super(output); this.format = format; } async start() { const release = await this.mutex.acquire(); this.writer = await this.output._getRootWriter(!!this.format._options.appendOnly); this.ebmlWriter = new EBMLWriter(this.writer); this.writeEBMLHeader(); this.createSegmentInfo(); this.createCues(); await this.writer.flush(); release(); } private writeEBMLHeader() { if (this.format._options.onEbmlHeader) { this.writer.startTrackingWrites(); } const ebmlHeader: EBML = { id: EBMLId.EBML, data: [ { id: EBMLId.EBMLVersion, data: 1 }, { id: EBMLId.EBMLReadVersion, data: 1 }, { id: EBMLId.EBMLMaxIDLength, data: 4 }, { id: EBMLId.EBMLMaxSizeLength, data: 8 }, { id: EBMLId.DocType, data: this.format instanceof WebMOutputFormat ? 'webm' : 'matroska' }, { id: EBMLId.DocTypeVersion, data: 2 }, { id: EBMLId.DocTypeReadVersion, data: 2 }, ] }; this.ebmlWriter.writeEBML(ebmlHeader); if (this.format._options.onEbmlHeader) { const { data, start } = this.writer.stopTrackingWrites(); // start should be 0 this.format._options.onEbmlHeader(data, start); } } /** * Creates a SeekHead element which is positioned near the start of the file and allows the media player to seek to * relevant sections more easily. Since we don't know the positions of those sections yet, we'll set them later. */ private maybeCreateSeekHead(writeOffsets: boolean) { if (this.format._options.appendOnly) { return; } const kaxCues = new Uint8Array([0x1c, 0x53, 0xbb, 0x6b]); const kaxInfo = new Uint8Array([0x15, 0x49, 0xa9, 0x66]); const kaxTracks = new Uint8Array([0x16, 0x54, 0xae, 0x6b]); const kaxAttachments = new Uint8Array([0x19, 0x41, 0xa4, 0x69]); const kaxTags = new Uint8Array([0x12, 0x54, 0xc3, 0x67]); const seekHead = { id: EBMLId.SeekHead, data: [ { id: EBMLId.Seek, data: [ { id: EBMLId.SeekID, data: kaxCues }, { id: EBMLId.SeekPosition, size: 5, data: writeOffsets ? this.ebmlWriter.offsets.get(this.cues!)! - this.segmentDataOffset : 0, }, ] }, { id: EBMLId.Seek, data: [ { id: EBMLId.SeekID, data: kaxInfo }, { id: EBMLId.SeekPosition, size: 5, data: writeOffsets ? this.ebmlWriter.offsets.get(this.segmentInfo!)! - this.segmentDataOffset : 0, }, ] }, { id: EBMLId.Seek, data: [ { id: EBMLId.SeekID, data: kaxTracks }, { id: EBMLId.SeekPosition, size: 5, data: writeOffsets ? this.ebmlWriter.offsets.get(this.tracksElement!)! - this.segmentDataOffset : 0, }, ] }, this.attachmentsElement ? { id: EBMLId.Seek, data: [ { id: EBMLId.SeekID, data: kaxAttachments }, { id: EBMLId.SeekPosition, size: 5, data: writeOffsets ? this.ebmlWriter.offsets.get(this.attachmentsElement)! - this.segmentDataOffset : 0, }, ] } : null, this.tagsElement ? { id: EBMLId.Seek, data: [ { id: EBMLId.SeekID, data: kaxTags }, { id: EBMLId.SeekPosition, size: 5, data: writeOffsets ? this.ebmlWriter.offsets.get(this.tagsElement)! - this.segmentDataOffset : 0, }, ] } : null, ] }; this.seekHead = seekHead; } private createSegmentInfo() { const segmentDuration: EBML = { id: EBMLId.Duration, data: new EBMLFloat64(0) }; this.segmentDuration = segmentDuration; const segmentInfo: EBML = { id: EBMLId.Info, data: [ { id: EBMLId.TimestampScale, data: 1e6 }, { id: EBMLId.MuxingApp, data: APP_NAME }, { id: EBMLId.WritingApp, data: APP_NAME }, !this.format._options.appendOnly ? segmentDuration : null, ] }; this.segmentInfo = segmentInfo; } private createTracks() { const tracksElement = { id: EBMLId.Tracks, data: [] as EBML[] }; this.tracksElement = tracksElement; for (const trackData of this.trackDatas) { const codecId = CODEC_STRING_MAP[trackData.track.source._codec]; assert(codecId); let seekPreRollNs = 0; if (trackData.type === 'audio' && trackData.track.source._codec === 'opus') { seekPreRollNs = 1e6 * 80; // In "Matroska ticks" (nanoseconds) const description = trackData.info.decoderConfig.description; if (description) { const bytes = toUint8Array(description); const header = parseOpusIdentificationHeader(bytes); // Use the preSkip value from the header seekPreRollNs = Math.round(1e9 * (header.preSkip / OPUS_SAMPLE_RATE)); } } tracksElement.data.push({ id: EBMLId.TrackEntry, data: [ { id: EBMLId.TrackNumber, data: trackData.track.id }, { id: EBMLId.TrackUID, data: trackData.track.id }, { id: EBMLId.TrackType, data: TRACK_TYPE_MAP[trackData.type] }, trackData.track.metadata.disposition?.default === false ? { id: EBMLId.FlagDefault, data: 0 } : null, trackData.track.metadata.disposition?.forced ? { id: EBMLId.FlagForced, data: 1 } : null, trackData.track.metadata.disposition?.hearingImpaired ? { id: EBMLId.FlagHearingImpaired, data: 1 } : null, trackData.track.metadata.disposition?.visuallyImpaired ? { id: EBMLId.FlagVisualImpaired, data: 1 } : null, trackData.track.metadata.disposition?.original ? { id: EBMLId.FlagOriginal, data: 1 } : null, trackData.track.metadata.disposition?.commentary ? { id: EBMLId.FlagCommentary, data: 1 } : null, { id: EBMLId.FlagLacing, data: 0 }, { id: EBMLId.Language, data: trackData.track.metadata.languageCode ?? UNDETERMINED_LANGUAGE }, { id: EBMLId.CodecID, data: codecId }, { id: EBMLId.CodecDelay, data: 0 }, { id: EBMLId.SeekPreRoll, data: seekPreRollNs }, trackData.track.metadata.name !== undefined ? { id: EBMLId.Name, data: new EBMLUnicodeString(trackData.track.metadata.name) } : null, (trackData.type === 'video' ? this.videoSpecificTrackInfo(trackData) : null), (trackData.type === 'audio' ? this.audioSpecificTrackInfo(trackData) : null), (trackData.type === 'subtitle' ? this.subtitleSpecificTrackInfo(trackData) : null), ] }); } } private videoSpecificTrackInfo(trackData: MatroskaVideoTrackData) { const { frameRate, rotation } = trackData.track.metadata; const elements: EBMLElement['data'] = [ (trackData.info.decoderConfig.description ? { id: EBMLId.CodecPrivate, data: toUint8Array(trackData.info.decoderConfig.description), } : null), (frameRate ? { id: EBMLId.DefaultDuration, data: 1e9 / frameRate, } : null), ]; // Convert from clockwise to counter-clockwise const flippedRotation = rotation ? normalizeRotation(-rotation) : 0; const hasNonSquarePixelAspectRatio = !!trackData.info.aspectRatio && ( trackData.info.aspectRatio.num * trackData.info.height !== trackData.info.aspectRatio.den * trackData.info.width ); const colorSpace = trackData.info.decoderConfig.colorSpace; const videoElement: EBMLElement = { id: EBMLId.Video, data: [ { id: EBMLId.PixelWidth, data: trackData.info.width }, { id: EBMLId.PixelHeight, data: trackData.info.height }, (hasNonSquarePixelAspectRatio ? { id: EBMLId.DisplayWidth, data: trackData.info.aspectRatio!.num } : null), (hasNonSquarePixelAspectRatio ? { id: EBMLId.DisplayHeight, data: trackData.info.aspectRatio!.den } : null), (hasNonSquarePixelAspectRatio ? { id: EBMLId.DisplayUnit, data: 3 } : null), // 3 = display aspect ratio trackData.info.alphaMode ? { id: EBMLId.AlphaMode, data: 1 } : null, (colorSpaceIsComplete(colorSpace) ? { id: EBMLId.Colour, data: [ { id: EBMLId.MatrixCoefficients, data: MATRIX_COEFFICIENTS_MAP[colorSpace.matrix], }, { id: EBMLId.TransferCharacteristics, data: TRANSFER_CHARACTERISTICS_MAP[colorSpace.transfer], }, { id: EBMLId.Primaries, data: COLOR_PRIMARIES_MAP[colorSpace.primaries], }, { id: EBMLId.Range, data: colorSpace.fullRange ? 2 : 1, }, ], } : null), (flippedRotation ? { id: EBMLId.Projection, data: [ { id: EBMLId.ProjectionType, data: 0, // rectangular }, { id: EBMLId.ProjectionPoseRoll, data: new EBMLFloat32((flippedRotation + 180) % 360 - 180), // [0, 270] -> [-180, 90] }, ], } : null), ] }; elements.push(videoElement); return elements; } private audioSpecificTrackInfo(trackData: MatroskaAudioTrackData) { const pcmInfo = (PCM_AUDIO_CODECS as readonly string[]).includes(trackData.track.source._codec) ? parsePcmCodec(trackData.track.source._codec as PcmAudioCodec) : null; return [ (trackData.info.decoderConfig.description ? { id: EBMLId.CodecPrivate, data: toUint8Array(trackData.info.decoderConfig.description), } : null), { id: EBMLId.Audio, data: [ { id: EBMLId.SamplingFrequency, data: new EBMLFloat32(trackData.info.sampleRate) }, { id: EBMLId.Channels, data: trackData.info.numberOfChannels }, pcmInfo ? { id: EBMLId.BitDepth, data: 8 * pcmInfo.sampleSize } : null, ] }, ]; } private subtitleSpecificTrackInfo(trackData: MatroskaSubtitleTrackData) { return [ { id: EBMLId.CodecPrivate, data: textEncoder.encode(trackData.info.config.description) }, ]; } private maybeCreateTags() { const simpleTags: EBMLElement[] = []; const addSimpleTag = (key: string, value: string | Uint8Array) => { simpleTags.push({ id: EBMLId.SimpleTag, data: [ { id: EBMLId.TagName, data: new EBMLUnicodeString(key) }, typeof value === 'string' ? { id: EBMLId.TagString, data: new EBMLUnicodeString(value) } : { id: EBMLId.TagBinary, data: value }, ] }); }; const metadataTags = this.output._metadataTags; const writtenTags = new Set(); for (const { key, value } of keyValueIterator(metadataTags)) { switch (key) { case 'title': { addSimpleTag('TITLE', value); writtenTags.add('TITLE'); }; break; case 'description': { addSimpleTag('DESCRIPTION', value); writtenTags.add('DESCRIPTION'); }; break; case 'artist': { addSimpleTag('ARTIST', value); writtenTags.add('ARTIST'); }; break; case 'album': { addSimpleTag('ALBUM', value); writtenTags.add('ALBUM'); }; break; case 'albumArtist': { addSimpleTag('ALBUM_ARTIST', value); writtenTags.add('ALBUM_ARTIST'); }; break; case 'genre': { addSimpleTag('GENRE', value); writtenTags.add('GENRE'); }; break; case 'comment': { addSimpleTag('COMMENT', value); writtenTags.add('COMMENT'); }; break; case 'lyrics': { addSimpleTag('LYRICS', value); writtenTags.add('LYRICS'); }; break; case 'date': { addSimpleTag('DATE', value.toISOString().slice(0, 10)); writtenTags.add('DATE'); }; break; case 'trackNumber': { const string = metadataTags.tracksTotal !== undefined ? `${value}/${metadataTags.tracksTotal}` : value.toString(); addSimpleTag('PART_NUMBER', string); writtenTags.add('PART_NUMBER'); }; break; case 'discNumber': { const string = metadataTags.discsTotal !== undefined ? `${value}/${metadataTags.discsTotal}` : value.toString(); addSimpleTag('DISC', string); writtenTags.add('DISC'); }; break; case 'tracksTotal': case 'discsTotal': { // Handled with trackNumber and discNumber respectively }; break; case 'images': case 'raw': { // Handled elsewhere }; break; default: assertNever(key); } } if (metadataTags.raw) { for (const key in metadataTags.raw) { const value = metadataTags.raw[key]!; if (value == null || writtenTags.has(key)) { continue; } if (typeof value === 'string' || value instanceof Uint8Array) { addSimpleTag(key, value); } } } if (simpleTags.length === 0) { return; } this.tagsElement = { id: EBMLId.Tags, data: [{ id: EBMLId.Tag, data: [ { id: EBMLId.Targets, data: [ { id: EBMLId.TargetTypeValue, data: 50 }, { id: EBMLId.TargetType, data: 'MOVIE' }, ] }, ...simpleTags, ] }], }; } private maybeCreateAttachments() { const metadataTags = this.output._metadataTags; const elements: EBMLElement[] = []; const existingFileUids = new Set(); const images = metadataTags.images ?? []; for (const image of images) { let imageName = image.name; if (imageName === undefined) { const baseName = image.kind === 'coverFront' ? 'cover' : image.kind === 'coverBack' ? 'back' : 'image'; imageName = baseName + (imageMimeTypeToExtension(image.mimeType) ?? ''); } let fileUid: bigint; while (true) { // Generate a random 64-bit unsigned integer fileUid = 0n; for (let i = 0; i < 8; i++) { fileUid <<= 8n; fileUid |= BigInt(Math.floor(Math.random() * 256)); } if (fileUid !== 0n && !existingFileUids.has(fileUid)) { break; } } existingFileUids.add(fileUid); elements.push({ id: EBMLId.AttachedFile, data: [ image.description !== undefined ? { id: EBMLId.FileDescription, data: new EBMLUnicodeString(image.description) } : null, { id: EBMLId.FileName, data: new EBMLUnicodeString(imageName) }, { id: EBMLId.FileMediaType, data: image.mimeType }, { id: EBMLId.FileData, data: image.data }, { id: EBMLId.FileUID, data: fileUid }, ], }); } // Add all AttachedFiles from the raw metadata for (const [key, value] of Object.entries(metadataTags.raw ?? {})) { if (!(value instanceof AttachedFile)) { continue; } const keyIsNumeric = /^\d+$/.test(key); if (!keyIsNumeric) { continue; } if (images.find(x => x.mimeType === value.mimeType && uint8ArraysAreEqual(x.data, value.data))) { // This attached file has very likely already been added as an image above // (happens when remuxing Matroska) continue; } elements.push({ id: EBMLId.AttachedFile, data: [ value.description !== undefined ? { id: EBMLId.FileDescription, data: new EBMLUnicodeString(value.description) } : null, { id: EBMLId.FileName, data: new EBMLUnicodeString(value.name ?? '') }, { id: EBMLId.FileMediaType, data: value.mimeType ?? '' }, { id: EBMLId.FileData, data: value.data }, { id: EBMLId.FileUID, data: BigInt(key) }, ], }); } if (elements.length === 0) { return; } this.attachmentsElement = { id: EBMLId.Attachments, data: elements }; } private createSegment() { this.createTracks(); this.maybeCreateTags(); this.maybeCreateAttachments(); this.maybeCreateSeekHead(false); const segment: EBML = { id: EBMLId.Segment, size: this.format._options.appendOnly ? -1 : SEGMENT_SIZE_BYTES, data: [ this.seekHead, // null if append-only this.segmentInfo, this.tracksElement, // Matroska spec says put this at the end of the file, but I think placing it before the first cluster // makes more sense, and FFmpeg agrees (argumentum ad ffmpegum fallacy) this.attachmentsElement, this.tagsElement, ], }; this.segment = segment; if (this.format._options.onSegmentHeader) { this.writer.startTrackingWrites(); } this.ebmlWriter.writeEBML(segment); if (this.format._options.onSegmentHeader) { const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onSegmentHeader(data, start); } } private createCues() { this.cues = { id: EBMLId.Cues, data: [] }; } private get segmentDataOffset() { assert(this.segment); return this.ebmlWriter.dataOffsets.get(this.segment)!; } private allTracksAreKnown() { for (const track of this.output._tracks) { if (!track.source._closed && !this.trackDatas.some(x => x.track === track)) { return false; // We haven't seen a sample from this open track yet } } return true; } async getMimeType() { await this.allTracksKnown.promise; const codecStrings = this.trackDatas.map((trackData) => { if (trackData.type === 'video') { return trackData.info.decoderConfig.codec; } else if (trackData.type === 'audio') { return trackData.info.decoderConfig.codec; } else { const map: Record = { webvtt: 'wvtt', }; return map[trackData.track.source._codec]; } }); return buildMatroskaMimeType({ isWebM: this.format instanceof WebMOutputFormat, hasVideo: this.trackDatas.some(x => x.type === 'video'), hasAudio: this.trackDatas.some(x => x.type === 'audio'), codecStrings, }); } private getVideoTrackData(track: OutputVideoTrack, packet: EncodedPacket, meta?: EncodedVideoChunkMetadata) { const existingTrackData = this.trackDatas.find(x => x.track === track); if (existingTrackData) { return existingTrackData as MatroskaVideoTrackData; } validateVideoChunkMetadata(meta); assert(meta); assert(meta.decoderConfig); assert(meta.decoderConfig.codedWidth !== undefined); assert(meta.decoderConfig.codedHeight !== undefined); const displayAspectWidth = meta.decoderConfig.displayAspectWidth; const displayAspectHeight = meta.decoderConfig.displayAspectHeight; const aspectRatio = displayAspectWidth === undefined || displayAspectHeight === undefined ? null : simplifyRational({ num: displayAspectWidth, den: displayAspectHeight, }); const newTrackData: MatroskaVideoTrackData = { track, type: 'video', info: { width: meta.decoderConfig.codedWidth, height: meta.decoderConfig.codedHeight, aspectRatio, decoderConfig: meta.decoderConfig, alphaMode: !!packet.sideData.alpha, // The first packet determines if this track has alpha or not }, chunkQueue: [], lastWrittenMsTimestamp: null, closed: false, }; if (track.source._codec === 'vp9') { // https://www.webmproject.org/docs/container specifies that VP9 "SHOULD" make use of the CodecPrivate // field. Since WebCodecs makes no use of the description field for VP9, we need to derive it ourselves: newTrackData.info.decoderConfig = { ...newTrackData.info.decoderConfig, description: new Uint8Array( generateVp9CodecConfigurationFromCodecString(newTrackData.info.decoderConfig.codec), ), }; } else if (track.source._codec === 'av1') { // Per https://github.com/ietf-wg-cellar/matroska-specification/blob/master/codec/av1.md, AV1 requires // CodecPrivate to be set, but WebCodecs makes no use of the description field for AV1. Thus, let's derive // it ourselves: newTrackData.info.decoderConfig = { ...newTrackData.info.decoderConfig, description: new Uint8Array( generateAv1CodecConfigurationFromCodecString(newTrackData.info.decoderConfig.codec), ), }; } this.trackDatas.push(newTrackData); this.trackDatas.sort((a, b) => a.track.id - b.track.id); if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } return newTrackData; } private getAudioTrackData(track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata) { const existingTrackData = this.trackDatas.find(x => x.track === track); if (existingTrackData) { return existingTrackData as MatroskaAudioTrackData; } validateAudioChunkMetadata(meta); assert(meta); assert(meta.decoderConfig); const decoderConfig = { ...meta.decoderConfig }; let requiresAdtsStripping = false; if (track.source._codec === 'aac' && !decoderConfig.description) { // Matroska stores raw AAC with AudioSpecificConfig in CodecPrivate, not ADTS-wrapped data. // Parse the first packet to extract the AudioSpecificConfig. const adtsFrame = readAdtsFrameHeader(FileSlice.tempFromBytes(packet.data)); if (!adtsFrame) { throw new Error( 'Couldn\'t parse ADTS header from the AAC packet. Make sure the packets are in ADTS format' + ' (as specified in ISO 13818-7) when not providing a description, or provide a description' + ' (must be an AudioSpecificConfig as specified in ISO 14496-3) and ensure the packets' + ' are raw AAC data.', ); } const sampleRate = aacFrequencyTable[adtsFrame.samplingFrequencyIndex]; const numberOfChannels = aacChannelMap[adtsFrame.channelConfiguration]; if (sampleRate === undefined || numberOfChannels === undefined) { throw new Error('Invalid ADTS frame header.'); } decoderConfig.description = buildAacAudioSpecificConfig({ objectType: adtsFrame.objectType, sampleRate, numberOfChannels, }); requiresAdtsStripping = true; } const newTrackData: MatroskaAudioTrackData = { track, type: 'audio', info: { numberOfChannels: meta.decoderConfig.numberOfChannels, sampleRate: meta.decoderConfig.sampleRate, decoderConfig, requiresAdtsStripping, }, chunkQueue: [], lastWrittenMsTimestamp: null, closed: false, }; this.trackDatas.push(newTrackData); this.trackDatas.sort((a, b) => a.track.id - b.track.id); if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } return newTrackData; } private getSubtitleTrackData(track: OutputSubtitleTrack, meta?: SubtitleMetadata) { const existingTrackData = this.trackDatas.find(x => x.track === track); if (existingTrackData) { return existingTrackData as MatroskaAudioTrackData; } validateSubtitleMetadata(meta); assert(meta); assert(meta.config); const newTrackData: MatroskaSubtitleTrackData = { track, type: 'subtitle', info: { config: meta.config, }, chunkQueue: [], lastWrittenMsTimestamp: null, closed: false, }; this.trackDatas.push(newTrackData); this.trackDatas.sort((a, b) => a.track.id - b.track.id); if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } return newTrackData; } async addEncodedVideoPacket(track: OutputVideoTrack, packet: EncodedPacket, meta?: EncodedVideoChunkMetadata) { const release = await this.mutex.acquire(); try { const trackData = this.getVideoTrackData(track, packet, meta); const isKeyFrame = packet.type === 'key'; this.validateTimestamp(trackData.track, packet.timestamp, isKeyFrame); let timestamp = packet.timestamp; let duration = packet.duration; if (track.metadata.frameRate !== undefined) { // Constrain the time values to the frame rate timestamp = roundToDivisor(timestamp, track.metadata.frameRate); duration = roundToDivisor(duration, track.metadata.frameRate); } const additions = trackData.info.alphaMode ? packet.sideData.alpha ?? null : null; const videoChunk = this.createInternalChunk(packet.data, timestamp, duration, packet.type, additions); if (track.source._codec === 'vp9') this.fixVP9ColorSpace(trackData, videoChunk); trackData.chunkQueue.push(videoChunk); await this.interleaveChunks(); } finally { release(); } } async addEncodedAudioPacket(track: OutputAudioTrack, packet: EncodedPacket, meta?: EncodedAudioChunkMetadata) { const release = await this.mutex.acquire(); try { const trackData = this.getAudioTrackData(track, packet, meta); let packetData = packet.data; if (trackData.info.requiresAdtsStripping) { const adtsFrame = readAdtsFrameHeader(FileSlice.tempFromBytes(packetData)); if (!adtsFrame) { throw new Error('Expected ADTS frame, didn\'t get one.'); } const headerLength = adtsFrame.crcCheck === null ? MIN_ADTS_FRAME_HEADER_SIZE : MAX_ADTS_FRAME_HEADER_SIZE; packetData = packetData.subarray(headerLength); } const isKeyFrame = packet.type === 'key'; this.validateTimestamp(trackData.track, packet.timestamp, isKeyFrame); const audioChunk = this.createInternalChunk(packetData, packet.timestamp, packet.duration, packet.type); trackData.chunkQueue.push(audioChunk); await this.interleaveChunks(); } finally { release(); } } async addSubtitleCue(track: OutputSubtitleTrack, cue: SubtitleCue, meta?: SubtitleMetadata) { const release = await this.mutex.acquire(); try { const trackData = this.getSubtitleTrackData(track, meta); this.validateTimestamp(trackData.track, cue.timestamp, true); let bodyText = cue.text; const timestampMs = Math.round(cue.timestamp * 1000); // Replace in-body timestamps so that they're relative to the cue start time inlineTimestampRegex.lastIndex = 0; bodyText = bodyText.replace(inlineTimestampRegex, (match) => { const time = parseSubtitleTimestamp(match.slice(1, -1)); const offsetTime = time - timestampMs; return `<${formatSubtitleTimestamp(offsetTime)}>`; }); const body = textEncoder.encode(bodyText); const additions = `${cue.settings ?? ''}\n${cue.identifier ?? ''}\n${cue.notes ?? ''}`; const subtitleChunk = this.createInternalChunk( body, cue.timestamp, cue.duration, 'key', additions.trim() ? textEncoder.encode(additions) : null, ); trackData.chunkQueue.push(subtitleChunk); await this.interleaveChunks(); } finally { release(); } } private async interleaveChunks(isFinalCall = false) { if (!isFinalCall && !this.allTracksAreKnown()) { return; // We can't interleave yet as we don't yet know how many tracks we'll truly have } outer: while (true) { let trackWithMinTimestamp: MatroskaTrackData | null = null; let minTimestamp = Infinity; for (const trackData of this.trackDatas) { if (!isFinalCall && trackData.chunkQueue.length === 0 && !trackData.closed) { break outer; } if (trackData.chunkQueue.length > 0 && trackData.chunkQueue[0]!.timestamp < minTimestamp) { trackWithMinTimestamp = trackData; minTimestamp = trackData.chunkQueue[0]!.timestamp; } } if (!trackWithMinTimestamp) { break; } const chunk = trackWithMinTimestamp.chunkQueue.shift()!; this.writeBlock(trackWithMinTimestamp, chunk); } if (!isFinalCall) { await this.writer.flush(); } } /** * Due to [a bug in Chromium](https://bugs.chromium.org/p/chromium/issues/detail?id=1377842), VP9 streams often * lack color space information. This method patches in that information. */ private fixVP9ColorSpace( trackData: MatroskaVideoTrackData, chunk: InternalMediaChunk, ) { // http://downloads.webmproject.org/docs/vp9/vp9-bitstream_superframe-and-uncompressed-header_v1.0.pdf if (chunk.type !== 'key') return; if (!trackData.info.decoderConfig.colorSpace || !trackData.info.decoderConfig.colorSpace.matrix) return; const bitstream = new Bitstream(chunk.data); bitstream.skipBits(2); const profileLowBit = bitstream.readBits(1); const profileHighBit = bitstream.readBits(1); const profile = (profileHighBit << 1) + profileLowBit; if (profile === 3) bitstream.skipBits(1); const showExistingFrame = bitstream.readBits(1); if (showExistingFrame) return; const frameType = bitstream.readBits(1); if (frameType !== 0) return; // Just to be sure bitstream.skipBits(2); const syncCode = bitstream.readBits(24); if (syncCode !== 0x498342) return; if (profile >= 2) bitstream.skipBits(1); const colorSpaceID = { rgb: 7, bt709: 2, bt470bg: 1, smpte170m: 3, }[trackData.info.decoderConfig.colorSpace.matrix]; // The bitstream position is now at the start of the color space bits. // We can use the global writeBits function here as requested. writeBits(chunk.data, bitstream.pos, bitstream.pos + 3, colorSpaceID); } /** Converts a read-only external chunk into an internal one for easier use. */ private createInternalChunk( data: Uint8Array, timestamp: number, duration: number, type: 'key' | 'delta', additions: Uint8Array | null = null, ) { const internalChunk: InternalMediaChunk = { data, type, timestamp, duration, additions, }; return internalChunk; } /** Writes a block containing media data to the file. */ private writeBlock(trackData: MatroskaTrackData, chunk: InternalMediaChunk) { // Due to the interlacing algorithm, this code will be run once we've seen one chunk from every media track. if (!this.segment) { this.createSegment(); } const msTimestamp = Math.round(1000 * chunk.timestamp); // We wanna only finalize this cluster (and begin a new one) if we know that each track will be able to // start the new one with a key frame. const keyFrameQueuedEverywhere = this.trackDatas.every((otherTrackData) => { if (trackData === otherTrackData) { return chunk.type === 'key'; } const firstQueuedSample = otherTrackData.chunkQueue[0]; if (firstQueuedSample) { return firstQueuedSample.type === 'key'; } return otherTrackData.closed; }); let shouldCreateNewCluster = false; if (!this.currentCluster) { shouldCreateNewCluster = true; } else { assert(this.currentClusterStartMsTimestamp !== null); assert(this.currentClusterMaxMsTimestamp !== null); const relativeTimestamp = msTimestamp - this.currentClusterStartMsTimestamp; shouldCreateNewCluster = ( keyFrameQueuedEverywhere // This check is required because that means there is already a block with this timestamp in the // CURRENT chunk, meaning that starting the next cluster at the same timestamp is forbidden (since // the already-written block would belong into it instead). && msTimestamp > this.currentClusterMaxMsTimestamp && relativeTimestamp >= 1000 * (this.format._options.minimumClusterDuration ?? 1) ) // The cluster would exceed its maximum allowed length. This puts us in an unfortunate position and forces // us to begin the next cluster with a delta frame. Although this is undesirable, it is not forbidden by the // spec and is supported by players. || relativeTimestamp > MAX_CLUSTER_TIMESTAMP_MS; } if (shouldCreateNewCluster) { this.createNewCluster(msTimestamp); } const relativeTimestamp = msTimestamp - this.currentClusterStartMsTimestamp!; if (relativeTimestamp < MIN_CLUSTER_TIMESTAMP_MS) { // The block lies too far in the past, it's not representable within this cluster return; } const prelude = new Uint8Array(4); const view = new DataView(prelude.buffer); // 0x80 to indicate it's the last byte of a multi-byte number view.setUint8(0, 0x80 | trackData.track.id); view.setInt16(1, relativeTimestamp, false); const msDuration = Math.round(1000 * chunk.duration); if (!chunk.additions) { // No additions, we can write out a SimpleBlock view.setUint8(3, Number(chunk.type === 'key') << 7); // Flags (keyframe flag only present for SimpleBlock) const simpleBlock = { id: EBMLId.SimpleBlock, data: [ prelude, chunk.data, ] }; this.ebmlWriter.writeEBML(simpleBlock); } else { const blockGroup = { id: EBMLId.BlockGroup, data: [ { id: EBMLId.Block, data: [ prelude, chunk.data, ] }, chunk.type === 'delta' ? { id: EBMLId.ReferenceBlock, data: new EBMLSignedInt(trackData.lastWrittenMsTimestamp! - msTimestamp), } : null, chunk.additions ? { id: EBMLId.BlockAdditions, data: [ { id: EBMLId.BlockMore, data: [ { id: EBMLId.BlockAddID, data: 1 }, // Some players expect BlockAddID to come first { id: EBMLId.BlockAdditional, data: chunk.additions }, ] }, ] } : null, msDuration > 0 ? { id: EBMLId.BlockDuration, data: msDuration } : null, ] }; this.ebmlWriter.writeEBML(blockGroup); } this.startTimestamp = Math.min(this.startTimestamp, msTimestamp); this.endTimestamp = Math.max(this.endTimestamp, msTimestamp + msDuration); trackData.lastWrittenMsTimestamp = msTimestamp; if (!this.trackDatasInCurrentCluster.has(trackData)) { this.trackDatasInCurrentCluster.set(trackData, { firstMsTimestamp: msTimestamp, }); } this.currentClusterMaxMsTimestamp = Math.max(this.currentClusterMaxMsTimestamp!, msTimestamp); } /** Creates a new Cluster element to contain media chunks. */ private createNewCluster(msTimestamp: number) { if (this.currentCluster) { this.finalizeCurrentCluster(); } if (this.format._options.onCluster) { this.writer.startTrackingWrites(); } this.currentCluster = { id: EBMLId.Cluster, size: this.format._options.appendOnly ? -1 : CLUSTER_SIZE_BYTES, data: [ { id: EBMLId.Timestamp, data: msTimestamp }, ], }; this.ebmlWriter.writeEBML(this.currentCluster); this.currentClusterStartMsTimestamp = msTimestamp; this.currentClusterMaxMsTimestamp = msTimestamp; this.trackDatasInCurrentCluster.clear(); } private finalizeCurrentCluster() { assert(this.currentCluster); if (!this.format._options.appendOnly) { const clusterSize = this.writer.getPos() - this.ebmlWriter.dataOffsets.get(this.currentCluster)!; const endPos = this.writer.getPos(); // Write the size now that we know it this.writer.seek(this.ebmlWriter.offsets.get(this.currentCluster)! + 4); this.ebmlWriter.writeVarInt(clusterSize, CLUSTER_SIZE_BYTES); this.writer.seek(endPos); } if (this.format._options.onCluster) { assert(this.currentClusterStartMsTimestamp !== null); const { data, start } = this.writer.stopTrackingWrites(); this.format._options.onCluster(data, start, this.currentClusterStartMsTimestamp / 1000); } const clusterOffsetFromSegment = this.ebmlWriter.offsets.get(this.currentCluster)! - this.segmentDataOffset; // Group tracks by their first timestamp and create a CuePoint for each unique timestamp const groupedByTimestamp = new Map(); for (const [trackData, { firstMsTimestamp }] of this.trackDatasInCurrentCluster) { if (!groupedByTimestamp.has(firstMsTimestamp)) { groupedByTimestamp.set(firstMsTimestamp, []); } groupedByTimestamp.get(firstMsTimestamp)!.push(trackData); } const groupedAndSortedByTimestamp = [...groupedByTimestamp.entries()].sort((a, b) => a[0] - b[0]); // Add CuePoints to the Cues element for better seeking for (const [msTimestamp, trackDatas] of groupedAndSortedByTimestamp) { assert(this.cues); (this.cues.data as EBML[]).push({ id: EBMLId.CuePoint, data: [ { id: EBMLId.CueTime, data: msTimestamp }, // Create CueTrackPositions for each track that starts at this timestamp ...trackDatas.map((trackData) => { return { id: EBMLId.CueTrackPositions, data: [ { id: EBMLId.CueTrack, data: trackData.track.id }, { id: EBMLId.CueClusterPosition, data: clusterOffsetFromSegment }, ] }; }), ] }); } } // eslint-disable-next-line @typescript-eslint/no-misused-promises override async onTrackClose(track: OutputTrack) { const release = await this.mutex.acquire(); const trackData = this.trackDatas.find(x => x.track === track); if (trackData) { trackData.closed = true; } if (this.allTracksAreKnown()) { this.allTracksKnown.resolve(); } // Since a track is now closed, we may be able to write out chunks that were previously waiting await this.interleaveChunks(); release(); } /** Finalizes the file, making it ready for use. Must be called after all media chunks have been added. */ async finalize() { const release = await this.mutex.acquire(); this.allTracksKnown.resolve(); for (const trackData of this.trackDatas) { trackData.closed = true; } if (!this.segment) { this.createSegment(); } // Flush any remaining queued chunks to the file await this.interleaveChunks(true); if (this.currentCluster) { this.finalizeCurrentCluster(); } assert(this.cues); this.ebmlWriter.writeEBML(this.cues); if (!this.format._options.appendOnly) { // Write the Segment size const segmentSize = this.writer.getPos() - this.segmentDataOffset; this.writer.seek(this.ebmlWriter.offsets.get(this.segment!)! + 4); this.ebmlWriter.writeVarInt(segmentSize, SEGMENT_SIZE_BYTES); // Write the duration of the media to the Segment const duration = this.startTimestamp === Infinity ? 0 : this.endTimestamp - this.startTimestamp; this.segmentDuration!.data = new EBMLFloat64(duration); this.writer.seek(this.ebmlWriter.offsets.get(this.segmentDuration!)!); this.ebmlWriter.writeEBML(this.segmentDuration); // Fill in SeekHead position data and write it again assert(this.seekHead); this.writer.seek(this.ebmlWriter.offsets.get(this.seekHead)!); this.maybeCreateSeekHead(true); this.ebmlWriter.writeEBML(this.seekHead); } release(); } } ===== src/matroska/matroska-demuxer.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { TrackType } from '../output'; import { extractAv1CodecInfoFromPacket, extractAvcDecoderConfigurationRecord, extractHevcDecoderConfigurationRecord, extractVp9CodecInfoFromPacket, } from '../codec-data'; import { AacCodecInfo, AudioCodec, extractAudioCodecString, extractVideoCodecString, MediaCodec, OPUS_SAMPLE_RATE, VideoCodec, } from '../codec'; import { Demuxer } from '../demuxer'; import { Input } from '../input'; import { Logging } from '../logging'; import { InputAudioTrackBacking, InputTrackBacking, InputVideoTrackBacking, } from '../input-track'; import { AttachedFile, DEFAULT_TRACK_DISPOSITION, MetadataTags, TrackDisposition } from '../metadata'; import { PacketRetrievalOptions } from '../media-sink'; import { assert, binarySearchLessOrEqual, COLOR_PRIMARIES_MAP_INVERSE, findLastIndex, isIso639Dash2LanguageCode, last, MATRIX_COEFFICIENTS_MAP_INVERSE, normalizeRotation, Rotation, roundIfAlmostInteger, TRANSFER_CHARACTERISTICS_MAP_INVERSE, UNDETERMINED_LANGUAGE, } from '../misc'; import { EncodedPacket, EncodedPacketSideData, PLACEHOLDER_DATA } from '../packet'; import { assertDefinedSize, CODEC_STRING_MAP, EBMLId, LEVEL_0_AND_1_EBML_IDS, LEVEL_1_EBML_IDS, MAX_HEADER_SIZE, MIN_HEADER_SIZE, readAsciiString, readUnicodeString, readElementHeader, readElementId, readFloat, readUnsignedInt, readVarInt, resync, searchForNextElementId, readUnsignedBigInt, } from './ebml'; import { buildMatroskaMimeType } from './matroska-misc'; import { FileSlice, readBytes, Reader, readI16Be, readU8 } from '../reader'; type Segment = { seekHeadSeen: boolean; infoSeen: boolean; tracksSeen: boolean; cuesSeen: boolean; attachmentsSeen: boolean; tagsSeen: boolean; timestampScale: number; timestampFactor: number; duration: number; seekEntries: SeekEntry[]; tracks: InternalTrack[]; cuePoints: CuePoint[]; dataStartPos: number; elementEndPos: number | null; clusterSeekStartPos: number; /** * Caches the last cluster that was read. Based on the assumption that there will be multiple reads to the * same cluster in quick succession. */ lastReadCluster: Cluster | null; metadataTags: MetadataTags; metadataTagsCollected: boolean; }; type SeekEntry = { id: number; segmentPosition: number; }; type Cluster = { segment: Segment; elementStartPos: number; elementEndPos: number; dataStartPos: number; timestamp: number; trackData: Map; }; type ClusterTrackData = { track: InternalTrack; startTimestamp: number; endTimestamp: number; firstKeyFrameTimestamp: number | null; blocks: ClusterBlock[]; presentationTimestamps: { timestamp: number; blockIndex: number; }[]; }; enum BlockLacing { None, Xiph, FixedSize, Ebml, } type ClusterBlock = { timestamp: number; duration: number; isKeyFrame: boolean; data: Uint8Array; lacing: BlockLacing; decoded: boolean; mainAdditional: Uint8Array | null; }; type CuePoint = { time: number; trackId: number; clusterPosition: number; }; enum ContentEncodingScope { Block = 1, Private = 2, Next = 4, } enum ContentCompAlgo { Zlib, Bzlib, lzo1x, HeaderStripping, } type DecodingInstruction = { order: number; scope: ContentEncodingScope; data: { type: 'decompress'; algorithm: ContentCompAlgo | null; settings: Uint8Array | null; } | { type: 'decrypt'; // Don't store more yet since this operation is unsupported } | null; }; type InternalTrack = { id: number; demuxer: MatroskaDemuxer; segment: Segment; /** * List of all encountered cluster offsets alongside their timestamps. This list never gets truncated, but memory * consumption should be negligible. */ clusterPositionCache: { elementStartPos: number; startTimestamp: number; }[]; cuePoints: CuePoint[]; disposition: TrackDisposition; trackBacking: InputTrackBacking | null; codecId: string | null; codecPrivate: Uint8Array | null; defaultDuration: number | null; defaultDurationNs: number | null; name: string | null; languageCode: string; hasLanguageBcp47: boolean; decodingInstructions: DecodingInstruction[]; info: | null | { type: 'video'; width: number; height: number; displayWidth: number | null; displayHeight: number | null; displayUnit: number | null; squarePixelWidth: number; squarePixelHeight: number; rotation: Rotation; codec: VideoCodec | null; codecDescription: Uint8Array | null; colorSpace: VideoColorSpaceInit | null; alphaMode: boolean; } | { type: 'audio'; numberOfChannels: number; sampleRate: number; bitDepth: number; codec: AudioCodec | null; codecDescription: Uint8Array | null; aacCodecInfo: AacCodecInfo | null; }; }; type InternalVideoTrack = InternalTrack & { info: { type: 'video' } }; type InternalAudioTrack = InternalTrack & { info: { type: 'audio' } }; const METADATA_ELEMENTS = [ { id: EBMLId.SeekHead, flag: 'seekHeadSeen' }, { id: EBMLId.Info, flag: 'infoSeen' }, { id: EBMLId.Tracks, flag: 'tracksSeen' }, { id: EBMLId.Cues, flag: 'cuesSeen' }, ] as const; const MAX_RESYNC_LENGTH = 10 * 2 ** 20; // 10 MiB export class MatroskaDemuxer extends Demuxer { reader: Reader; readMetadataPromise: Promise | null = null; segments: Segment[] = []; currentSegment: Segment | null = null; currentTrack: InternalTrack | null = null; currentCluster: Cluster | null = null; currentBlock: ClusterBlock | null = null; currentBlockAdditional: { addId: number; data: Uint8Array | null; } | null = null; currentCueTime: number | null = null; currentDecodingInstruction: DecodingInstruction | null = null; currentTagTargetIsMovie: boolean = true; currentSimpleTagName: string | null = null; currentAttachedFile: { fileUid: bigint | null; fileName: string | null; fileMediaType: string | null; fileData: Uint8Array | null; fileDescription: string | null; } | null = null; isWebM = false; constructor(input: Input) { super(input); this.reader = input._reader; } async getTrackBackings() { await this.readMetadata(); return this.segments.flatMap(segment => segment.tracks.map(track => track.trackBacking!)); } override async getMimeType() { await this.readMetadata(); const backings = await this.getTrackBackings(); const codecStrings = await Promise.all(backings.map( x => x.getDecoderConfig().then(c => c?.codec ?? null), )); return buildMatroskaMimeType({ isWebM: this.isWebM, hasVideo: this.segments.some(segment => segment.tracks.some(x => x.info?.type === 'video')), hasAudio: this.segments.some(segment => segment.tracks.some(x => x.info?.type === 'audio')), codecStrings: codecStrings.filter(Boolean) as string[], }); } async getMetadataTags() { await this.readMetadata(); // Load metadata tags from each segment lazily (only once) for (const segment of this.segments) { if (!segment.metadataTagsCollected) { if (this.reader.fileSize !== null) { await this.loadSegmentMetadata(segment); } else { // The seeking would be too crazy, let's not } segment.metadataTagsCollected = true; } } // This is kinda handwavy, and how we handle multiple segments isn't suuuuper well-defined anyway; so we just // shallow-merge metadata tags from all (usually just one) segments. let metadataTags: MetadataTags = {}; for (const segment of this.segments) { metadataTags = { ...metadataTags, ...segment.metadataTags }; } return metadataTags; } readMetadata() { return this.readMetadataPromise ??= (async () => { let currentPos = 0; // Loop over all top-level elements in the file while (true) { let slice = this.reader.requestSliceRange(currentPos, MIN_HEADER_SIZE, MAX_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const header = readElementHeader(slice); if (!header) { break; // Zero padding at the end of the file triggers this, for example } const id = header.id; let size = header.size; const dataStartPos = slice.filePos; if (id === EBMLId.EBML) { assertDefinedSize(size); let slice = this.reader.requestSlice(dataStartPos, size); if (slice instanceof Promise) slice = await slice; if (!slice) break; this.readContiguousElements(slice); } else if (id === EBMLId.Segment) { // Segment found! await this.readSegment(dataStartPos, size); if (size === undefined) { // Segment sizes can be undefined (common in livestreamed files), so assume this is the last // and only segment break; } if (this.reader.fileSize === null) { break; // Stop at the first segment } } else if (id === EBMLId.Cluster) { if (this.reader.fileSize === null) { break; // Shouldn't be reached anyway, since we stop at the first segment } // Clusters are not a top-level element in Matroska, but some files contain a Segment whose size // doesn't contain any of the clusters that follow it. In the case, we apply the following logic: if // we find a top-level cluster, attribute it to the previous segment. if (size === undefined) { // Just in case this is one of those weird sizeless clusters, let's do our best and still try to // determine its size. const nextElementPos = await searchForNextElementId( this.reader, dataStartPos, LEVEL_0_AND_1_EBML_IDS, this.reader.fileSize, ); size = nextElementPos.pos - dataStartPos; } const lastSegment = last(this.segments); if (lastSegment) { // Extend the previous segment's size lastSegment.elementEndPos = dataStartPos + size; } } assertDefinedSize(size); currentPos = dataStartPos + size; } })(); } async readSegment(segmentDataStart: number, dataSize: number | undefined) { this.currentSegment = { seekHeadSeen: false, infoSeen: false, tracksSeen: false, cuesSeen: false, tagsSeen: false, attachmentsSeen: false, timestampScale: -1, timestampFactor: -1, duration: -1, seekEntries: [], tracks: [], cuePoints: [], dataStartPos: segmentDataStart, elementEndPos: dataSize === undefined ? null // Assume it goes until the end of the file : segmentDataStart + dataSize, clusterSeekStartPos: segmentDataStart, lastReadCluster: null, metadataTags: {}, metadataTagsCollected: false, }; this.segments.push(this.currentSegment); let currentPos = segmentDataStart; while (this.currentSegment.elementEndPos === null || currentPos < this.currentSegment.elementEndPos) { let slice = this.reader.requestSliceRange(currentPos, MIN_HEADER_SIZE, MAX_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const elementStartPos = currentPos; const header = readElementHeader(slice); if (!header || (!LEVEL_1_EBML_IDS.includes(header.id) && header.id !== EBMLId.Void)) { // Potential junk. Let's try to resync const nextPos = await resync( this.reader, elementStartPos, LEVEL_1_EBML_IDS, Math.min(this.currentSegment.elementEndPos ?? Infinity, elementStartPos + MAX_RESYNC_LENGTH), ); if (nextPos) { currentPos = nextPos; continue; } else { break; // Resync failed } } const { id, size } = header; const dataStartPos = slice.filePos; const metadataElementIndex = METADATA_ELEMENTS.findIndex(x => x.id === id); if (metadataElementIndex !== -1) { const field = METADATA_ELEMENTS[metadataElementIndex]!.flag; this.currentSegment[field] = true; assertDefinedSize(size); let slice = this.reader.requestSlice(dataStartPos, size); if (slice instanceof Promise) slice = await slice; if (slice) { this.readContiguousElements(slice); } } else if (id === EBMLId.Tags || id === EBMLId.Attachments) { // Metadata found at the beginning of the segment, great, let's parse it if (id === EBMLId.Tags) { this.currentSegment.tagsSeen = true; } else { this.currentSegment.attachmentsSeen = true; } assertDefinedSize(size); let slice = this.reader.requestSlice(dataStartPos, size); if (slice instanceof Promise) slice = await slice; if (slice) { this.readContiguousElements(slice); } } else if (id === EBMLId.Cluster) { this.currentSegment.clusterSeekStartPos = elementStartPos; break; // Stop at the first cluster } if (size === undefined) { break; } else { currentPos = dataStartPos + size; } } // Sort the seek entries by file position so reading them exhibits a sequential pattern this.currentSegment.seekEntries.sort((a, b) => a.segmentPosition - b.segmentPosition); if (this.reader.fileSize !== null) { // Use the seek head to read missing metadata elements for (const seekEntry of this.currentSegment.seekEntries) { const target = METADATA_ELEMENTS.find(x => x.id === seekEntry.id); if (!target) { continue; } if (this.currentSegment[target.flag]) continue; let slice = this.reader.requestSliceRange( segmentDataStart + seekEntry.segmentPosition, MIN_HEADER_SIZE, MAX_HEADER_SIZE, ); if (slice instanceof Promise) slice = await slice; if (!slice) continue; const header = readElementHeader(slice); if (!header) continue; const { id, size } = header; if (id !== target.id) continue; assertDefinedSize(size); this.currentSegment[target.flag] = true; let dataSlice = this.reader.requestSlice(slice.filePos, size); if (dataSlice instanceof Promise) dataSlice = await dataSlice; if (!dataSlice) continue; this.readContiguousElements(dataSlice); } } if (this.currentSegment.timestampScale === -1) { // TimestampScale element is missing. Technically an invalid file, but let's default to the typical value, // which is 1e6. this.currentSegment.timestampScale = 1e6; this.currentSegment.timestampFactor = 1e9 / 1e6; } // Compute default duration for all tracks now that we have the timestamp factor for (const track of this.currentSegment.tracks) { if (track.defaultDurationNs !== null) { track.defaultDuration = (this.currentSegment.timestampFactor * track.defaultDurationNs) / 1e9; } } // Now, let's distribute the cue points to the tracks const idToTrack = new Map(this.currentSegment.tracks.map(x => [x.id, x])); // Assign cue points to their respective tracks for (const cuePoint of this.currentSegment.cuePoints) { const track = idToTrack.get(cuePoint.trackId); if (track) { track.cuePoints.push(cuePoint); } } for (const track of this.currentSegment.tracks) { // Sort cue points by time track.cuePoints.sort((a, b) => a.time - b.time); // Remove multiple cue points for the same time for (let i = 0; i < track.cuePoints.length - 1; i++) { const cuePoint1 = track.cuePoints[i]!; const cuePoint2 = track.cuePoints[i + 1]!; if (cuePoint1.time === cuePoint2.time) { track.cuePoints.splice(i + 1, 1); i--; } } } let trackWithMostCuePoints: InternalTrack | null = null; let maxCuePointCount = -Infinity; for (const track of this.currentSegment.tracks) { if (track.cuePoints.length > maxCuePointCount) { maxCuePointCount = track.cuePoints.length; trackWithMostCuePoints = track; } } // For every track that has received 0 cue points (can happen, often only the video track receives cue points), // we still want to have better seeking. Therefore, let's give it the cue points of the track with the most cue // points, which should provide us with the most fine-grained seeking. for (const track of this.currentSegment.tracks) { if (track.cuePoints.length === 0) { track.cuePoints = trackWithMostCuePoints!.cuePoints; } } this.currentSegment = null; } async readCluster(startPos: number, segment: Segment) { if (segment.lastReadCluster?.elementStartPos === startPos) { return segment.lastReadCluster; } let headerSlice = this.reader.requestSliceRange(startPos, MIN_HEADER_SIZE, MAX_HEADER_SIZE); if (headerSlice instanceof Promise) headerSlice = await headerSlice; assert(headerSlice); const elementStartPos = startPos; const elementHeader = readElementHeader(headerSlice); assert(elementHeader); const id = elementHeader.id; assert(id === EBMLId.Cluster); let size = elementHeader.size; const dataStartPos = headerSlice.filePos; if (size === undefined) { // The cluster's size is undefined (can happen in livestreamed files). We'd still like to know the size of // it, so we have no other choice but to iterate over the EBML structure until we find an element at level // 0 or 1, indicating the end of the cluster (all elements inside the cluster are at level 2). const nextElementPos = await searchForNextElementId( this.reader, dataStartPos, LEVEL_0_AND_1_EBML_IDS, segment.elementEndPos, ); size = nextElementPos.pos - dataStartPos; } // Load the entire cluster let dataSlice = this.reader.requestSlice(dataStartPos, size); if (dataSlice instanceof Promise) dataSlice = await dataSlice; const cluster: Cluster = { segment, elementStartPos, elementEndPos: dataStartPos + size, dataStartPos, timestamp: -1, trackData: new Map(), }; this.currentCluster = cluster; if (dataSlice) { // Read the children of the cluster, stopping early at level 0 or 1 EBML elements. We do this because some // clusters have incorrect sizes that are too large const endPos = this.readContiguousElements(dataSlice, LEVEL_0_AND_1_EBML_IDS); cluster.elementEndPos = endPos; } for (const [, trackData] of cluster.trackData) { const track = trackData.track; // This must hold, as track datas only get created if a block for that track is encountered assert(trackData.blocks.length > 0); let hasLacedBlocks = false; for (let i = 0; i < trackData.blocks.length; i++) { const block = trackData.blocks[i]!; block.timestamp += cluster.timestamp; hasLacedBlocks ||= block.lacing !== BlockLacing.None; } trackData.presentationTimestamps = trackData.blocks .map((block, i) => ({ timestamp: block.timestamp, blockIndex: i })) .sort((a, b) => a.timestamp - b.timestamp); for (let i = 0; i < trackData.presentationTimestamps.length; i++) { const currentEntry = trackData.presentationTimestamps[i]!; const currentBlock = trackData.blocks[currentEntry.blockIndex]!; if (trackData.firstKeyFrameTimestamp === null && currentBlock.isKeyFrame) { trackData.firstKeyFrameTimestamp = currentBlock.timestamp; } if (i < trackData.presentationTimestamps.length - 1) { // Update block durations based on presentation order const nextEntry = trackData.presentationTimestamps[i + 1]!; currentBlock.duration = nextEntry.timestamp - currentBlock.timestamp; } else if (currentBlock.duration === 0) { if (track.defaultDuration != null) { if (currentBlock.lacing === BlockLacing.None) { currentBlock.duration = track.defaultDuration; } else { // Handled by the lace resolution code } } } } if (hasLacedBlocks) { // Perform lace resolution. Here, we expand each laced block into multiple blocks where each contains // one frame of the lace. We do this after determining block timestamps so we can properly distribute // the block's duration across the laced frames. this.expandLacedBlocks(trackData.blocks, track); // Recompute since blocks have changed trackData.presentationTimestamps = trackData.blocks .map((block, i) => ({ timestamp: block.timestamp, blockIndex: i })) .sort((a, b) => a.timestamp - b.timestamp); } const firstBlock = trackData.blocks[trackData.presentationTimestamps[0]!.blockIndex]!; const lastBlock = trackData.blocks[last(trackData.presentationTimestamps)!.blockIndex]!; trackData.startTimestamp = firstBlock.timestamp; trackData.endTimestamp = lastBlock.timestamp + lastBlock.duration; // Let's remember that a cluster with a given timestamp is here, speeding up future lookups if no cues exist const insertionIndex = binarySearchLessOrEqual( track.clusterPositionCache, trackData.startTimestamp, x => x.startTimestamp, ); if ( insertionIndex === -1 || track.clusterPositionCache[insertionIndex]!.elementStartPos !== elementStartPos ) { track.clusterPositionCache.splice(insertionIndex + 1, 0, { elementStartPos: cluster.elementStartPos, startTimestamp: trackData.startTimestamp, }); } } segment.lastReadCluster = cluster; return cluster; } getTrackDataInCluster(cluster: Cluster, trackNumber: number) { let trackData = cluster.trackData.get(trackNumber); if (!trackData) { const track = cluster.segment.tracks.find(x => x.id === trackNumber); if (!track) { return null; } trackData = { track, startTimestamp: 0, endTimestamp: 0, firstKeyFrameTimestamp: null, blocks: [], presentationTimestamps: [], }; cluster.trackData.set(trackNumber, trackData); } return trackData; } expandLacedBlocks(blocks: ClusterBlock[], track: InternalTrack) { // https://www.matroska.org/technical/notes.html#block-lacing for (let blockIndex = 0; blockIndex < blocks.length; blockIndex++) { const originalBlock = blocks[blockIndex]!; if (originalBlock.lacing === BlockLacing.None) { continue; } // Decode the block data if it hasn't been decoded yet (needed for lacing expansion) if (!originalBlock.decoded) { originalBlock.data = this.decodeBlockData(track, originalBlock.data); originalBlock.decoded = true; } const slice = FileSlice.tempFromBytes(originalBlock.data); const frameSizes: number[] = []; const frameCount = readU8(slice) + 1; switch (originalBlock.lacing) { case BlockLacing.Xiph: { let totalUsedSize = 0; // Xiph lacing, just like in Ogg for (let i = 0; i < frameCount - 1; i++) { let frameSize = 0; while (slice.bufferPos < slice.length) { const value = readU8(slice); frameSize += value; if (value < 255) { frameSizes.push(frameSize); totalUsedSize += frameSize; break; } } } // Compute the last frame's size from whatever's left frameSizes.push(slice.length - (slice.bufferPos + totalUsedSize)); }; break; case BlockLacing.FixedSize: { // Fixed size lacing: all frames have same size const totalDataSize = slice.length - 1; // Minus the frame count byte const frameSize = Math.floor(totalDataSize / frameCount); for (let i = 0; i < frameCount; i++) { frameSizes.push(frameSize); } }; break; case BlockLacing.Ebml: { // EBML lacing: first size absolute, subsequent ones are coded as signed differences from the last const firstResult = readVarInt(slice); assert(firstResult !== null); // Assume it's not an invalid VINT let currentSize = firstResult; frameSizes.push(currentSize); let totalUsedSize = currentSize; for (let i = 1; i < frameCount - 1; i++) { const startPos = slice.bufferPos; const diffResult = readVarInt(slice); assert(diffResult !== null); const unsignedDiff = diffResult; const width = slice.bufferPos - startPos; const bias = (1 << (width * 7 - 1)) - 1; // Typo-corrected version of 2^((7*n)-1)^-1 const diff = unsignedDiff - bias; currentSize += diff; frameSizes.push(currentSize); totalUsedSize += currentSize; } // Compute the last frame's size from whatever's left frameSizes.push(slice.length - (slice.bufferPos + totalUsedSize)); }; break; default: assert(false); } assert(frameSizes.length === frameCount); blocks.splice(blockIndex, 1); // Remove the original block const blockDuration = originalBlock.duration || frameCount * (track.defaultDuration ?? 0); // Now, let's insert each frame as its own block for (let i = 0; i < frameCount; i++) { const frameSize = frameSizes[i]!; const frameData = readBytes(slice, frameSize); // Distribute timestamps evenly across the block duration const frameTimestamp = originalBlock.timestamp + (blockDuration * i / frameCount); const frameDuration = blockDuration / frameCount; blocks.splice(blockIndex + i, 0, { timestamp: frameTimestamp, duration: frameDuration, isKeyFrame: originalBlock.isKeyFrame, data: frameData, lacing: BlockLacing.None, decoded: true, mainAdditional: originalBlock.mainAdditional, }); } blockIndex += frameCount; // Skip the blocks we just added blockIndex--; } } async loadSegmentMetadata(segment: Segment) { for (const seekEntry of segment.seekEntries) { if (seekEntry.id === EBMLId.Tags && !segment.tagsSeen) { // We need to load the tags } else if (seekEntry.id === EBMLId.Attachments && !segment.attachmentsSeen) { // We need to load the attachments } else { continue; } let slice = this.reader.requestSliceRange( segment.dataStartPos + seekEntry.segmentPosition, MIN_HEADER_SIZE, MAX_HEADER_SIZE, ); if (slice instanceof Promise) slice = await slice; if (!slice) continue; const header = readElementHeader(slice); if (!header || header.id !== seekEntry.id) continue; const { size } = header; assertDefinedSize(size); assert(!this.currentSegment); this.currentSegment = segment; let dataSlice = this.reader.requestSlice(slice.filePos, size); if (dataSlice instanceof Promise) dataSlice = await dataSlice; if (dataSlice) { this.readContiguousElements(dataSlice); } this.currentSegment = null; // Mark as seen if (seekEntry.id === EBMLId.Tags) { segment.tagsSeen = true; } else if (seekEntry.id === EBMLId.Attachments) { segment.attachmentsSeen = true; } } } readContiguousElements(slice: FileSlice, stopIds?: number[]) { while (slice.remainingLength >= MIN_HEADER_SIZE) { const startPos = slice.filePos; const foundElement = this.traverseElement(slice, stopIds); if (!foundElement) { return startPos; } } return slice.filePos; } traverseElement(slice: FileSlice, stopIds?: number[]): boolean { const header = readElementHeader(slice); if (!header) { return false; } if (stopIds && stopIds.includes(header.id)) { return false; } const { id, size } = header; const dataStartPos = slice.filePos; assertDefinedSize(size); switch (id) { case EBMLId.DocType: { this.isWebM = readAsciiString(slice, size) === 'webm'; }; break; case EBMLId.Seek: { if (!this.currentSegment) break; const seekEntry: SeekEntry = { id: -1, segmentPosition: -1 }; this.currentSegment.seekEntries.push(seekEntry); this.readContiguousElements(slice.slice(dataStartPos, size)); if (seekEntry.id === -1 || seekEntry.segmentPosition === -1) { this.currentSegment.seekEntries.pop(); } }; break; case EBMLId.SeekID: { const lastSeekEntry = this.currentSegment?.seekEntries[this.currentSegment.seekEntries.length - 1]; if (!lastSeekEntry) break; lastSeekEntry.id = readUnsignedInt(slice, size); }; break; case EBMLId.SeekPosition: { const lastSeekEntry = this.currentSegment?.seekEntries[this.currentSegment.seekEntries.length - 1]; if (!lastSeekEntry) break; lastSeekEntry.segmentPosition = readUnsignedInt(slice, size); }; break; case EBMLId.TimestampScale: { if (!this.currentSegment) break; this.currentSegment.timestampScale = readUnsignedInt(slice, size); this.currentSegment.timestampFactor = 1e9 / this.currentSegment.timestampScale; }; break; case EBMLId.Duration: { if (!this.currentSegment) break; this.currentSegment.duration = readFloat(slice, size); }; break; case EBMLId.TrackEntry: { if (!this.currentSegment) break; this.currentTrack = { id: -1, segment: this.currentSegment, demuxer: this, clusterPositionCache: [], cuePoints: [], disposition: { ...DEFAULT_TRACK_DISPOSITION, primary: false, }, trackBacking: null, codecId: null, codecPrivate: null, defaultDuration: null, defaultDurationNs: null, name: null, languageCode: 'eng', // The default in Matroska hasLanguageBcp47: false, decodingInstructions: [], info: null, }; this.readContiguousElements(slice.slice(dataStartPos, size)); // Check if track was disabled during parsing (e.g., by FlagEnabled being 0) if (!this.currentTrack) { break; } if (this.currentTrack.decodingInstructions.some((instruction) => { return instruction.data?.type !== 'decompress' || instruction.scope !== ContentEncodingScope.Block || instruction.data.algorithm !== ContentCompAlgo.HeaderStripping; })) { Logging._warn(`Track #${this.currentTrack.id} has an unsupported content encoding; dropping.`); this.currentTrack = null; } if ( this.currentTrack && this.currentTrack.id !== -1 && this.currentTrack.codecId && this.currentTrack.info ) { const slashIndex = this.currentTrack.codecId.indexOf('/'); const codecIdWithoutSuffix = slashIndex === -1 ? this.currentTrack.codecId : this.currentTrack.codecId.slice(0, slashIndex); if ( this.currentTrack.info.type === 'video' && this.currentTrack.info.width !== -1 && this.currentTrack.info.height !== -1 ) { this.currentTrack.info.squarePixelWidth = this.currentTrack.info.width; this.currentTrack.info.squarePixelHeight = this.currentTrack.info.height; if ( this.currentTrack.info.displayWidth !== null && this.currentTrack.info.displayHeight !== null ) { const num = this.currentTrack.info.displayWidth * this.currentTrack.info.height; const den = this.currentTrack.info.displayHeight * this.currentTrack.info.width; if (num > 0 && den > 0) { if (num > den) { this.currentTrack.info.squarePixelWidth = Math.round( this.currentTrack.info.width * num / den, ); } else { this.currentTrack.info.squarePixelHeight = Math.round( this.currentTrack.info.height * den / num, ); } } } if (this.currentTrack.codecId === CODEC_STRING_MAP.avc) { this.currentTrack.info.codec = 'avc'; this.currentTrack.info.codecDescription = this.currentTrack.codecPrivate; } else if (this.currentTrack.codecId === CODEC_STRING_MAP.hevc) { this.currentTrack.info.codec = 'hevc'; this.currentTrack.info.codecDescription = this.currentTrack.codecPrivate; } else if (codecIdWithoutSuffix === CODEC_STRING_MAP.vp8) { this.currentTrack.info.codec = 'vp8'; } else if (codecIdWithoutSuffix === CODEC_STRING_MAP.vp9) { this.currentTrack.info.codec = 'vp9'; } else if (codecIdWithoutSuffix === CODEC_STRING_MAP.av1) { this.currentTrack.info.codec = 'av1'; } const videoTrack = this.currentTrack as InternalVideoTrack; this.currentTrack.trackBacking = new MatroskaVideoTrackBacking(videoTrack); this.currentSegment.tracks.push(this.currentTrack); } else if (this.currentTrack.info.type === 'audio') { if (codecIdWithoutSuffix === CODEC_STRING_MAP.aac) { this.currentTrack.info.codec = 'aac'; this.currentTrack.info.aacCodecInfo = { isMpeg2: this.currentTrack.codecId.includes('MPEG2'), objectType: null, }; this.currentTrack.info.codecDescription = this.currentTrack.codecPrivate; } else if (this.currentTrack.codecId === CODEC_STRING_MAP.mp3) { this.currentTrack.info.codec = 'mp3'; } else if (codecIdWithoutSuffix === CODEC_STRING_MAP.opus) { this.currentTrack.info.codec = 'opus'; this.currentTrack.info.codecDescription = this.currentTrack.codecPrivate; this.currentTrack.info.sampleRate = OPUS_SAMPLE_RATE; // Always the same } else if (codecIdWithoutSuffix === CODEC_STRING_MAP.vorbis) { this.currentTrack.info.codec = 'vorbis'; this.currentTrack.info.codecDescription = this.currentTrack.codecPrivate; } else if (codecIdWithoutSuffix === CODEC_STRING_MAP.flac) { this.currentTrack.info.codec = 'flac'; this.currentTrack.info.codecDescription = this.currentTrack.codecPrivate; } else if (codecIdWithoutSuffix === CODEC_STRING_MAP.ac3) { this.currentTrack.info.codec = 'ac3'; this.currentTrack.info.codecDescription = this.currentTrack.codecPrivate; } else if (codecIdWithoutSuffix === CODEC_STRING_MAP.eac3) { this.currentTrack.info.codec = 'eac3'; this.currentTrack.info.codecDescription = this.currentTrack.codecPrivate; } else if (this.currentTrack.codecId === 'A_PCM/INT/LIT') { if (this.currentTrack.info.bitDepth === 8) { this.currentTrack.info.codec = 'pcm-u8'; } else if (this.currentTrack.info.bitDepth === 16) { this.currentTrack.info.codec = 'pcm-s16'; } else if (this.currentTrack.info.bitDepth === 24) { this.currentTrack.info.codec = 'pcm-s24'; } else if (this.currentTrack.info.bitDepth === 32) { this.currentTrack.info.codec = 'pcm-s32'; } } else if (this.currentTrack.codecId === 'A_PCM/INT/BIG') { if (this.currentTrack.info.bitDepth === 8) { this.currentTrack.info.codec = 'pcm-u8'; } else if (this.currentTrack.info.bitDepth === 16) { this.currentTrack.info.codec = 'pcm-s16be'; } else if (this.currentTrack.info.bitDepth === 24) { this.currentTrack.info.codec = 'pcm-s24be'; } else if (this.currentTrack.info.bitDepth === 32) { this.currentTrack.info.codec = 'pcm-s32be'; } } else if (this.currentTrack.codecId === 'A_PCM/FLOAT/IEEE') { if (this.currentTrack.info.bitDepth === 32) { this.currentTrack.info.codec = 'pcm-f32'; } else if (this.currentTrack.info.bitDepth === 64) { this.currentTrack.info.codec = 'pcm-f64'; } } const audioTrack = this.currentTrack as InternalAudioTrack; this.currentTrack.trackBacking = new MatroskaAudioTrackBacking(audioTrack); this.currentSegment.tracks.push(this.currentTrack); } } this.currentTrack = null; }; break; case EBMLId.TrackNumber: { if (!this.currentTrack) break; this.currentTrack.id = readUnsignedInt(slice, size); }; break; case EBMLId.TrackType: { if (!this.currentTrack) break; const type = readUnsignedInt(slice, size); if (type === 1) { this.currentTrack.info = { type: 'video', width: -1, height: -1, displayWidth: null, displayHeight: null, displayUnit: null, squarePixelWidth: -1, squarePixelHeight: -1, rotation: 0, codec: null, codecDescription: null, colorSpace: null, alphaMode: false, }; } else if (type === 2) { this.currentTrack.info = { type: 'audio', numberOfChannels: 1, // Default value sampleRate: 8000, // Default value bitDepth: -1, codec: null, codecDescription: null, aacCodecInfo: null, }; } }; break; case EBMLId.FlagEnabled: { if (!this.currentTrack) break; const enabled = readUnsignedInt(slice, size); if (!enabled) { this.currentTrack = null; } }; break; case EBMLId.FlagDefault: { if (!this.currentTrack) break; this.currentTrack.disposition.default = !!readUnsignedInt(slice, size); }; break; case EBMLId.FlagForced: { if (!this.currentTrack) break; this.currentTrack.disposition.forced = !!readUnsignedInt(slice, size); }; break; case EBMLId.FlagOriginal: { if (!this.currentTrack) break; this.currentTrack.disposition.original = !!readUnsignedInt(slice, size); }; break; case EBMLId.FlagHearingImpaired: { if (!this.currentTrack) break; this.currentTrack.disposition.hearingImpaired = !!readUnsignedInt(slice, size); }; break; case EBMLId.FlagVisualImpaired: { if (!this.currentTrack) break; this.currentTrack.disposition.visuallyImpaired = !!readUnsignedInt(slice, size); }; break; case EBMLId.FlagCommentary: { if (!this.currentTrack) break; this.currentTrack.disposition.commentary = !!readUnsignedInt(slice, size); }; break; case EBMLId.CodecID: { if (!this.currentTrack) break; this.currentTrack.codecId = readAsciiString(slice, size); }; break; case EBMLId.CodecPrivate: { if (!this.currentTrack) break; this.currentTrack.codecPrivate = readBytes(slice, size); }; break; case EBMLId.DefaultDuration: { if (!this.currentTrack) break; this.currentTrack.defaultDurationNs = readUnsignedInt(slice, size); }; break; case EBMLId.Name: { if (!this.currentTrack) break; this.currentTrack.name = readUnicodeString(slice, size); }; break; case EBMLId.Language: { if (!this.currentTrack) break; if (this.currentTrack.hasLanguageBcp47) { // LanguageBCP47 was present, which takes precedence break; } this.currentTrack.languageCode = readAsciiString(slice, size); if (!isIso639Dash2LanguageCode(this.currentTrack.languageCode)) { this.currentTrack.languageCode = UNDETERMINED_LANGUAGE; } }; break; case EBMLId.LanguageBCP47: { if (!this.currentTrack) break; const bcp47 = readAsciiString(slice, size); const languageSubtag = bcp47.split('-')[0]; if (languageSubtag) { // Technically invalid, for now: The language subtag might be a language code from ISO 639-1, // ISO 639-2, ISO 639-3, ISO 639-5 or some other thing (source: Wikipedia). But, `languageCode` is // documented as ISO 639-2. Changing the definition would be a breaking change. This will get // cleaned up in the future by defining languageCode to be BCP 47 instead. this.currentTrack.languageCode = languageSubtag; } else { this.currentTrack.languageCode = UNDETERMINED_LANGUAGE; } this.currentTrack.hasLanguageBcp47 = true; }; break; case EBMLId.Video: { if (this.currentTrack?.info?.type !== 'video') break; this.readContiguousElements(slice.slice(dataStartPos, size)); }; break; case EBMLId.PixelWidth: { if (this.currentTrack?.info?.type !== 'video') break; this.currentTrack.info.width = readUnsignedInt(slice, size); }; break; case EBMLId.PixelHeight: { if (this.currentTrack?.info?.type !== 'video') break; this.currentTrack.info.height = readUnsignedInt(slice, size); }; break; case EBMLId.DisplayWidth: { if (this.currentTrack?.info?.type !== 'video') break; this.currentTrack.info.displayWidth = readUnsignedInt(slice, size); }; break; case EBMLId.DisplayHeight: { if (this.currentTrack?.info?.type !== 'video') break; this.currentTrack.info.displayHeight = readUnsignedInt(slice, size); }; break; case EBMLId.DisplayUnit: { if (this.currentTrack?.info?.type !== 'video') break; this.currentTrack.info.displayUnit = readUnsignedInt(slice, size); }; break; case EBMLId.AlphaMode: { if (this.currentTrack?.info?.type !== 'video') break; this.currentTrack.info.alphaMode = readUnsignedInt(slice, size) === 1; }; break; case EBMLId.Colour: { if (this.currentTrack?.info?.type !== 'video') break; this.currentTrack.info.colorSpace = {}; this.readContiguousElements(slice.slice(dataStartPos, size)); }; break; case EBMLId.MatrixCoefficients: { if (this.currentTrack?.info?.type !== 'video' || !this.currentTrack.info.colorSpace) break; const matrixCoefficients = readUnsignedInt(slice, size); const mapped = MATRIX_COEFFICIENTS_MAP_INVERSE[matrixCoefficients] ?? null; this.currentTrack.info.colorSpace.matrix = mapped as VideoColorSpaceInit['matrix']; }; break; case EBMLId.Range: { if (this.currentTrack?.info?.type !== 'video' || !this.currentTrack.info.colorSpace) break; this.currentTrack.info.colorSpace.fullRange = readUnsignedInt(slice, size) === 2; }; break; case EBMLId.TransferCharacteristics: { if (this.currentTrack?.info?.type !== 'video' || !this.currentTrack.info.colorSpace) break; const transferCharacteristics = readUnsignedInt(slice, size); const mapped = TRANSFER_CHARACTERISTICS_MAP_INVERSE[transferCharacteristics] ?? null; this.currentTrack.info.colorSpace.transfer = mapped as VideoColorSpaceInit['transfer']; }; break; case EBMLId.Primaries: { if (this.currentTrack?.info?.type !== 'video' || !this.currentTrack.info.colorSpace) break; const primaries = readUnsignedInt(slice, size); const mapped = COLOR_PRIMARIES_MAP_INVERSE[primaries] ?? null; this.currentTrack.info.colorSpace.primaries = mapped as VideoColorSpaceInit['primaries']; }; break; case EBMLId.Projection: { if (this.currentTrack?.info?.type !== 'video') break; this.readContiguousElements(slice.slice(dataStartPos, size)); }; break; case EBMLId.ProjectionPoseRoll: { if (this.currentTrack?.info?.type !== 'video') break; const rotation = readFloat(slice, size); const flippedRotation = -rotation; // Convert counter-clockwise to clockwise try { this.currentTrack.info.rotation = normalizeRotation(flippedRotation); } catch { // It wasn't a valid rotation } }; break; case EBMLId.Audio: { if (this.currentTrack?.info?.type !== 'audio') break; this.readContiguousElements(slice.slice(dataStartPos, size)); }; break; case EBMLId.SamplingFrequency: { if (this.currentTrack?.info?.type !== 'audio') break; this.currentTrack.info.sampleRate = readFloat(slice, size); }; break; case EBMLId.Channels: { if (this.currentTrack?.info?.type !== 'audio') break; this.currentTrack.info.numberOfChannels = readUnsignedInt(slice, size); }; break; case EBMLId.BitDepth: { if (this.currentTrack?.info?.type !== 'audio') break; this.currentTrack.info.bitDepth = readUnsignedInt(slice, size); }; break; case EBMLId.CuePoint: { if (!this.currentSegment) break; this.readContiguousElements(slice.slice(dataStartPos, size)); this.currentCueTime = null; }; break; case EBMLId.CueTime: { this.currentCueTime = readUnsignedInt(slice, size); }; break; case EBMLId.CueTrackPositions: { if (this.currentCueTime === null) break; assert(this.currentSegment); const cuePoint: CuePoint = { time: this.currentCueTime, trackId: -1, clusterPosition: -1 }; this.currentSegment.cuePoints.push(cuePoint); this.readContiguousElements(slice.slice(dataStartPos, size)); if (cuePoint.trackId === -1 || cuePoint.clusterPosition === -1) { this.currentSegment.cuePoints.pop(); } }; break; case EBMLId.CueTrack: { const lastCuePoint = this.currentSegment?.cuePoints[this.currentSegment.cuePoints.length - 1]; if (!lastCuePoint) break; lastCuePoint.trackId = readUnsignedInt(slice, size); }; break; case EBMLId.CueClusterPosition: { const lastCuePoint = this.currentSegment?.cuePoints[this.currentSegment.cuePoints.length - 1]; if (!lastCuePoint) break; assert(this.currentSegment); lastCuePoint.clusterPosition = this.currentSegment.dataStartPos + readUnsignedInt(slice, size); }; break; case EBMLId.Timestamp: { if (!this.currentCluster) break; this.currentCluster.timestamp = readUnsignedInt(slice, size); }; break; case EBMLId.SimpleBlock: { if (!this.currentCluster) break; const trackNumber = readVarInt(slice); if (trackNumber === null) break; const trackData = this.getTrackDataInCluster(this.currentCluster, trackNumber); if (!trackData) break; // Not a track we care about const relativeTimestamp = readI16Be(slice); const flags = readU8(slice); const lacing = (flags >> 1) & 0x3 as BlockLacing; // If the block is laced, we'll expand it later let isKeyFrame = !!(flags & 0x80); if (trackData.track.info?.type === 'audio' && trackData.track.info.codec) { // Some files don't mark their audio packets as key packets (I'm looking at you, Firefox). But, we // can fix this in most cases: if we recognize the codec of the track, then we know every packet is // necessarily a key packet, no matter what the container says. // https://github.com/Vanilagy/mediabunny/issues/192 isKeyFrame = true; } const blockData = readBytes(slice, size - (slice.filePos - dataStartPos)); const hasDecodingInstructions = trackData.track.decodingInstructions.length > 0; trackData.blocks.push({ timestamp: relativeTimestamp, // We'll add the cluster's timestamp to this later duration: 0, // Will set later isKeyFrame, data: blockData, lacing, decoded: !hasDecodingInstructions, mainAdditional: null, }); }; break; case EBMLId.BlockGroup: { if (!this.currentCluster) break; this.readContiguousElements(slice.slice(dataStartPos, size)); this.currentBlock = null; }; break; case EBMLId.Block: { if (!this.currentCluster) break; const trackNumber = readVarInt(slice); if (trackNumber === null) break; const trackData = this.getTrackDataInCluster(this.currentCluster, trackNumber); if (!trackData) break; const relativeTimestamp = readI16Be(slice); const flags = readU8(slice); const lacing = (flags >> 1) & 0x3 as BlockLacing; // If the block is laced, we'll expand it later const blockData = readBytes(slice, size - (slice.filePos - dataStartPos)); const hasDecodingInstructions = trackData.track.decodingInstructions.length > 0; this.currentBlock = { timestamp: relativeTimestamp, // We'll add the cluster's timestamp to this later duration: 0, // Will set later isKeyFrame: true, data: blockData, lacing, decoded: !hasDecodingInstructions, mainAdditional: null, }; trackData.blocks.push(this.currentBlock); }; break; case EBMLId.BlockAdditions: { this.readContiguousElements(slice.slice(dataStartPos, size)); }; break; case EBMLId.BlockMore: { if (!this.currentBlock) break; this.currentBlockAdditional = { addId: 1, data: null, }; this.readContiguousElements(slice.slice(dataStartPos, size)); if (this.currentBlockAdditional.data && this.currentBlockAdditional.addId === 1) { this.currentBlock.mainAdditional = this.currentBlockAdditional.data; } this.currentBlockAdditional = null; }; break; case EBMLId.BlockAdditional: { if (!this.currentBlockAdditional) break; this.currentBlockAdditional.data = readBytes(slice, size); }; break; case EBMLId.BlockAddID: { if (!this.currentBlockAdditional) break; this.currentBlockAdditional.addId = readUnsignedInt(slice, size); }; break; case EBMLId.BlockDuration: { if (!this.currentBlock) break; this.currentBlock.duration = readUnsignedInt(slice, size); }; break; case EBMLId.ReferenceBlock: { if (!this.currentBlock) break; this.currentBlock.isKeyFrame = false; // We ignore the actual value here, we just use the reference as an indicator for "not a key frame". // This is in line with FFmpeg's behavior. }; break; case EBMLId.Tag: { this.currentTagTargetIsMovie = true; this.readContiguousElements(slice.slice(dataStartPos, size)); }; break; case EBMLId.Targets: { this.readContiguousElements(slice.slice(dataStartPos, size)); }; break; case EBMLId.TargetTypeValue: { const targetTypeValue = readUnsignedInt(slice, size); if (targetTypeValue !== 50) { this.currentTagTargetIsMovie = false; } }; break; case EBMLId.TagTrackUID: case EBMLId.TagEditionUID: case EBMLId.TagChapterUID: case EBMLId.TagAttachmentUID: { this.currentTagTargetIsMovie = false; }; break; case EBMLId.SimpleTag: { if (!this.currentTagTargetIsMovie) break; this.currentSimpleTagName = null; this.readContiguousElements(slice.slice(dataStartPos, size)); }; break; case EBMLId.TagName: { this.currentSimpleTagName = readUnicodeString(slice, size); }; break; case EBMLId.TagString: { if (!this.currentSimpleTagName) break; const value = readUnicodeString(slice, size); this.processTagValue(this.currentSimpleTagName, value); }; break; case EBMLId.TagBinary: { if (!this.currentSimpleTagName) break; const value = readBytes(slice, size); this.processTagValue(this.currentSimpleTagName, value); }; break; case EBMLId.AttachedFile: { if (!this.currentSegment) break; this.currentAttachedFile = { fileUid: null, fileName: null, fileMediaType: null, fileData: null, fileDescription: null, }; this.readContiguousElements(slice.slice(dataStartPos, size)); const tags = this.currentSegment.metadataTags; if (this.currentAttachedFile.fileUid && this.currentAttachedFile.fileData) { // All attached files get surfaced in the `raw` metadata tags tags.raw ??= {}; tags.raw[this.currentAttachedFile.fileUid.toString()] = new AttachedFile( this.currentAttachedFile.fileData, this.currentAttachedFile.fileMediaType ?? undefined, this.currentAttachedFile.fileName ?? undefined, this.currentAttachedFile.fileDescription ?? undefined, ); } // Only process image attachments if (this.currentAttachedFile.fileMediaType?.startsWith('image/') && this.currentAttachedFile.fileData) { const fileName = this.currentAttachedFile.fileName; let kind: 'coverFront' | 'coverBack' | 'unknown' = 'unknown'; if (fileName) { const lowerName = fileName.toLowerCase(); if (lowerName.startsWith('cover.')) { kind = 'coverFront'; } else if (lowerName.startsWith('back.')) { kind = 'coverBack'; } } tags.images ??= []; tags.images.push({ data: this.currentAttachedFile.fileData, mimeType: this.currentAttachedFile.fileMediaType, kind, name: this.currentAttachedFile.fileName ?? undefined, description: this.currentAttachedFile.fileDescription ?? undefined, }); } this.currentAttachedFile = null; }; break; case EBMLId.FileUID: { if (!this.currentAttachedFile) break; this.currentAttachedFile.fileUid = readUnsignedBigInt(slice, size); }; break; case EBMLId.FileName: { if (!this.currentAttachedFile) break; this.currentAttachedFile.fileName = readUnicodeString(slice, size); }; break; case EBMLId.FileMediaType: { if (!this.currentAttachedFile) break; this.currentAttachedFile.fileMediaType = readAsciiString(slice, size); }; break; case EBMLId.FileData: { if (!this.currentAttachedFile) break; this.currentAttachedFile.fileData = readBytes(slice, size); }; break; case EBMLId.FileDescription: { if (!this.currentAttachedFile) break; this.currentAttachedFile.fileDescription = readUnicodeString(slice, size); }; break; case EBMLId.ContentEncodings: { if (!this.currentTrack) break; this.readContiguousElements(slice.slice(dataStartPos, size)); // "**MUST** start with the `ContentEncoding` with the highest `ContentEncodingOrder`" this.currentTrack.decodingInstructions.sort((a, b) => b.order - a.order); }; break; case EBMLId.ContentEncoding: { this.currentDecodingInstruction = { order: 0, scope: ContentEncodingScope.Block, data: null, }; this.readContiguousElements(slice.slice(dataStartPos, size)); if (this.currentDecodingInstruction.data) { this.currentTrack!.decodingInstructions.push(this.currentDecodingInstruction); } this.currentDecodingInstruction = null; }; break; case EBMLId.ContentEncodingOrder: { if (!this.currentDecodingInstruction) break; this.currentDecodingInstruction.order = readUnsignedInt(slice, size); }; break; case EBMLId.ContentEncodingScope: { if (!this.currentDecodingInstruction) break; this.currentDecodingInstruction.scope = readUnsignedInt(slice, size); }; break; case EBMLId.ContentCompression: { if (!this.currentDecodingInstruction) break; this.currentDecodingInstruction.data = { type: 'decompress', algorithm: ContentCompAlgo.Zlib, settings: null, }; this.readContiguousElements(slice.slice(dataStartPos, size)); }; break; case EBMLId.ContentCompAlgo: { if (this.currentDecodingInstruction?.data?.type !== 'decompress') break; this.currentDecodingInstruction.data.algorithm = readUnsignedInt(slice, size); }; break; case EBMLId.ContentCompSettings: { if (this.currentDecodingInstruction?.data?.type !== 'decompress') break; this.currentDecodingInstruction.data.settings = readBytes(slice, size); }; break; case EBMLId.ContentEncryption: { if (!this.currentDecodingInstruction) break; this.currentDecodingInstruction.data = { type: 'decrypt', }; }; break; } slice.filePos = dataStartPos + size; return true; } decodeBlockData(track: InternalTrack, rawData: Uint8Array) { assert(track.decodingInstructions.length > 0); // This method shouldn't be called otherwise let currentData = rawData; for (const instruction of track.decodingInstructions) { assert(instruction.data); switch (instruction.data.type) { case 'decompress': { switch (instruction.data.algorithm) { case ContentCompAlgo.HeaderStripping: { if (instruction.data.settings && instruction.data.settings.length > 0) { const prefix = instruction.data.settings; const newData = new Uint8Array(prefix.length + currentData.length); newData.set(prefix, 0); newData.set(currentData, prefix.length); currentData = newData; } }; break; default: { // Unhandled }; } }; break; default: { // Unhandled }; } } return currentData; } processTagValue(name: string, value: string | Uint8Array) { if (!this.currentSegment?.metadataTags) return; const metadataTags = this.currentSegment.metadataTags; metadataTags.raw ??= {}; metadataTags.raw[name] ??= value; if (typeof value === 'string') { switch (name.toLowerCase()) { case 'title': { metadataTags.title ??= value; }; break; case 'description': { metadataTags.description ??= value; }; break; case 'artist': { metadataTags.artist ??= value; }; break; case 'album': { metadataTags.album ??= value; }; break; case 'album_artist': { metadataTags.albumArtist ??= value; }; break; case 'genre': { metadataTags.genre ??= value; }; break; case 'comment': { metadataTags.comment ??= value; }; break; case 'lyrics': { metadataTags.lyrics ??= value; }; break; case 'date': { const date = new Date(value); if (!Number.isNaN(date.getTime())) { metadataTags.date ??= date; } }; break; case 'track_number': case 'part_number': { const parts = value.split('/'); const trackNum = Number.parseInt(parts[0]!, 10); const tracksTotal = parts[1] && Number.parseInt(parts[1], 10); if (Number.isInteger(trackNum) && trackNum > 0) { metadataTags.trackNumber ??= trackNum; } if (tracksTotal && Number.isInteger(tracksTotal) && tracksTotal > 0) { metadataTags.tracksTotal ??= tracksTotal; } }; break; case 'disc_number': case 'disc': { const discParts = value.split('/'); const discNum = Number.parseInt(discParts[0]!, 10); const discsTotal = discParts[1] && Number.parseInt(discParts[1], 10); if (Number.isInteger(discNum) && discNum > 0) { metadataTags.discNumber ??= discNum; } if (discsTotal && Number.isInteger(discsTotal) && discsTotal > 0) { metadataTags.discsTotal ??= discsTotal; } }; break; } } } } abstract class MatroskaTrackBacking implements InputTrackBacking { packetToClusterLocation = new WeakMap(); constructor(public internalTrack: InternalTrack) {} abstract getType(): TrackType; abstract getDecoderConfig(): Promise; getId() { return this.internalTrack.id; } getNumber() { const demuxer = this.internalTrack.demuxer; const trackType = this.internalTrack.trackBacking!.getType(); let number = 0; for (const segment of demuxer.segments) { for (const track of segment.tracks) { if (track.trackBacking!.getType() === trackType) { number++; } if (track === this.internalTrack) { break; } } } return number; } getCodec(): MediaCodec | null { throw new Error('Not implemented on base class.'); } getInternalCodecId() { return this.internalTrack.codecId; } getName() { return this.internalTrack.name; } getLanguageCode() { return this.internalTrack.languageCode; } getTimeResolution() { return this.internalTrack.segment.timestampFactor; } isRelativeToUnixEpoch() { return false; } getUnixTimeForTimestamp() { return null; } getDisposition() { return this.internalTrack.disposition; } getPairingMask() { return 1n; } getBitrate() { return null; } getAverageBitrate() { return null; } async getDurationFromMetadata() { const segment = this.internalTrack.segment; if (segment.duration <= 0) { return null; } let endTimestamp = segment.duration / segment.timestampFactor; const firstPacket = await this.getFirstPacket({ metadataOnly: true }); endTimestamp += firstPacket?.timestamp ?? 0; return endTimestamp; } async getLiveRefreshInterval() { return null; } async getFirstPacket(options: PacketRetrievalOptions) { return this.performClusterLookup( null, (cluster) => { const trackData = cluster.trackData.get(this.internalTrack.id); if (trackData) { return { blockIndex: 0, correctBlockFound: true, }; } return { blockIndex: -1, correctBlockFound: false, }; }, -Infinity, // Use -Infinity as a search timestamp to avoid using the cues Infinity, options, ); } private intoTimescale(timestamp: number) { // Do a little rounding to catch cases where the result is very close to an integer. If it is, it's likely // that the number was originally an integer divided by the timescale. For stability, it's best // to return the integer in this case. return roundIfAlmostInteger(timestamp * this.internalTrack.segment.timestampFactor); } async getPacket(timestamp: number, options: PacketRetrievalOptions) { const timestampInTimescale = this.intoTimescale(timestamp); return this.performClusterLookup( null, (cluster) => { const trackData = cluster.trackData.get(this.internalTrack.id); if (!trackData) { return { blockIndex: -1, correctBlockFound: false }; } const index = binarySearchLessOrEqual( trackData.presentationTimestamps, timestampInTimescale, x => x.timestamp, ); const blockIndex = index !== -1 ? trackData.presentationTimestamps[index]!.blockIndex : -1; const correctBlockFound = index !== -1 && timestampInTimescale < trackData.endTimestamp; return { blockIndex, correctBlockFound }; }, timestampInTimescale, timestampInTimescale, options, ); } async getNextPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { const locationInCluster = this.packetToClusterLocation.get(packet); if (locationInCluster === undefined) { throw new Error('Packet was not created from this track.'); } return this.performClusterLookup( locationInCluster.cluster, (cluster) => { if (cluster === locationInCluster.cluster) { const trackData = cluster.trackData.get(this.internalTrack.id)!; if (locationInCluster.blockIndex + 1 < trackData.blocks.length) { // We can simply take the next block in the cluster return { blockIndex: locationInCluster.blockIndex + 1, correctBlockFound: true, }; } } else { const trackData = cluster.trackData.get(this.internalTrack.id); if (trackData) { return { blockIndex: 0, correctBlockFound: true, }; } } return { blockIndex: -1, correctBlockFound: false, }; }, -Infinity, // Use -Infinity as a search timestamp to avoid using the cues Infinity, options, ); } async getKeyPacket(timestamp: number, options: PacketRetrievalOptions) { const timestampInTimescale = this.intoTimescale(timestamp); return this.performClusterLookup( null, (cluster) => { const trackData = cluster.trackData.get(this.internalTrack.id); if (!trackData) { return { blockIndex: -1, correctBlockFound: false }; } const index = findLastIndex(trackData.presentationTimestamps, (x) => { const block = trackData.blocks[x.blockIndex]!; return block.isKeyFrame && x.timestamp <= timestampInTimescale; }); const blockIndex = index !== -1 ? trackData.presentationTimestamps[index]!.blockIndex : -1; const correctBlockFound = index !== -1 && timestampInTimescale < trackData.endTimestamp; return { blockIndex, correctBlockFound }; }, timestampInTimescale, timestampInTimescale, options, ); } async getNextKeyPacket(packet: EncodedPacket, options: PacketRetrievalOptions) { const locationInCluster = this.packetToClusterLocation.get(packet); if (locationInCluster === undefined) { throw new Error('Packet was not created from this track.'); } return this.performClusterLookup( locationInCluster.cluster, (cluster) => { if (cluster === locationInCluster.cluster) { const trackData = cluster.trackData.get(this.internalTrack.id)!; const nextKeyFrameIndex = trackData.blocks.findIndex( (x, i) => x.isKeyFrame && i > locationInCluster.blockIndex, ); if (nextKeyFrameIndex !== -1) { // We can simply take the next key frame in the cluster return { blockIndex: nextKeyFrameIndex, correctBlockFound: true, }; } } else { const trackData = cluster.trackData.get(this.internalTrack.id); if (trackData && trackData.firstKeyFrameTimestamp !== null) { const keyFrameIndex = trackData.blocks.findIndex(x => x.isKeyFrame); assert(keyFrameIndex !== -1); // There must be one return { blockIndex: keyFrameIndex, correctBlockFound: true, }; } } return { blockIndex: -1, correctBlockFound: false, }; }, -Infinity, // Use -Infinity as a search timestamp to avoid using the cues Infinity, options, ); } private async fetchPacketInCluster(cluster: Cluster, blockIndex: number, options: PacketRetrievalOptions) { if (blockIndex === -1) { return null; } const trackData = cluster.trackData.get(this.internalTrack.id)!; const block = trackData.blocks[blockIndex]; assert(block); // Perform lazy decoding if needed if (!block.decoded) { block.data = this.internalTrack.demuxer.decodeBlockData(this.internalTrack, block.data); block.decoded = true; } const data = options.metadataOnly ? PLACEHOLDER_DATA : block.data; const timestamp = block.timestamp / this.internalTrack.segment.timestampFactor; const duration = block.duration / this.internalTrack.segment.timestampFactor; const sideData: EncodedPacketSideData = {}; if (block.mainAdditional && this.internalTrack.info?.type === 'video' && this.internalTrack.info.alphaMode) { sideData.alpha = options.metadataOnly ? PLACEHOLDER_DATA : block.mainAdditional; sideData.alphaByteLength = block.mainAdditional.byteLength; } const packet = new EncodedPacket( data, block.isKeyFrame ? 'key' : 'delta', timestamp, duration, cluster.dataStartPos + blockIndex, block.data.byteLength, sideData, ); this.packetToClusterLocation.set(packet, { cluster, blockIndex }); return packet; } /** Looks for a packet in the clusters while trying to load as few clusters as possible to retrieve it. */ private async performClusterLookup( // The cluster where we start looking startCluster: Cluster | null, // This function returns the best-matching block in a given cluster getMatchInCluster: (cluster: Cluster) => { blockIndex: number; correctBlockFound: boolean }, // The timestamp with which we can search the lookup table searchTimestamp: number, // The timestamp for which we know the correct block will not come after it latestTimestamp: number, options: PacketRetrievalOptions, ): Promise { const { demuxer, segment } = this.internalTrack; let currentCluster: Cluster | null = null; let bestCluster: Cluster | null = null; let bestBlockIndex = -1; if (startCluster) { const { blockIndex, correctBlockFound } = getMatchInCluster(startCluster); if (correctBlockFound) { return this.fetchPacketInCluster(startCluster, blockIndex, options); } if (blockIndex !== -1) { bestCluster = startCluster; bestBlockIndex = blockIndex; } } // Search for a cue point; this way, we won't need to start searching from the start of the file // but can jump right into the correct cluster (or at least nearby). const cuePointIndex = binarySearchLessOrEqual( this.internalTrack.cuePoints, searchTimestamp, x => x.time, ); const cuePoint = cuePointIndex !== -1 ? this.internalTrack.cuePoints[cuePointIndex]! : null; // Also check the position cache const positionCacheIndex = binarySearchLessOrEqual( this.internalTrack.clusterPositionCache, searchTimestamp, x => x.startTimestamp, ); const positionCacheEntry = positionCacheIndex !== -1 ? this.internalTrack.clusterPositionCache[positionCacheIndex]! : null; const lookupEntryPosition = Math.max( cuePoint?.clusterPosition ?? 0, positionCacheEntry?.elementStartPos ?? 0, ) || null; let currentPos: number; if (!startCluster) { currentPos = lookupEntryPosition ?? segment.clusterSeekStartPos; } else { if (lookupEntryPosition === null || startCluster.elementStartPos >= lookupEntryPosition) { currentPos = startCluster.elementEndPos; currentCluster = startCluster; } else { // Use the lookup entry currentPos = lookupEntryPosition; } } while (segment.elementEndPos === null || currentPos <= segment.elementEndPos - MIN_HEADER_SIZE) { if (currentCluster) { const trackData = currentCluster.trackData.get(this.internalTrack.id); if (trackData && trackData.startTimestamp > latestTimestamp) { // We're already past the upper bound, no need to keep searching break; } } // Load the header let slice = demuxer.reader.requestSliceRange(currentPos, MIN_HEADER_SIZE, MAX_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const elementStartPos = currentPos; const elementHeader = readElementHeader(slice); if ( !elementHeader || (!LEVEL_1_EBML_IDS.includes(elementHeader.id) && elementHeader.id !== EBMLId.Void) ) { // There's an element here that shouldn't be here. Might be garbage. In this case, let's // try and resync to the next valid element. const nextPos = await resync( demuxer.reader, elementStartPos, LEVEL_1_EBML_IDS, Math.min(segment.elementEndPos ?? Infinity, elementStartPos + MAX_RESYNC_LENGTH), ); if (nextPos) { currentPos = nextPos; continue; } else { break; // Resync failed } } const id = elementHeader.id; let size = elementHeader.size; const dataStartPos = slice.filePos; if (id === EBMLId.Cluster) { currentCluster = await demuxer.readCluster(elementStartPos, segment); // readCluster computes the proper size even if it's undefined in the header, so let's use that instead size = currentCluster.elementEndPos - dataStartPos; const { blockIndex, correctBlockFound } = getMatchInCluster(currentCluster); if (correctBlockFound) { return this.fetchPacketInCluster(currentCluster, blockIndex, options); } if (blockIndex !== -1) { bestCluster = currentCluster; bestBlockIndex = blockIndex; } } if (size === undefined) { // Undefined element size (can happen in livestreamed files). In this case, we need to do some // searching to determine the actual size of the element. assert(id !== EBMLId.Cluster); // Undefined cluster sizes are fixed further up // Search for the next element at level 0 or 1 const nextElementPos = await searchForNextElementId( demuxer.reader, dataStartPos, LEVEL_0_AND_1_EBML_IDS, segment.elementEndPos, ); size = nextElementPos.pos - dataStartPos; } const endPos = dataStartPos + size; if (segment.elementEndPos === null) { // Check the next element. If it's a new segment, we know this segment ends here. The new // segment is just ignored, since we're likely in a livestreamed file and thus only care about // the first segment. let slice = demuxer.reader.requestSliceRange(endPos, MIN_HEADER_SIZE, MAX_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const elementId = readElementId(slice); if (elementId === EBMLId.Segment) { segment.elementEndPos = endPos; // We now know the segment's size break; } } currentPos = endPos; } // Catch faulty cue points if (cuePoint && (!bestCluster || bestCluster.elementStartPos < cuePoint.clusterPosition)) { // The cue point lied to us! We found a cue point but no cluster there that satisfied the match. In this // case, let's search again but using the cue point before that. const previousCuePoint = this.internalTrack.cuePoints[cuePointIndex - 1]; assert(!previousCuePoint || previousCuePoint.time < cuePoint.time); const newSearchTimestamp = previousCuePoint?.time ?? -Infinity; return this.performClusterLookup(null, getMatchInCluster, newSearchTimestamp, latestTimestamp, options); } if (bestCluster) { // If we finished looping but didn't find a perfect match, still return the best match we found return this.fetchPacketInCluster(bestCluster, bestBlockIndex, options); } return null; } } class MatroskaVideoTrackBacking extends MatroskaTrackBacking implements InputVideoTrackBacking { override internalTrack: InternalVideoTrack; decoderConfigPromise: Promise | null = null; constructor(internalTrack: InternalVideoTrack) { super(internalTrack); this.internalTrack = internalTrack; } getType() { return 'video' as const; } override getCodec(): VideoCodec | null { return this.internalTrack.info.codec; } getCodedWidth() { return this.internalTrack.info.width; } getCodedHeight() { return this.internalTrack.info.height; } getSquarePixelWidth() { return this.internalTrack.info.squarePixelWidth; } getSquarePixelHeight() { return this.internalTrack.info.squarePixelHeight; } getRotation() { return this.internalTrack.info.rotation; } async getColorSpace(): Promise { return { primaries: this.internalTrack.info.colorSpace?.primaries, transfer: this.internalTrack.info.colorSpace?.transfer, matrix: this.internalTrack.info.colorSpace?.matrix, fullRange: this.internalTrack.info.colorSpace?.fullRange, }; } async canBeTransparent() { return this.internalTrack.info.alphaMode; } async getDecoderConfig(): Promise { if (!this.internalTrack.info.codec) { return null; } return this.decoderConfigPromise ??= (async (): Promise => { let firstPacket: EncodedPacket | null = null; const needsPacketForAdditionalInfo = this.internalTrack.info.codec === 'vp9' || this.internalTrack.info.codec === 'av1' // Packets are in Annex B format: || (this.internalTrack.info.codec === 'avc' && !this.internalTrack.info.codecDescription) // Packets are in Annex B format: || (this.internalTrack.info.codec === 'hevc' && !this.internalTrack.info.codecDescription); if (needsPacketForAdditionalInfo) { firstPacket = await this.getFirstPacket({}); } const config: VideoDecoderConfig = { codec: extractVideoCodecString({ width: this.internalTrack.info.width, height: this.internalTrack.info.height, codec: this.internalTrack.info.codec, codecDescription: this.internalTrack.info.codecDescription, colorSpace: this.internalTrack.info.colorSpace, avcType: 1, // We don't know better (or do we?) so just assume 'avc1' avcCodecInfo: this.internalTrack.info.codec === 'avc' && firstPacket ? extractAvcDecoderConfigurationRecord(firstPacket.data) : null, hevcCodecInfo: this.internalTrack.info.codec === 'hevc' && firstPacket ? extractHevcDecoderConfigurationRecord(firstPacket.data) : null, vp9CodecInfo: this.internalTrack.info.codec === 'vp9' && firstPacket ? extractVp9CodecInfoFromPacket(firstPacket.data) : null, av1CodecInfo: this.internalTrack.info.codec === 'av1' && firstPacket ? extractAv1CodecInfoFromPacket(firstPacket.data) : null, }), codedWidth: this.internalTrack.info.width, codedHeight: this.internalTrack.info.height, description: this.internalTrack.info.codecDescription ?? undefined, colorSpace: this.internalTrack.info.colorSpace ?? undefined, }; if ( this.internalTrack.info.width !== this.internalTrack.info.squarePixelWidth || this.internalTrack.info.height !== this.internalTrack.info.squarePixelHeight ) { config.displayAspectWidth = this.internalTrack.info.squarePixelWidth; config.displayAspectHeight = this.internalTrack.info.squarePixelHeight; } return config; })(); } } class MatroskaAudioTrackBacking extends MatroskaTrackBacking implements InputAudioTrackBacking { override internalTrack: InternalAudioTrack; decoderConfig: AudioDecoderConfig | null = null; constructor(internalTrack: InternalAudioTrack) { super(internalTrack); this.internalTrack = internalTrack; } getType() { return 'audio' as const; } override getCodec(): AudioCodec | null { return this.internalTrack.info.codec; } getNumberOfChannels() { return this.internalTrack.info.numberOfChannels; } getSampleRate() { return this.internalTrack.info.sampleRate; } async getDecoderConfig(): Promise { if (!this.internalTrack.info.codec) { return null; } return this.decoderConfig ??= { codec: extractAudioCodecString({ codec: this.internalTrack.info.codec, codecDescription: this.internalTrack.info.codecDescription, aacCodecInfo: this.internalTrack.info.aacCodecInfo, }), numberOfChannels: this.internalTrack.info.numberOfChannels, sampleRate: this.internalTrack.info.sampleRate, description: this.internalTrack.info.codecDescription ?? undefined, }; } } ===== src/matroska/ebml.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { MediaCodec } from '../codec'; import { assert, assertNever, textDecoder, textEncoder } from '../misc'; import { FileSlice, readBytes, Reader, readF32Be, readF64Be, readU8 } from '../reader'; import { Writer } from '../writer'; export interface EBMLElement { id: number; size?: number; data: | number | bigint | string | Uint8Array | EBMLFloat32 | EBMLFloat64 | EBMLSignedInt | EBMLUnicodeString | (EBML | null)[]; } export type EBML = EBMLElement | Uint8Array | (EBML | null)[]; /** Wrapper around a number to be able to differentiate it in the writer. */ export class EBMLFloat32 { value: number; constructor(value: number) { this.value = value; } } /** Wrapper around a number to be able to differentiate it in the writer. */ export class EBMLFloat64 { value: number; constructor(value: number) { this.value = value; } } /** Wrapper around a number to be able to differentiate it in the writer. */ export class EBMLSignedInt { value: number; constructor(value: number) { this.value = value; } } export class EBMLUnicodeString { constructor(public value: string) {} } /** Defines some of the EBML IDs used by Matroska files. */ export enum EBMLId { EBML = 0x1a45dfa3, EBMLVersion = 0x4286, EBMLReadVersion = 0x42f7, EBMLMaxIDLength = 0x42f2, EBMLMaxSizeLength = 0x42f3, DocType = 0x4282, DocTypeVersion = 0x4287, DocTypeReadVersion = 0x4285, Void = 0xec, Segment = 0x18538067, SeekHead = 0x114d9b74, Seek = 0x4dbb, SeekID = 0x53ab, SeekPosition = 0x53ac, Duration = 0x4489, Info = 0x1549a966, TimestampScale = 0x2ad7b1, MuxingApp = 0x4d80, WritingApp = 0x5741, Tracks = 0x1654ae6b, TrackEntry = 0xae, TrackNumber = 0xd7, TrackUID = 0x73c5, TrackType = 0x83, FlagEnabled = 0xb9, FlagDefault = 0x88, FlagForced = 0x55aa, FlagOriginal = 0x55ae, FlagHearingImpaired = 0x55ab, FlagVisualImpaired = 0x55ac, FlagCommentary = 0x55af, FlagLacing = 0x9c, Name = 0x536e, Language = 0x22b59c, LanguageBCP47 = 0x22b59d, CodecID = 0x86, CodecPrivate = 0x63a2, CodecDelay = 0x56aa, SeekPreRoll = 0x56bb, DefaultDuration = 0x23e383, Video = 0xe0, PixelWidth = 0xb0, PixelHeight = 0xba, DisplayWidth = 0x54b0, DisplayHeight = 0x54ba, DisplayUnit = 0x54b2, AlphaMode = 0x53c0, Audio = 0xe1, SamplingFrequency = 0xb5, Channels = 0x9f, BitDepth = 0x6264, SimpleBlock = 0xa3, BlockGroup = 0xa0, Block = 0xa1, BlockAdditions = 0x75a1, BlockMore = 0xa6, BlockAdditional = 0xa5, BlockAddID = 0xee, BlockDuration = 0x9b, ReferenceBlock = 0xfb, Cluster = 0x1f43b675, Timestamp = 0xe7, Cues = 0x1c53bb6b, CuePoint = 0xbb, CueTime = 0xb3, CueTrackPositions = 0xb7, CueTrack = 0xf7, CueClusterPosition = 0xf1, Colour = 0x55b0, MatrixCoefficients = 0x55b1, TransferCharacteristics = 0x55ba, Primaries = 0x55bb, Range = 0x55b9, Projection = 0x7670, ProjectionType = 0x7671, ProjectionPoseRoll = 0x7675, Attachments = 0x1941a469, AttachedFile = 0x61a7, FileDescription = 0x467e, FileName = 0x466e, FileMediaType = 0x4660, FileData = 0x465c, FileUID = 0x46ae, Chapters = 0x1043a770, Tags = 0x1254c367, Tag = 0x7373, Targets = 0x63c0, TargetTypeValue = 0x68ca, TargetType = 0x63ca, TagTrackUID = 0x63c5, TagEditionUID = 0x63c9, TagChapterUID = 0x63c4, TagAttachmentUID = 0x63c6, SimpleTag = 0x67c8, TagName = 0x45a3, TagLanguage = 0x447a, TagString = 0x4487, TagBinary = 0x4485, ContentEncodings = 0x6d80, ContentEncoding = 0x6240, ContentEncodingOrder = 0x5031, ContentEncodingScope = 0x5032, ContentCompression = 0x5034, ContentCompAlgo = 0x4254, ContentCompSettings = 0x4255, ContentEncryption = 0x5035, } export const LEVEL_0_EBML_IDS: EBMLId[] = [ EBMLId.EBML, EBMLId.Segment, ]; // All the stuff that can appear in a segment, basically export const LEVEL_1_EBML_IDS: EBMLId[] = [ EBMLId.SeekHead, EBMLId.Info, EBMLId.Cluster, EBMLId.Tracks, EBMLId.Cues, EBMLId.Attachments, EBMLId.Chapters, EBMLId.Tags, ]; export const LEVEL_0_AND_1_EBML_IDS = [ ...LEVEL_0_EBML_IDS, ...LEVEL_1_EBML_IDS, ]; export const measureUnsignedInt = (value: number) => { if (value < (1 << 8)) { return 1; } else if (value < (1 << 16)) { return 2; } else if (value < (1 << 24)) { return 3; } else if (value < 2 ** 32) { return 4; } else if (value < 2 ** 40) { return 5; } else { return 6; } }; export const measureUnsignedBigInt = (value: bigint) => { if (value < (1n << 8n)) { return 1; } else if (value < (1n << 16n)) { return 2; } else if (value < (1n << 24n)) { return 3; } else if (value < (1n << 32n)) { return 4; } else if (value < (1n << 40n)) { return 5; } else if (value < (1n << 48n)) { return 6; } else if (value < (1n << 56n)) { return 7; } else { return 8; } }; export const measureSignedInt = (value: number) => { if (value >= -(1 << 6) && value < (1 << 6)) { return 1; } else if (value >= -(1 << 13) && value < (1 << 13)) { return 2; } else if (value >= -(1 << 20) && value < (1 << 20)) { return 3; } else if (value >= -(1 << 27) && value < (1 << 27)) { return 4; } else if (value >= -(2 ** 34) && value < 2 ** 34) { return 5; } else { return 6; } }; export const measureVarInt = (value: number) => { if (value < (1 << 7) - 1) { /** Top bit is set, leaving 7 bits to hold the integer, but we can't store * 127 because "all bits set to one" is a reserved value. Same thing for the * other cases below: */ return 1; } else if (value < (1 << 14) - 1) { return 2; } else if (value < (1 << 21) - 1) { return 3; } else if (value < (1 << 28) - 1) { return 4; } else if (value < 2 ** 35 - 1) { return 5; } else if (value < 2 ** 42 - 1) { return 6; } else { throw new Error('EBML varint size not supported ' + value); } }; export class EBMLWriter { helper = new Uint8Array(8); helperView = new DataView(this.helper.buffer); /** * Stores the position from the start of the file to where EBML elements have been written. This is used to * rewrite/edit elements that were already added before, and to measure sizes of things. */ offsets = new WeakMap(); /** Same as offsets, but stores position where the element's data starts (after ID and size fields). */ dataOffsets = new WeakMap(); constructor(private writer: Writer) {} writeByte(value: number) { this.helperView.setUint8(0, value); this.writer.write(this.helper.subarray(0, 1)); } writeFloat32(value: number) { this.helperView.setFloat32(0, value, false); this.writer.write(this.helper.subarray(0, 4)); } writeFloat64(value: number) { this.helperView.setFloat64(0, value, false); this.writer.write(this.helper); } writeUnsignedInt(value: number, width = measureUnsignedInt(value)) { let pos = 0; // Each case falls through: switch (width) { case 6: // Need to use division to access >32 bits of floating point var this.helperView.setUint8(pos++, (value / 2 ** 40) | 0); // eslint-disable-next-line no-fallthrough case 5: this.helperView.setUint8(pos++, (value / 2 ** 32) | 0); // eslint-disable-next-line no-fallthrough case 4: this.helperView.setUint8(pos++, value >> 24); // eslint-disable-next-line no-fallthrough case 3: this.helperView.setUint8(pos++, value >> 16); // eslint-disable-next-line no-fallthrough case 2: this.helperView.setUint8(pos++, value >> 8); // eslint-disable-next-line no-fallthrough case 1: this.helperView.setUint8(pos++, value); break; default: throw new Error('Bad unsigned int size ' + width); } this.writer.write(this.helper.subarray(0, pos)); } writeUnsignedBigInt(value: bigint, width = measureUnsignedBigInt(value)) { let pos = 0; for (let i = width - 1; i >= 0; i--) { this.helperView.setUint8(pos++, Number((value >> BigInt(i * 8)) & 0xffn)); } this.writer.write(this.helper.subarray(0, pos)); } writeSignedInt(value: number, width = measureSignedInt(value)) { if (value < 0) { // Two's complement stuff value += 2 ** (width * 8); } this.writeUnsignedInt(value, width); } writeVarInt(value: number, width = measureVarInt(value)) { let pos = 0; switch (width) { case 1: this.helperView.setUint8(pos++, (1 << 7) | value); break; case 2: this.helperView.setUint8(pos++, (1 << 6) | (value >> 8)); this.helperView.setUint8(pos++, value); break; case 3: this.helperView.setUint8(pos++, (1 << 5) | (value >> 16)); this.helperView.setUint8(pos++, value >> 8); this.helperView.setUint8(pos++, value); break; case 4: this.helperView.setUint8(pos++, (1 << 4) | (value >> 24)); this.helperView.setUint8(pos++, value >> 16); this.helperView.setUint8(pos++, value >> 8); this.helperView.setUint8(pos++, value); break; case 5: /** * JavaScript converts its doubles to 32-bit integers for bitwise * operations, so we need to do a division by 2^32 instead of a * right-shift of 32 to retain those top 3 bits */ this.helperView.setUint8(pos++, (1 << 3) | ((value / 2 ** 32) & 0x7)); this.helperView.setUint8(pos++, value >> 24); this.helperView.setUint8(pos++, value >> 16); this.helperView.setUint8(pos++, value >> 8); this.helperView.setUint8(pos++, value); break; case 6: this.helperView.setUint8(pos++, (1 << 2) | ((value / 2 ** 40) & 0x3)); this.helperView.setUint8(pos++, (value / 2 ** 32) | 0); this.helperView.setUint8(pos++, value >> 24); this.helperView.setUint8(pos++, value >> 16); this.helperView.setUint8(pos++, value >> 8); this.helperView.setUint8(pos++, value); break; default: throw new Error('Bad EBML varint size ' + width); } this.writer.write(this.helper.subarray(0, pos)); } writeAsciiString(str: string) { this.writer.write(new Uint8Array(str.split('').map(x => x.charCodeAt(0)))); } writeEBML(data: EBML | null) { if (data === null) return; if (data instanceof Uint8Array) { this.writer.write(data); } else if (Array.isArray(data)) { for (const elem of data) { this.writeEBML(elem); } } else { this.offsets.set(data, this.writer.getPos()); this.writeUnsignedInt(data.id); // ID field if (Array.isArray(data.data)) { const sizePos = this.writer.getPos(); const sizeSize = data.size === -1 ? 1 : (data.size ?? 4); if (data.size === -1) { // Write the reserved all-one-bits marker for unknown/unbounded size. this.writeByte(0xff); } else { this.writer.seek(this.writer.getPos() + sizeSize); } const startPos = this.writer.getPos(); this.dataOffsets.set(data, startPos); this.writeEBML(data.data); if (data.size !== -1) { const size = this.writer.getPos() - startPos; const endPos = this.writer.getPos(); this.writer.seek(sizePos); this.writeVarInt(size, sizeSize); this.writer.seek(endPos); } } else if (typeof data.data === 'number') { const size = data.size ?? measureUnsignedInt(data.data); this.writeVarInt(size); this.writeUnsignedInt(data.data, size); } else if (typeof data.data === 'bigint') { const size = data.size ?? measureUnsignedBigInt(data.data); this.writeVarInt(size); this.writeUnsignedBigInt(data.data, size); } else if (typeof data.data === 'string') { this.writeVarInt(data.data.length); this.writeAsciiString(data.data); } else if (data.data instanceof Uint8Array) { this.writeVarInt(data.data.byteLength, data.size); this.writer.write(data.data); } else if (data.data instanceof EBMLFloat32) { this.writeVarInt(4); this.writeFloat32(data.data.value); } else if (data.data instanceof EBMLFloat64) { this.writeVarInt(8); this.writeFloat64(data.data.value); } else if (data.data instanceof EBMLSignedInt) { const size = data.size ?? measureSignedInt(data.data.value); this.writeVarInt(size); this.writeSignedInt(data.data.value, size); } else if (data.data instanceof EBMLUnicodeString) { const bytes = textEncoder.encode(data.data.value); this.writeVarInt(bytes.length); this.writer.write(bytes); } else { assertNever(data.data); } } } } export const MAX_VAR_INT_SIZE = 8; export const MIN_HEADER_SIZE = 2; // 1-byte ID and 1-byte size export const MAX_HEADER_SIZE = 2 * MAX_VAR_INT_SIZE; // 8-byte ID and 8-byte size export const readVarIntSize = (slice: FileSlice) => { if (slice.remainingLength < 1) { return null; } const firstByte = readU8(slice); slice.skip(-1); if (firstByte === 0) { return null; // Invalid VINT } let width = 1; let mask = 0x80; while ((firstByte & mask) === 0) { width++; mask >>= 1; } // Check if we have enough bytes to read the full varint if (slice.remainingLength < width) { return null; } return width; }; export const readVarInt = (slice: FileSlice) => { if (slice.remainingLength < 1) { return null; } // Read the first byte to determine the width of the variable-length integer const firstByte = readU8(slice); if (firstByte === 0) { return null; // Invalid VINT } // Find the position of VINT_MARKER, which determines the width let width = 1; let mask = 1 << 7; while ((firstByte & mask) === 0) { width++; mask >>= 1; } if (slice.remainingLength < width - 1) { // Not enough bytes return null; } // First byte's value needs the marker bit cleared let value = firstByte & (mask - 1); // Read remaining bytes for (let i = 1; i < width; i++) { value *= 1 << 8; value += readU8(slice); } return value; }; export const readUnsignedInt = (slice: FileSlice, width: number) => { if (width < 1 || width > 8) { throw new Error('Bad unsigned int size ' + width); } let value = 0; // Read bytes from most significant to least significant for (let i = 0; i < width; i++) { value *= 1 << 8; value += readU8(slice); } return value; }; export const readUnsignedBigInt = (slice: FileSlice, width: number) => { if (width < 1) { throw new Error('Bad unsigned int size ' + width); } let value = 0n; for (let i = 0; i < width; i++) { value <<= 8n; value += BigInt(readU8(slice)); } return value; }; export const readSignedInt = (slice: FileSlice, width: number) => { let value = readUnsignedInt(slice, width); // If the highest bit is set, convert from two's complement if (value & (1 << (width * 8 - 1))) { value -= 2 ** (width * 8); } return value; }; export const readElementId = (slice: FileSlice) => { const size = readVarIntSize(slice); if (size === null) { return null; } if (slice.remainingLength < size) { return null; // It don't fit } const id = readUnsignedInt(slice, size); return id; }; /** Returns `undefined` to indicate the EBML undefined size. Returns `null` if the size couldn't be read. */ export const readElementSize = (slice: FileSlice): number | undefined | null => { // Need at least 1 byte to read the size if (slice.remainingLength < 1) { return null; } const firstByte = readU8(slice); if (firstByte === 0xff) { return undefined; } slice.skip(-1); const size = readVarInt(slice); if (size === null) { return null; } // In some (livestreamed) files, this is the value of the size field. While this technically is just a very // large number, it is intended to behave like the reserved size 0xFF, meaning the size is undefined. We // catch the number here. Note that it cannot be perfectly represented as a double, but the comparison works // nonetheless. // eslint-disable-next-line no-loss-of-precision if (size === 0x00ffffffffffffff) { return undefined; } return size; }; export const readElementHeader = (slice: FileSlice) => { assert(slice.remainingLength >= MIN_HEADER_SIZE); const id = readElementId(slice); if (id === null) { return null; } const size = readElementSize(slice); if (size === null) { return null; } return { id, size }; }; export const readAsciiString = (slice: FileSlice, length: number) => { const bytes = readBytes(slice, length); // Actual string length might be shorter due to null terminators let strLength = 0; while (strLength < length && bytes[strLength] !== 0) { strLength += 1; } return String.fromCharCode(...bytes.subarray(0, strLength)); }; export const readUnicodeString = (slice: FileSlice, length: number) => { const bytes = readBytes(slice, length); // Actual string length might be shorter due to null terminators let strLength = 0; while (strLength < length && bytes[strLength] !== 0) { strLength += 1; } return textDecoder.decode(bytes.subarray(0, strLength)); }; export const readFloat = (slice: FileSlice, width: number) => { if (width === 0) { return 0; } if (width !== 4 && width !== 8) { throw new Error('Bad float size ' + width); } return width === 4 ? readF32Be(slice) : readF64Be(slice); }; /** Returns the byte offset in the file of the next element with a matching ID. */ export const searchForNextElementId = async ( reader: Reader, startPos: number, ids: EBMLId[], until: number | null, ): Promise<{ pos: number; found: boolean }> => { const idsSet = new Set(ids); let currentPos = startPos; while (until === null || currentPos < until) { let slice = reader.requestSliceRange(currentPos, MIN_HEADER_SIZE, MAX_HEADER_SIZE); if (slice instanceof Promise) slice = await slice; if (!slice) break; const elementHeader = readElementHeader(slice); if (!elementHeader) { break; } if (idsSet.has(elementHeader.id)) { return { pos: currentPos, found: true }; } assertDefinedSize(elementHeader.size); currentPos = slice.filePos + elementHeader.size; } return { pos: (until !== null && until > currentPos) ? until : currentPos, found: false }; }; /** Searches for the next occurrence of an element ID using a naive byte-wise search. */ export const resync = async (reader: Reader, startPos: number, ids: EBMLId[], until: number) => { const CHUNK_SIZE = 2 ** 16; // So we don't need to grab thousands of slices const idsSet = new Set(ids); let currentPos = startPos; while (currentPos < until) { let slice = reader.requestSliceRange(currentPos, 0, Math.min(CHUNK_SIZE, until - currentPos)); if (slice instanceof Promise) slice = await slice; if (!slice) break; if (slice.length < MAX_VAR_INT_SIZE) break; for (let i = 0; i < slice.length - MAX_VAR_INT_SIZE; i++) { slice.filePos = currentPos; const elementId = readElementId(slice); if (elementId !== null && idsSet.has(elementId)) { return currentPos; } currentPos++; } } return null; }; export const CODEC_STRING_MAP: Partial> = { 'avc': 'V_MPEG4/ISO/AVC', 'hevc': 'V_MPEGH/ISO/HEVC', 'vp8': 'V_VP8', 'vp9': 'V_VP9', 'av1': 'V_AV1', 'aac': 'A_AAC', 'mp3': 'A_MPEG/L3', 'opus': 'A_OPUS', 'vorbis': 'A_VORBIS', 'flac': 'A_FLAC', 'ac3': 'A_AC3', 'eac3': 'A_EAC3', 'pcm-u8': 'A_PCM/INT/LIT', 'pcm-s16': 'A_PCM/INT/LIT', 'pcm-s16be': 'A_PCM/INT/BIG', 'pcm-s24': 'A_PCM/INT/LIT', 'pcm-s24be': 'A_PCM/INT/BIG', 'pcm-s32': 'A_PCM/INT/LIT', 'pcm-s32be': 'A_PCM/INT/BIG', 'pcm-f32': 'A_PCM/FLOAT/IEEE', 'pcm-f64': 'A_PCM/FLOAT/IEEE', 'webvtt': 'S_TEXT/WEBVTT', }; export function assertDefinedSize(size: number | undefined): asserts size is number { if (size === undefined) { throw new Error('Undefined element size is used in a place where it is not supported.'); } }; ===== src/misc.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { Bitstream } from '../shared/bitstream'; import { Logging } from './logging'; export function assert(x: unknown): asserts x { if (!x) { throw new Error('Assertion failed.'); } } /** * Represents a clockwise rotation in degrees. * @group Miscellaneous * @public */ export type Rotation = 0 | 90 | 180 | 270; export const normalizeRotation = (rotation: number) => { const mappedRotation = (rotation % 360 + 360) % 360; if (mappedRotation === 0 || mappedRotation === 90 || mappedRotation === 180 || mappedRotation === 270) { return mappedRotation as Rotation; } else { throw new Error(`Invalid rotation ${rotation}.`); } }; export type TransformationMatrix = [number, number, number, number, number, number, number, number, number]; export const last = (arr: T[]) => { return arr && arr[arr.length - 1]; }; export const isU32 = (value: number) => { return value >= 0 && value < 2 ** 32; }; /** Reads an exponential-Golomb universal code from a Bitstream. */ export const readExpGolomb = (bitstream: Bitstream) => { let leadingZeroBits = 0; while (bitstream.readBits(1) === 0 && leadingZeroBits < 32) { leadingZeroBits++; } if (leadingZeroBits >= 32) { throw new Error('Invalid exponential-Golomb code.'); } const result = (1 << leadingZeroBits) - 1 + bitstream.readBits(leadingZeroBits); return result; }; /** Reads a signed exponential-Golomb universal code from a Bitstream. */ export const readSignedExpGolomb = (bitstream: Bitstream) => { const codeNum = readExpGolomb(bitstream); return ((codeNum & 1) === 0) ? -(codeNum >> 1) : ((codeNum + 1) >> 1); }; export const writeBits = (bytes: Uint8Array, start: number, end: number, value: number) => { for (let i = start; i < end; i++) { const byteIndex = Math.floor(i / 8); let byte = bytes[byteIndex]!; const bitIndex = 0b111 - (i & 0b111); byte &= ~(1 << bitIndex); byte |= ((value & (1 << (end - i - 1))) >> (end - i - 1)) << bitIndex; bytes[byteIndex] = byte; } }; export const toUint8Array = (source: AllowSharedBufferSource): Uint8Array => { if (source.constructor === Uint8Array) { // We want a true Uint8Array, not something that extends it like Buffer return source; } else if (ArrayBuffer.isView(source)) { return new Uint8Array(source.buffer, source.byteOffset, source.byteLength); } else { return new Uint8Array(source); } }; export const toDataView = (source: AllowSharedBufferSource): DataView => { if (source.constructor === DataView) { return source; } else if (ArrayBuffer.isView(source)) { return new DataView(source.buffer, source.byteOffset, source.byteLength); } else { return new DataView(source); } }; export const textDecoder = /* #__PURE__ */ new TextDecoder(); export const textEncoder = /* #__PURE__ */ new TextEncoder(); export const isIso88591Compatible = (text: string) => { for (let i = 0; i < text.length; i++) { const code = text.charCodeAt(i); if (code > 255) { return false; } } return true; }; const invertObject = (object: Record) => { return Object.fromEntries(Object.entries(object).map(([key, value]) => [value, key])) as Record; }; // For the color space mappings, see Rec. ITU-T H.273. export const COLOR_PRIMARIES_MAP = { bt709: 1, // ITU-R BT.709 bt470bg: 5, // ITU-R BT.470BG smpte170m: 6, // ITU-R BT.601 525 - SMPTE 170M bt2020: 9, // ITU-R BT.202 smpte432: 12, // SMPTE EG 432-1 }; export const COLOR_PRIMARIES_MAP_INVERSE = /* #__PURE__ */ invertObject(COLOR_PRIMARIES_MAP); export const TRANSFER_CHARACTERISTICS_MAP = { 'bt709': 1, // ITU-R BT.709 'smpte170m': 6, // SMPTE 170M 'linear': 8, // Linear transfer characteristics 'iec61966-2-1': 13, // IEC 61966-2-1 'pq': 16, // Rec. ITU-R BT.2100-2 perceptual quantization (PQ) system 'hlg': 18, // Rec. ITU-R BT.2100-2 hybrid loggamma (HLG) system }; export const TRANSFER_CHARACTERISTICS_MAP_INVERSE = /* #__PURE__ */ invertObject(TRANSFER_CHARACTERISTICS_MAP); export const MATRIX_COEFFICIENTS_MAP = { 'rgb': 0, // Identity 'bt709': 1, // ITU-R BT.709 'bt470bg': 5, // ITU-R BT.470BG 'smpte170m': 6, // SMPTE 170M 'bt2020-ncl': 9, // ITU-R BT.2020-2 (non-constant luminance) }; export const MATRIX_COEFFICIENTS_MAP_INVERSE = /* #__PURE__ */ invertObject(MATRIX_COEFFICIENTS_MAP); export type RequiredNonNull = { [K in keyof T]-?: NonNullable; }; export const colorSpaceIsComplete = ( colorSpace: VideoColorSpaceInit | undefined, ): colorSpace is RequiredNonNull => { return ( !!colorSpace && !!colorSpace.primaries && !!colorSpace.transfer && !!colorSpace.matrix && colorSpace.fullRange !== undefined ); }; export const isAllowSharedBufferSource = (x: unknown) => { return ( x instanceof ArrayBuffer || (typeof SharedArrayBuffer !== 'undefined' && x instanceof SharedArrayBuffer) || ArrayBuffer.isView(x) ); }; export class AsyncMutex { currentPromise = Promise.resolve(); pending = 0; async acquire() { let resolver: () => void; const nextPromise = new Promise((resolve) => { let resolved = false; resolver = () => { if (resolved) { return; } resolve(); this.pending--; resolved = true; }; }); const currentPromiseAlias = this.currentPromise; this.currentPromise = nextPromise; this.pending++; await currentPromiseAlias; return resolver!; } } export const HEX_STRING_REGEX = /^[0-9a-fA-F]+$/; export const bytesToHexString = (bytes: Uint8Array) => { return [...bytes].map(x => x.toString(16).padStart(2, '0')).join(''); }; export const hexStringToBytes = (hexString: string) => { assert(hexString.length % 2 === 0); const bytes = new Uint8Array(hexString.length / 2); for (let i = 0; i < hexString.length; i += 2) { bytes[i / 2] = parseInt(hexString.slice(i, i + 2), 16); } return bytes; }; export const reverseBitsU32 = (x: number): number => { x = ((x >> 1) & 0x55555555) | ((x & 0x55555555) << 1); x = ((x >> 2) & 0x33333333) | ((x & 0x33333333) << 2); x = ((x >> 4) & 0x0f0f0f0f) | ((x & 0x0f0f0f0f) << 4); x = ((x >> 8) & 0x00ff00ff) | ((x & 0x00ff00ff) << 8); x = ((x >> 16) & 0x0000ffff) | ((x & 0x0000ffff) << 16); return x >>> 0; // Ensure it's treated as an unsigned 32-bit integer }; /** Returns the smallest index i such that val[i] === key, or -1 if no such index exists. */ export const binarySearchExact = (arr: T[], key: number, valueGetter: (x: T) => number): number => { let low = 0; let high = arr.length - 1; let ans = -1; while (low <= high) { const mid = (low + high) >> 1; const midVal = valueGetter(arr[mid]!); if (midVal === key) { ans = mid; high = mid - 1; // Continue searching left to find the lowest index } else if (midVal < key) { low = mid + 1; } else { high = mid - 1; } } return ans; }; /** Returns the largest index i such that val[i] <= key, or -1 if no such index exists. */ export const binarySearchLessOrEqual = (arr: T[], key: number, valueGetter: (x: T) => number) => { let low = 0; let high = arr.length - 1; let ans = -1; while (low <= high) { const mid = (low + (high - low + 1) / 2) | 0; const midVal = valueGetter(arr[mid]!); if (midVal <= key) { ans = mid; low = mid + 1; } else { high = mid - 1; } } return ans; }; /** Assumes the array is already sorted. */ export const insertSorted = (arr: T[], item: T, valueGetter: (x: T) => number) => { const insertionIndex = binarySearchLessOrEqual(arr, valueGetter(item), valueGetter); arr.splice(insertionIndex + 1, 0, item); // This even behaves correctly for the -1 case }; export const promiseWithResolvers = () => { let resolve: (value: T) => void; let reject: (reason: unknown) => void; const promise = new Promise((res, rej) => { resolve = res; reject = rej; }); return { promise, resolve: resolve!, reject: reject! }; }; export const removeItem = (arr: T[], item: T) => { const index = arr.indexOf(item); if (index !== -1) { arr.splice(index, 1); } }; export const findLast = (arr: T[], predicate: (x: T) => boolean) => { for (let i = arr.length - 1; i >= 0; i--) { if (predicate(arr[i]!)) { return arr[i]; } } return undefined; }; export const findLastIndex = (arr: T[], predicate: (x: T) => boolean) => { for (let i = arr.length - 1; i >= 0; i--) { if (predicate(arr[i]!)) { return i; } } return -1; }; /** * Sync or async iterable. * @group Miscellaneous * @public */ export type AnyIterable = | Iterable | AsyncIterable; export const toAsyncIterator = async function* (source: AnyIterable): AsyncGenerator { if (Symbol.iterator in source) { // @ts-expect-error Trust me yield* source[Symbol.iterator](); } else { // @ts-expect-error Trust me yield* source[Symbol.asyncIterator](); } }; export const validateAnyIterable = (iterable: AnyIterable) => { if (!(Symbol.iterator in iterable) && !(Symbol.asyncIterator in iterable)) { throw new TypeError('Argument must be an iterable or async iterable.'); } }; export const assertNever = (x: never) => { // eslint-disable-next-line @typescript-eslint/restrict-template-expressions throw new Error(`Unexpected value: ${x}`); }; export const getUint24 = (view: DataView, byteOffset: number, littleEndian: boolean) => { const byte1 = view.getUint8(byteOffset); const byte2 = view.getUint8(byteOffset + 1); const byte3 = view.getUint8(byteOffset + 2); if (littleEndian) { return byte1 | (byte2 << 8) | (byte3 << 16); } else { return (byte1 << 16) | (byte2 << 8) | byte3; } }; export const getInt24 = (view: DataView, byteOffset: number, littleEndian: boolean) => { // The left shift pushes the most significant bit into the sign bit region, and the subsequent right shift // then correctly interprets the sign bit. return getUint24(view, byteOffset, littleEndian) << 8 >> 8; }; export const setUint24 = (view: DataView, byteOffset: number, value: number, littleEndian: boolean) => { // Ensure the value is within 24-bit unsigned range (0 to 16777215) value = value >>> 0; // Convert to unsigned 32-bit value = value & 0xFFFFFF; // Mask to 24 bits if (littleEndian) { view.setUint8(byteOffset, value & 0xFF); view.setUint8(byteOffset + 1, (value >>> 8) & 0xFF); view.setUint8(byteOffset + 2, (value >>> 16) & 0xFF); } else { view.setUint8(byteOffset, (value >>> 16) & 0xFF); view.setUint8(byteOffset + 1, (value >>> 8) & 0xFF); view.setUint8(byteOffset + 2, value & 0xFF); } }; export const setInt24 = (view: DataView, byteOffset: number, value: number, littleEndian: boolean) => { // Ensure the value is within 24-bit signed range (-8388608 to 8388607) value = clamp(value, -8388608, 8388607); // Convert negative values to their 24-bit representation if (value < 0) { value = (value + 0x1000000) & 0xFFFFFF; } setUint24(view, byteOffset, value, littleEndian); }; export const setInt64 = (view: DataView, byteOffset: number, value: number, littleEndian: boolean) => { if (littleEndian) { view.setUint32(byteOffset + 0, value, true); view.setInt32(byteOffset + 4, Math.floor(value / 2 ** 32), true); } else { view.setInt32(byteOffset + 0, Math.floor(value / 2 ** 32), true); view.setUint32(byteOffset + 4, value, true); } }; /** * Calls a function on each value spat out by an async generator. The reason for writing this manually instead of * using a generator function is that the generator function queues return() calls - here, we forward them immediately. */ export const mapAsyncGenerator = ( generator: AsyncGenerator, map: (t: T) => U, ): AsyncGenerator => { return { async next() { const result = await generator.next(); if (result.done) { return { value: undefined, done: true }; } else { return { value: map(result.value), done: false }; } }, return() { return generator.return() as ReturnType['return']>; }, throw(error) { return generator.throw(error) as ReturnType['throw']>; }, [Symbol.asyncIterator]() { return this; }, }; }; export const clamp = (value: number, min: number, max: number) => { return Math.max(min, Math.min(max, value)); }; export const UNDETERMINED_LANGUAGE = 'und'; export const roundIfAlmostInteger = (value: number) => { const rounded = Math.round(value); if (Math.abs(value / rounded - 1) < 10 * Number.EPSILON) { return rounded; } else { return value; } }; export const roundToMultiple = (value: number, multiple: number) => { return Math.round(value / multiple) * multiple; }; export const roundToDivisor = (value: number, multiple: number) => { return Math.round(value * multiple) / multiple; }; export const floorToMultiple = (value: number, multiple: number) => { return Math.floor(value / multiple) * multiple; }; export const floorToDivisor = (value: number, multiple: number) => { return Math.floor(value * multiple) / multiple; }; export const ilog = (x: number) => { let ret = 0; while (x) { ret++; x >>= 1; } return ret; }; const ISO_639_2_REGEX = /^[a-z]{3}$/; export const isIso639Dash2LanguageCode = (x: string) => { return ISO_639_2_REGEX.test(x); }; // Since the result will be truncated, add a bit of eps to compensate for floating point errors export const SECOND_TO_MICROSECOND_FACTOR = 1e6 * (1 + Number.EPSILON); /** * Sets all keys K of T to be required. * @group Miscellaneous * @public */ export type SetRequired = T & Required>; /** * Recursively makes all properties of T readonly. * @group Miscellaneous * @public */ export type DeepReadonly = T extends object ? { readonly [K in keyof T]: DeepReadonly; } : T; /** * Sets all keys K of T to be optional. * @group Miscellaneous * @public */ export type SetOptional = Omit & Partial>; /** * Merges two RequestInit objects with special handling for headers. * Headers are merged case-insensitively, but original casing is preserved. * init2 headers take precedence and will override case-insensitive matches from init1. */ export const mergeRequestInit = (init1: RequestInit, init2: RequestInit): RequestInit => { const merged: RequestInit = { ...init1, ...init2 }; // Special handling for headers if (init1.headers || init2.headers) { const headers1 = init1.headers ? normalizeHeaders(init1.headers) : {}; const headers2 = init2.headers ? normalizeHeaders(init2.headers) : {}; const mergedHeaders = { ...headers1 }; // For each header in headers2, check if a case-insensitive match exists in mergedHeaders Object.entries(headers2).forEach(([key2, value2]) => { const existingKey = Object.keys(mergedHeaders).find( key1 => key1.toLowerCase() === key2.toLowerCase(), ); if (existingKey) { delete mergedHeaders[existingKey]; } mergedHeaders[key2] = value2; }); merged.headers = mergedHeaders; } return merged; }; /** Normalizes HeadersInit to a Record format. */ export const normalizeHeaders = (headers: HeadersInit): Record => { if (headers instanceof Headers) { const result: Record = {}; headers.forEach((value, key) => { result[key] = value; }); return result; } if (Array.isArray(headers)) { const result: Record = {}; headers.forEach(([key, value]) => { result[key] = value; }); return result; } return headers; }; export const retriedFetch = async ( fetchFn: typeof fetch, url: string | URL | Request, requestInit: RequestInit, getRetryDelay: (previousAttempts: number, error: unknown, url: string | URL | Request) => number | null, shouldStop: () => boolean, ) => { let attempts = 0; while (true) { try { return await fetchFn(url, requestInit); } catch (error) { if (shouldStop()) { throw error; } attempts++; const retryDelayInSeconds = getRetryDelay(attempts, error, url); if (retryDelayInSeconds === null) { throw error; } Logging._error('Retrying failed fetch. Error:', error); if (!Number.isFinite(retryDelayInSeconds) || retryDelayInSeconds < 0) { throw new TypeError('Retry delay must be a non-negative finite number.'); } if (retryDelayInSeconds > 0) { await wait(1000 * retryDelayInSeconds); } if (shouldStop()) { throw error; } } } }; export const computeRationalApproximation = (x: number, maxDenominator: number): Rational => { // Handle negative numbers const sign = x < 0 ? -1 : 1; x = Math.abs(x); let prevNumerator = 0, prevDenominator = 1; let currNumerator = 1, currDenominator = 0; // Continued fraction algorithm let remainder = x; while (true) { const integer = Math.floor(remainder); // Calculate next convergent const nextNumerator = integer * currNumerator + prevNumerator; const nextDenominator = integer * currDenominator + prevDenominator; if (nextDenominator > maxDenominator) { return { num: sign * currNumerator, den: currDenominator, }; } prevNumerator = currNumerator; prevDenominator = currDenominator; currNumerator = nextNumerator; currDenominator = nextDenominator; remainder = 1 / (remainder - integer); // Guard against precision issues if (!isFinite(remainder)) { break; } } return { num: sign * currNumerator, den: currDenominator, }; }; export class CallSerializer { currentPromise = Promise.resolve(); call(fn: () => Promise | void) { return this.currentPromise = this.currentPromise.then(fn); } } let isWebKitCache: boolean | null = null; export const isWebKit = () => { if (isWebKitCache !== null) { return isWebKitCache; } // This even returns true for WebKit-wrapping browsers such as Chrome on iOS return isWebKitCache = !!( typeof navigator !== 'undefined' && ( // eslint-disable-next-line @typescript-eslint/no-deprecated navigator.vendor?.match(/apple/i) // Or, in workers: || (/AppleWebKit/.test(navigator.userAgent) && !/Chrome/.test(navigator.userAgent)) || /\b(iPad|iPhone|iPod)\b/.test(navigator.userAgent) ) ); }; let isFirefoxCache: boolean | null = null; export const isFirefox = () => { if (isFirefoxCache !== null) { return isFirefoxCache; } return isFirefoxCache = typeof navigator !== 'undefined' && navigator.userAgent?.includes('Firefox'); }; let isChromiumCache: boolean | null = null; export const isChromium = () => { if (isChromiumCache !== null) { return isChromiumCache; } return isChromiumCache = !!( typeof navigator !== 'undefined' // eslint-disable-next-line @typescript-eslint/no-deprecated && (navigator.vendor?.includes('Google Inc') || /Chrome/.test(navigator.userAgent)) ); }; let chromiumVersionCache: number | null = null; export const getChromiumVersion = () => { if (chromiumVersionCache !== null) { return chromiumVersionCache; } if (typeof navigator === 'undefined') { return null; } const match = /\bChrome\/(\d+)/.exec(navigator.userAgent); if (!match) { return null; } return chromiumVersionCache = Number(match[1]!); }; /** * T or a promise that resolves to T. * @group Miscellaneous * @public */ export type MaybePromise = T | Promise; /** Acts like `??` except the condition is -1 and not null/undefined. */ export const coalesceIndex = (a: number, b: number) => { return a !== -1 ? a : b; }; export const closedIntervalsOverlap = (startA: number, endA: number, startB: number, endB: number) => { return startA <= endB && startB <= endA; }; type KeyValuePair> = { [K in keyof T]-?: { key: K; value: T[K] extends infer R | undefined ? R : T[K]; } }[keyof T]; export const keyValueIterator = function* >(object: T) { for (const key in object) { const value = object[key]; if (value === undefined) { continue; } yield { key, value } as KeyValuePair; } }; export const imageMimeTypeToExtension = (mimeType: string) => { switch (mimeType.toLowerCase()) { case 'image/jpeg': case 'image/jpg': return '.jpg'; case 'image/png': return '.png'; case 'image/gif': return '.gif'; case 'image/webp': return '.webp'; case 'image/bmp': return '.bmp'; case 'image/svg+xml': return '.svg'; case 'image/tiff': return '.tiff'; case 'image/avif': return '.avif'; case 'image/x-icon': case 'image/vnd.microsoft.icon': return '.ico'; default: return null; } }; export const base64ToBytes = (base64: string) => { const decoded = atob(base64); const bytes = new Uint8Array(decoded.length); for (let i = 0; i < decoded.length; i++) { bytes[i] = decoded.charCodeAt(i); } return bytes; }; export const bytesToBase64 = (bytes: Uint8Array) => { let string = ''; for (let i = 0; i < bytes.length; i++) { string += String.fromCharCode(bytes[i]!); } return btoa(string); }; export const uint8ArraysAreEqual = (a: Uint8Array, b: Uint8Array) => { if (a.length !== b.length) { return false; } for (let i = 0; i < a.length; i++) { if (a[i] !== b[i]) { return false; } } return true; }; export const polyfillSymbolDispose = () => { // https://www.typescriptlang.org/docs/handbook/release-notes/typescript-5-2.html // @ts-expect-error Readonly Symbol.dispose ??= Symbol('Symbol.dispose'); }; export const isNumber = (x: unknown) => { return typeof x === 'number' && !Number.isNaN(x); }; /** * A path to a file. File paths can be relative or absolute, and be local paths or full URLs. Paths must be POSIX-like, * using `/` as the separator. * * Examples of valid paths: * - `'video.mp4'` * - `'path/to/video.mp4'` * - `'./video.mp4'` * - `'../video.mp4'` * - `'/path/to/video.mp4'` * - `'https://example.com/video.mp4'` * - `'file:///home/user/video.mp4'` * - `'video.mp4?key=foo'` * * @group Miscellaneous * @public */ export type FilePath = string; export const joinPaths = (basePath: FilePath, relativePath: FilePath) => { // If relativePath is a full URL with protocol, return it as-is if (relativePath.includes('://')) { return relativePath; } // Strip query parameters from URL base paths so their contents don't mess up the join if (basePath.includes('://')) { const queryIndex = basePath.indexOf('?'); if (queryIndex !== -1) { basePath = basePath.slice(0, queryIndex); } } let result: string; if (relativePath.startsWith('/')) { const protocolIndex = basePath.indexOf('://'); if (protocolIndex === -1) { result = relativePath; } else { const pathStart = basePath.indexOf('/', protocolIndex + 3); if (pathStart === -1) { result = basePath + relativePath; } else { result = basePath.slice(0, pathStart) + relativePath; } } } else { const lastSlash = basePath.lastIndexOf('/'); if (lastSlash === -1) { result = relativePath; } else { result = basePath.slice(0, lastSlash + 1) + relativePath; } } // Normalize ./ and ../ let prefix = ''; const protocolIndex = result.indexOf('://'); if (protocolIndex !== -1) { const pathStart = result.indexOf('/', protocolIndex + 3); if (pathStart !== -1) { prefix = result.slice(0, pathStart); result = result.slice(pathStart); } } const segments = result.split('/'); const normalized: string[] = []; for (const segment of segments) { if (segment === '..') { normalized.pop(); } else if (segment !== '.') { normalized.push(segment); } } return prefix + normalized.join('/'); }; export const arrayCount = (array: T[], predicate: (item: T) => boolean) => { let count = 0; for (let i = 0; i < array.length; i++) { if (predicate(array[i]!)) { count++; } } return count; }; export const arrayArgmin = (array: T[], getValue: (item: T) => number): number => { let minIndex = -1; let minValue = Infinity; for (let i = 0; i < array.length; i++) { const value = getValue(array[i]!); if (value < minValue) { minValue = value; minIndex = i; } } return minIndex; }; export const arrayArgmax = (array: T[], getValue: (item: T) => number): number => { let maxIndex = -1; let maxValue = -Infinity; for (let i = 0; i < array.length; i++) { const value = getValue(array[i]!); if (value > maxValue) { maxValue = value; maxIndex = i; } } return maxIndex; }; /** * A rational number; a ratio of two integers. * @group Miscellaneous * @public */ export type Rational = { /** The numerator of the rational number. */ num: number; /** The denominator of the rational number. */ den: number; }; export const simplifyRational = (rational: Rational): Rational => { assert(Number.isInteger(rational.num)); assert(Number.isInteger(rational.den)); assert(rational.den !== 0); let a = Math.abs(rational.num); let b = Math.abs(rational.den); // Euclidean algorithm while (b !== 0) { const t = a % b; a = b; b = t; } const gcd = a || 1; return { num: rational.num / gcd, den: rational.den / gcd, }; }; /** * Specifies a rectangular region where all quantities must be non-negative integers. * @group Miscellaneous * @public */ export type Rectangle = { /** The distance in pixels to the left edge of the rectangle. */ left: number; /** The distance in pixels to the top edge of the rectangle. */ top: number; /** The width in pixels of the rectangle. */ width: number; /** The height in pixels of the rectangle. */ height: number; }; export const validateRectangle = (rect: Rectangle, propertyPath: string) => { if (typeof rect !== 'object' || !rect) { throw new TypeError(`${propertyPath} must be an object.`); } if (!Number.isInteger(rect.left) || rect.left < 0) { throw new TypeError(`${propertyPath}.left must be a non-negative integer.`); } if (!Number.isInteger(rect.top) || rect.top < 0) { throw new TypeError(`${propertyPath}.top must be a non-negative integer.`); } if (!Number.isInteger(rect.width) || rect.width < 0) { throw new TypeError(`${propertyPath}.width must be a non-negative integer.`); } if (!Number.isInteger(rect.height) || rect.height < 0) { throw new TypeError(`${propertyPath}.height must be a non-negative integer.`); } }; export type NonFunctionKeys = { [K in keyof T]-?: T[K] extends ((...args: never[]) => unknown) ? never : K }[keyof T]; export type UnthrottledTimerHandle = { id: ReturnType | number; }; type UnthrottledTimerMessage = | { type: 'set-timeout'; timerId: number; delay: number } | { type: 'set-interval'; timerId: number; delay: number } | { type: 'clear-timeout'; timerId: number } | { type: 'clear-interval'; timerId: number }; type UnthrottledTimerEvent = { type: 'fire'; timerId: number }; let unthrottledTimerWorker: Worker | undefined; let nextUnthrottledTimerId = 1; const unthrottledTimeoutCallbacks = new Map void>(); const unthrottledIntervalCallbacks = new Map void>(); const shouldUseNativeTimers = () => { return typeof window === 'undefined'; }; const unthrottledTimerWorkerMain = () => { const timeoutHandles = new Map>(); const intervalHandles = new Map>(); self.onmessage = (event: MessageEvent) => { const message = event.data; switch (message.type) { case 'set-timeout': { const handle = setTimeout(() => { timeoutHandles.delete(message.timerId); self.postMessage({ type: 'fire', timerId: message.timerId }); }, message.delay); timeoutHandles.set(message.timerId, handle); }; break; case 'set-interval': { const handle = setInterval(() => { self.postMessage({ type: 'fire', timerId: message.timerId }); }, message.delay); intervalHandles.set(message.timerId, handle); }; break; case 'clear-timeout': { const handle = timeoutHandles.get(message.timerId); if (handle !== undefined) { clearTimeout(handle); timeoutHandles.delete(message.timerId); } }; break; case 'clear-interval': { const handle = intervalHandles.get(message.timerId); if (handle !== undefined) { clearInterval(handle); intervalHandles.delete(message.timerId); } }; break; } }; }; const getUnthrottledTimerWorker = () => { if (unthrottledTimerWorker) { return unthrottledTimerWorker; } const workerSource = `(${unthrottledTimerWorkerMain.toString()})();`; const workerURL = URL.createObjectURL(new Blob([workerSource], { type: 'text/javascript' })); unthrottledTimerWorker = new Worker(workerURL); URL.revokeObjectURL(workerURL); unthrottledTimerWorker.onmessage = (event: MessageEvent) => { const message = event.data; const timeoutCallback = unthrottledTimeoutCallbacks.get(message.timerId); if (timeoutCallback) { unthrottledTimeoutCallbacks.delete(message.timerId); timeoutCallback(); return; } const intervalCallback = unthrottledIntervalCallbacks.get(message.timerId); if (intervalCallback) { intervalCallback(); } }; return unthrottledTimerWorker; }; export const setTimeoutUnthrottled = ( // eslint-disable-next-line @typescript-eslint/no-unsafe-function-type callback: Function, delay: number, ): UnthrottledTimerHandle => { if (shouldUseNativeTimers()) { return { id: setTimeout(callback, delay) }; } const timerId = nextUnthrottledTimerId++; unthrottledTimeoutCallbacks.set(timerId, () => { (callback as () => void)(); }); getUnthrottledTimerWorker().postMessage({ type: 'set-timeout', timerId, delay, } satisfies UnthrottledTimerMessage); return { id: timerId }; }; export const clearTimeoutUnthrottled = (timer: UnthrottledTimerHandle) => { if (shouldUseNativeTimers()) { clearTimeout(timer.id); return; } assert(typeof timer.id === 'number'); unthrottledTimeoutCallbacks.delete(timer.id); getUnthrottledTimerWorker().postMessage({ type: 'clear-timeout', timerId: timer.id, } satisfies UnthrottledTimerMessage); }; export const setIntervalUnthrottled = ( // eslint-disable-next-line @typescript-eslint/no-unsafe-function-type callback: Function, delay: number, ): UnthrottledTimerHandle => { if (shouldUseNativeTimers()) { return { id: setInterval(callback, delay) }; } const timerId = nextUnthrottledTimerId++; unthrottledIntervalCallbacks.set(timerId, () => { (callback as () => void)(); }); getUnthrottledTimerWorker().postMessage({ type: 'set-interval', timerId, delay, } satisfies UnthrottledTimerMessage); return { id: timerId }; }; export const clearIntervalUnthrottled = (timer: UnthrottledTimerHandle) => { if (shouldUseNativeTimers()) { clearInterval(timer.id); return; } assert(typeof timer.id === 'number'); unthrottledIntervalCallbacks.delete(timer.id); getUnthrottledTimerWorker().postMessage({ type: 'clear-interval', timerId: timer.id, } satisfies UnthrottledTimerMessage); }; export const wait = (ms: number) => { return new Promise(resolve => setTimeout(resolve, ms)); }; export const rejectAfter = (ms: number, message = 'Promise rejected') => { return new Promise((_, reject) => { setTimeout(() => reject(new Error(message)), ms); }); }; export const toArray = (x: T | T[]) => { if (Array.isArray(x)) { return x; } else { return [x]; } }; /** * Options for {@link EventEmitter.on}. * * @group Miscellaneous * @public */ export type EventListenerOptions = { /** If `true`, the listener will be automatically removed after being called once. Defaults to `false`. */ once?: boolean; }; /** * A class that manages event listeners and dispatches events to them. * * @group Miscellaneous * @public */ export class EventEmitter> { /** @internal */ _listeners = new Map unknown; once: boolean }>>(); /** Registers a listener for the given event. Returns a function that, when called, removes the listener again. */ on( event: K, listener: (data: TEvents[K]) => unknown, options?: EventListenerOptions, ): () => void { if (!this._listeners.has(event)) { this._listeners.set(event, new Set()); } const entry = { fn: listener as (data: never) => void, once: options?.once ?? false }; this._listeners.get(event)!.add(entry); return () => { this._listeners.get(event)?.delete(entry); }; } /** @internal */ _emit( ...args: TEvents[K] extends void ? [event: K] : [event: K, data: TEvents[K]] ): void { const [event, data] = args; const listeners = this._listeners.get(event); if (!listeners) { return; } for (const entry of listeners) { try { (entry.fn as (data: unknown) => void)(data); } catch (error) { // Deliberately not routed through Logging here: Logging emits via an EventEmitter, so a throwing // log listener would recurse straight back into this handler. console.error(error); } if (entry.once) { listeners.delete(entry); } } } } export const ceilToMultipleOfTwo = (value: number) => Math.ceil(value / 2) * 2; /** * Utility class for running async functions in parallel up to a certain level of parallelism. Can be used to apply * backpressure only if the concurrency level would be exceeded. * * @group Miscellaneous * @public */ export class ConcurrentRunner { /** @internal */ _queue: Promise[] = []; /** @internal */ _errored = false; /** * The maximum number of in-flight promises. You can also think of it as the "high water mark". * You can set this value to dynamically change the level of parallelism. */ parallelism: number; constructor(parallelism: number) { this.parallelism = parallelism; } /** Whether any function has errored. The runner is effectively bricked if this is `true`, by design. */ get errored() { return this._errored; } /** The number of tasks currently running. */ get inFlightCount() { return this._queue.length; } /** * Schedules an async function to be run. If the maximum allowed level of parallelism has not yet been reached, * the function will be executed immediately and `run()` will resolve immediately. Otherwise, the function will be * called as soon as any currently-running function finishes, and `run()` will only resolve then. * * Throws if the runner is errored. */ async run(fn: () => Promise) { if (this._errored) { await Promise.race(this._queue); // Will surface the error } while (this._queue.length >= this.parallelism) { await Promise.race(this._queue); } const promise = fn(); this._queue.push(promise); void promise .then(() => removeItem(this._queue, promise)) .catch(() => this._errored = true); } /** Waits for all currently running functions to finish. Throws if the runner is errored. */ async flush() { await Promise.all(this._queue); } } export const isRecordStringString = (value: unknown): value is Record => { return value !== null && typeof value === 'object' && Object.getPrototypeOf(value) === Object.prototype && Object.values(value).every(x => typeof x === 'string'); }; ===== src/logging.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { EventEmitter, type EventListenerOptions } from './misc'; /** * Controls how much information Mediabunny prints to the console. Higher levels include all lower levels. * * @group Logging * @public */ export enum LogLevel { /** Nothing is printed to the console. */ Silent = 0, /** Only errors are printed. */ Errors = 1, /** Errors and warnings are printed. */ Warnings = 2, /** Errors, warnings, and informational messages are printed. */ Info = 3, } /** * The events emitted by {@link Logging}. Each event carries the same arguments that were passed to the corresponding * log call. * * @group Logging * @public */ export type LoggingEvents = { /** Emitted before an error is logged. */ error: unknown[]; /** Emitted before a warning is logged. */ warn: unknown[]; /** Emitted before an informational message is logged. */ info: unknown[]; }; /** * Mediabunny's central logging singleton. Use {@link Logging.level} to control how much is printed to the console, * and subscribe to log events using {@link Logging.on}. * * Having manual control over logging is useful for command-line applications where you want full say over the output. * * @group Logging * @public */ export class Logging { private constructor() {} /** @internal */ static _level: LogLevel = LogLevel.Info; /** @internal */ static _emitterInstance: EventEmitter | null = null; /** The current log level. Defaults to {@link LogLevel.Info}. */ static get level() { return Logging._level; } static set level(value: LogLevel) { if ( value !== LogLevel.Silent && value !== LogLevel.Errors && value !== LogLevel.Warnings && value !== LogLevel.Info ) { throw new TypeError('Invalid log level. Use one of the values of the LogLevel enum.'); } Logging._level = value; } /** @internal */ static get _emitter() { // Created lazily to avoid touching the EventEmitter binding at module-eval time return Logging._emitterInstance ??= new EventEmitter(); } /** Registers a listener for a log event. Returns a function that, when called, removes the listener again. */ static on( event: K, listener: (data: LoggingEvents[K]) => unknown, options?: EventListenerOptions, ) { return Logging._emitter.on(event, listener, options); } /** @internal */ static _error(...args: unknown[]) { Logging._emitter._emit('error', args); if (Logging._level >= LogLevel.Errors) { console.error(...args); } } /** @internal */ static _warn(...args: unknown[]) { Logging._emitter._emit('warn', args); if (Logging._level >= LogLevel.Warnings) { console.warn(...args); } } /** @internal */ static _info(...args: unknown[]) { Logging._emitter._emit('info', args); if (Logging._level >= LogLevel.Info) { console.info(...args); } } } ===== src/media-source.ts ===== /*! * Copyright (c) 2026-present, Vanilagy and contributors * * This Source Code Form is subject to the terms of the Mozilla Public * License, v. 2.0. If a copy of the MPL was not distributed with this * file, You can obtain one at https://mozilla.org/MPL/2.0/. */ import { buildAacAudioSpecificConfig, parseAacAudioSpecificConfig } from '../shared/aac-misc'; import { AUDIO_CODECS, AudioCodec, MediaCodec, parsePcmCodec, PCM_AUDIO_CODECS, PcmAudioCodec, SUBTITLE_CODECS, SubtitleCodec, VIDEO_CODECS, VideoCodec, } from './codec'; import { OutputAudioTrack, OutputSubtitleTrack, OutputTrack, OutputVideoTrack } from './output'; import { assert, assertNever, binarySearchLessOrEqual, CallSerializer, clamp, clearIntervalUnthrottled, floorToDivisor, last, promiseWithResolvers, roundToDivisor, setInt24, setIntervalUnthrottled, setUint24, toUint8Array, UnthrottledTimerHandle, } from './misc'; import { Muxer } from './muxer'; import { SubtitleParser } from './subtitles'; import { toAlaw, toUlaw } from './pcm'; import { CustomVideoEncoder, CustomAudioEncoder, customVideoEncoders, customAudioEncoders, } from './custom-coder'; import { EncodedPacket, EncodedPacketSideData, PacketType } from './packet'; import { AudioSample, audioSampleToInterleavedFormat, toInterleavedAudioFormat, VideoSample, VideoSamplePixelFormat, } from './sample'; import { AudioEncodingConfig, buildAudioEncoderConfig, buildVideoEncoderConfig, validateAudioEncodingConfig, validateVideoEncodingConfig, VideoEncodingConfig, } from './encode'; import { AudioResampler } from './resample'; import { determineVideoPacketType } from './codec-data'; import { Logging } from './logging'; /** * Base class for media sources. Media sources are used to add media samples to an output file. * @group Media sources * @public */ export abstract class MediaSource { /** @internal */ abstract readonly _codec: MediaCodec; /** @internal */ _connectedTrack: OutputTrack | null = null; /** @internal */ _closingPromise: Promise | null = null; /** @internal */ _closed = false; /** @internal */ _ensureValidAdd() { if (!this._connectedTrack) { throw new Error('Source is not connected to an output track.'); } if (this._connectedTrack.output.state === 'canceled') { throw new Error('Output has been canceled.'); } if (this._connectedTrack.output.state === 'finalizing' || this._connectedTrack.output.state === 'finalized') { throw new Error('Output has been finalized.'); } if (this._connectedTrack.output.state === 'pending') { throw new Error('Output has not started.'); } if (this._closed) { throw new Error('Source is closed.'); } } /** @internal */ async _start() {} /** @internal */ // eslint-disable-next-line @typescript-eslint/no-unused-vars async _flushAndClose(forceClose: boolean) {} /** * Closes this source. This prevents future samples from being added and signals to the output file that no further * samples will come in for this track. Calling `.close()` is optional but recommended after adding the * last sample - for improved performance and reduced memory usage. */ close() { if (this._closingPromise) { return; } const connectedTrack = this._connectedTrack; if (!connectedTrack) { throw new Error('Cannot call close without connecting the source to an output track.'); } if (connectedTrack.output.state === 'pending') { throw new Error('Cannot call close before output has been started.'); } this._closingPromise = (async () => { await this._flushAndClose(false); this._closed = true; if (connectedTrack.output.state === 'finalizing' || connectedTrack.output.state === 'finalized') { return; } connectedTrack.output._muxer.onTrackClose(connectedTrack); })(); } /** @internal */ async _flushOrWaitForOngoingClose(forceClose: boolean) { return this._closingPromise ??= (async () => { await this._flushAndClose(forceClose); this._closed = true; })(); } } /** * Base class for video sources - sources for video tracks. * @group Media sources * @public */ export abstract class VideoSource extends MediaSource { /** @internal */ override _connectedTrack: OutputVideoTrack | null = null; /** @internal */ override readonly _codec: VideoCodec; /** Internal constructor. */ constructor(codec: VideoCodec) { super(); if (!VIDEO_CODECS.includes(codec)) { throw new TypeError(`Invalid video codec '${codec}'. Must be one of: ${VIDEO_CODECS.join(', ')}.`); } this._codec = codec; } } const maybeEnsureIsKeyPacket = (track: OutputVideoTrack, packet: EncodedPacket) => { if (track.metadata.hasOnlyKeyPackets && packet.type !== 'key') { throw new Error('Cannot add non-key packets to a hasOnlyKeyPackets video track.'); } }; /** * The most basic video source; can be used to directly pipe encoded packets into the output file. * @group Media sources * @public */ export class EncodedVideoPacketSource extends VideoSource { /** Creates a new {@link EncodedVideoPacketSource} whose packets are encoded using `codec`. */ constructor(codec: VideoCodec) { super(codec); } /** * Adds an encoded packet to the output video track. Packets must be added in *decode order*, while a packet's * timestamp must be its *presentation timestamp*. B-frames are handled automatically. * * @param meta - Additional metadata from the encoder. You should pass this for the first call, including a valid * decoder config. * * @returns A Promise that resolves once the output is ready to receive more samples. You should await this Promise * to respect writer and encoder backpressure. */ add(packet: EncodedPacket, meta?: EncodedVideoChunkMetadata) { if (!(packet instanceof EncodedPacket)) { throw new TypeError('packet must be an EncodedPacket.'); } if (packet.isMetadataOnly) { throw new TypeError('Metadata-only packets cannot be added.'); } if (meta !== undefined && (!meta || typeof meta !== 'object')) { throw new TypeError('meta, when provided, must be an object.'); } this._ensureValidAdd(); maybeEnsureIsKeyPacket(this._connectedTrack!, packet); return this._connectedTrack!.output._muxer.addEncodedVideoPacket(this._connectedTrack!, packet, meta); } } class VideoEncoderWrapper { private ensureEncoderPromise: Promise | null = null; private encoderInitialized = false; private encoder: VideoEncoder | null = null; private muxer: Muxer | null = null; private lastMultipleOfKeyFrameInterval = -1; private emittedEncoderPackets = 0; // Tracks the input dimensions of the first frame private codedWidth: number | null = null; private codedHeight: number | null = null; // Tracks the output dimensions of the first frame (used to lock dimensions for fill/contain/cover) private outputWidth: number | null = null; private outputHeight: number | null = null; // Frame rate normalization state private frameRateLastSample: VideoSample | null = null; private frameRateLastTimestamp: number | null = null; private frameRateLastEndTimestamp: number | null = null; // VideoEncoder converts everything to microseconds, so we need to do some bookkeeping to restore the original // timing information private preciseTimings: { microsecondTimestamp: number; timestamp: number; duration: number; timestampIsValid: boolean; durationIsValid: boolean; }[] = []; private customEncoder: CustomVideoEncoder | null = null; private customEncoderCallSerializer = new CallSerializer(); private customEncoderQueueSize = 0; // Alpha stuff private alphaEncoder: VideoEncoder | null = null; private splitter: ColorAlphaSplitter | null = null; private splitterCreationFailed = false; private alphaFrameQueue: (VideoFrame | null)[] = []; /** * Encoders typically throw their errors "out of band", meaning asynchronously in some other execution context. * However, we want to surface these errors to the user within the normal control flow, so they don't go uncaught. * So, we keep track of the encoder error and throw it as soon as we get the chance. */ private error: Error | null = null; private closed = false; private lastMuxerPromise: Promise = Promise.resolve(); constructor(private source: VideoSource, private encodingConfig: VideoEncodingConfig) {} async add(videoSample: VideoSample, shouldClose: boolean, encodeOptions?: VideoEncoderEncodeOptions) { const originalSample = videoSample; try { this.checkForEncoderError(); this.source._ensureValidAdd(); const config = this.encodingConfig; const sizeChangeBehavior = config.sizeChangeBehavior ?? 'deny'; let isSizeChange = false; // Ensure video sample size remains constant or handle the change if (this.codedWidth !== null && this.codedHeight !== null) { if (videoSample.codedWidth !== this.codedWidth || videoSample.codedHeight !== this.codedHeight) { isSizeChange = true; if (sizeChangeBehavior === 'deny') { throw new Error( `Video sample size must remain constant. Expected ${this.codedWidth}x${this.codedHeight},` + ` got ${videoSample.codedWidth}x${videoSample.codedHeight}. To allow the sample size to` + ` change over time, set \`sizeChangeBehavior\` to a value other than 'deny' in the` + ` encoding options.`, ); } } } else { this.codedWidth = videoSample.codedWidth; this.codedHeight = videoSample.codedHeight; } // Determine if we need to apply transformations via canvas const hasTransformConfig = config.transform?.width !== undefined || config.transform?.height !== undefined || config.transform?.rotate !== undefined || config.transform?.crop !== undefined || config.transform?.force === true; const needsTransform = hasTransformConfig || (isSizeChange && sizeChangeBehavior !== 'passThrough'); if (needsTransform) { let targetWidth = config.transform?.width; let targetHeight = config.transform?.height; let appliedFit: 'fill' | 'contain' | 'cover' = config.transform?.fit ?? 'fill'; // If the size changed and behavior is fill/contain/cover, lock to the original output dimensions if (isSizeChange && sizeChangeBehavior !== 'passThrough') { assert(this.outputWidth); assert(this.outputHeight); assert(sizeChangeBehavior !== 'deny'); targetWidth = this.outputWidth!; targetHeight = this.outputHeight!; appliedFit = sizeChangeBehavior; } const transformed = await videoSample.transform({ width: targetWidth, height: targetHeight, roundDimensionsTo: 2, crop: config.transform?.crop, rotate: config.transform?.rotate, fit: appliedFit, alpha: config.alpha, }); // Save the output dimensions of the first frame if (this.outputWidth === null || this.outputHeight === null) { this.outputWidth = transformed.displayWidth; this.outputHeight = transformed.displayHeight; } if (shouldClose) { videoSample.close(); } videoSample = transformed; shouldClose = true; } else { // If no canvas is needed, we still need to record the output dimensions for the first frame if (this.outputWidth === null || this.outputHeight === null) { this.outputWidth = videoSample.codedWidth; this.outputHeight = videoSample.codedHeight; } } const frameRate = config.transform?.frameRate; if (frameRate !== undefined) { // Apply frame rate normalization const originalEndTimestamp = videoSample.timestamp + videoSample.duration; const alignedTimestamp = floorToDivisor(videoSample.timestamp, frameRate); if (this.frameRateLastSample !== null) { if (alignedTimestamp <= this.frameRateLastTimestamp!) { // Same frame rate slot, replace stored sample with the newer one this.frameRateLastSample.close(); this.frameRateLastSample = videoSample.clone(); this.frameRateLastEndTimestamp = originalEndTimestamp; return; } else { // Pad the gap by repeating the previous frame await this.padFrameRate(alignedTimestamp, encodeOptions); } } // Clone if the sample is still the user's, to avoid mutating externally-owned data if (videoSample === originalSample) { videoSample = videoSample.clone(); shouldClose = true; } videoSample.setTimestamp(alignedTimestamp); videoSample.setDuration(1 / frameRate); this.frameRateLastSample?.close(); this.frameRateLastSample = videoSample.clone(); this.frameRateLastTimestamp = alignedTimestamp; this.frameRateLastEndTimestamp = originalEndTimestamp; } await this.processAndEncode(videoSample, encodeOptions); } finally { if (shouldClose) { videoSample.close(); } } } /** * Runs the process function (if any) and encodes the resulting samples. */ private async processAndEncode( videoSample: VideoSample, encodeOptions?: VideoEncoderEncodeOptions, ) { const config = this.encodingConfig; let samplesToEncode: VideoSample[]; // Apply the user-defined process function, if any if (config.transform?.process) { let processed = config.transform.process(videoSample); if (processed instanceof Promise) { processed = await processed; } if (processed === null) { return; } if (!Array.isArray(processed)) { processed = [processed]; } samplesToEncode = processed.map((x) => { if (x instanceof VideoSample) { return x; } if (typeof VideoFrame !== 'undefined' && x instanceof VideoFrame) { return new VideoSample(x); } // Calling the VideoSample constructor here will automatically handle input validation for us // (it throws for any non-legal argument). return new VideoSample(x as CanvasImageSource, { timestamp: videoSample.timestamp, duration: videoSample.duration, }); }); } else { samplesToEncode = [videoSample]; } try { for (const sampleToEncode of samplesToEncode) { if (!this.encoderInitialized) { if (!this.ensureEncoderPromise) { this.ensureEncoder(sampleToEncode); } // No, this "if" statement is not useless. Sometimes, the above call to // `ensureEncoder` might have synchronously completed and the encoder is // already initialized. In this case, we don't need to await the promise // anymore. This also fixes nasty async race condition bugs when multiple // code paths are calling this method: It's important that the call that // initialized the encoder go through this code first. if (!this.encoderInitialized) { await this.ensureEncoderPromise; } } assert(this.encoderInitialized); if (this.closed) { break; } const keyFrameInterval = this.encodingConfig.keyFrameInterval ?? 2; const multipleOfKeyFrameInterval = Math.floor(sampleToEncode.timestamp / keyFrameInterval); const mergedEncodeOptions = { ...sampleToEncode.encodeOptions, ...encodeOptions }; const finalEncodeOptions = { ...mergedEncodeOptions, keyFrame: mergedEncodeOptions.keyFrame !== undefined ? mergedEncodeOptions.keyFrame // Ensure a key frame every keyFrameInterval seconds. It is important that all video tracks // follow the same "key frame" rhythm, because aligned key frames are required to start new // fragments in ISOBMFF or clusters in Matroska (or at least desirable). : keyFrameInterval === 0 || multipleOfKeyFrameInterval !== this.lastMultipleOfKeyFrameInterval, }; this.lastMultipleOfKeyFrameInterval = multipleOfKeyFrameInterval; this.encodingConfig.onEncodedSample?.(sampleToEncode); if (this.customEncoder) { this.customEncoderQueueSize++; // We clone the sample so it cannot be closed on us from the outside before it reaches the encoder const clonedSample = sampleToEncode.clone(); const promise = this.customEncoderCallSerializer .call(() => this.customEncoder!.encode(clonedSample, finalEncodeOptions)) .then(() => this.customEncoderQueueSize--) .catch((error: Error) => this.error ??= error) .finally(() => { clonedSample.close(); }); if (this.customEncoderQueueSize >= 4) { await promise; } } else { assert(this.encoder); const videoFrame = sampleToEncode.toVideoFrame(); const preciseTimingIndex = binarySearchLessOrEqual( this.preciseTimings, videoFrame.timestamp, x => x.microsecondTimestamp, ); const existingEntry = preciseTimingIndex !== -1 ? this.preciseTimings[preciseTimingIndex] : null; if (existingEntry && existingEntry.microsecondTimestamp === videoFrame.timestamp) { if (existingEntry.timestamp !== sampleToEncode.timestamp) { // Mapping isn't unique, can't use the timestamp existingEntry.timestampIsValid = false; } if (existingEntry.duration !== sampleToEncode.duration) { // Mapping isn't unique, can't use the duration existingEntry.durationIsValid = false; } } else { this.preciseTimings.splice(preciseTimingIndex + 1, 0, { microsecondTimestamp: videoFrame.timestamp, timestamp: sampleToEncode.timestamp, duration: sampleToEncode.duration, timestampIsValid: true, durationIsValid: true, }); // Make sure it doesn't grow indefinitely if (this.preciseTimings.length > 128) { this.preciseTimings.shift(); } } if (!this.alphaEncoder) { // No alpha encoder, simple case this.encoder.encode(videoFrame, finalEncodeOptions); videoFrame.close(); } else { // We're expected to encode alpha as well const frameDefinitelyHasNoAlpha = !!videoFrame.format && !videoFrame.format.includes('A'); if (frameDefinitelyHasNoAlpha || this.splitterCreationFailed) { this.alphaFrameQueue.push(null); this.encoder.encode(videoFrame, finalEncodeOptions); videoFrame.close(); } else { const width = videoFrame.displayWidth; const height = videoFrame.displayHeight; if (!this.splitter) { this.splitter = new ColorAlphaSplitter(width, height); } // The splitter takes ownership, so no need to close the frames ourselves const { colorFrame, alphaFrame } = await this.splitter.update(videoFrame); this.alphaFrameQueue.push(alphaFrame); this.encoder.encode(colorFrame, finalEncodeOptions); colorFrame.close(); } } // We need to do this after sending the frame to the encoder as the frame otherwise might be closed if (this.encoder.encodeQueueSize >= 4) { await new Promise(resolve => this.encoder!.addEventListener('dequeue', resolve, { once: true }), ); } } await this.lastMuxerPromise; // Allow the writer to apply backpressure } } finally { for (const sample of samplesToEncode) { if (sample !== videoSample) { sample.close(); } } } } /** Repeats the last frame rate sample to fill the gap up to the given timestamp. */ private async padFrameRate(until: number, encodeOptions?: VideoEncoderEncodeOptions) { const frameRate = this.encodingConfig.transform!.frameRate!; assert(this.frameRateLastSample); const frameDifference = Math.round((until - this.frameRateLastTimestamp!) * frameRate); for (let i = 1; i < frameDifference; i++) { const sample = this.frameRateLastSample.clone(); sample.setTimestamp(this.frameRateLastTimestamp! + i / frameRate); sample.setDuration(1 / frameRate); await this.processAndEncode(sample, encodeOptions); sample.close(); } } private ensureEncoder(videoSample: VideoSample) { this.ensureEncoderPromise = (async () => { const encoderConfig = buildVideoEncoderConfig({ ...this.encodingConfig, width: videoSample.codedWidth, height: videoSample.codedHeight, squarePixelWidth: videoSample.squarePixelWidth, squarePixelHeight: videoSample.squarePixelHeight, framerate: this.source._connectedTrack?.metadata.frameRate, }); this.encodingConfig.onEncoderConfig?.(encoderConfig); const MatchingCustomEncoder = customVideoEncoders.find(x => x.supports( this.encodingConfig.codec, encoderConfig, )); if (MatchingCustomEncoder) { // @ts-expect-error "Can't create instance of abstract class 🤓" this.customEncoder = new MatchingCustomEncoder() as CustomVideoEncoder; // @ts-expect-error It's technically readonly this.customEncoder.codec = this.encodingConfig.codec; // @ts-expect-error It's technically readonly this.customEncoder.config = encoderConfig; // @ts-expect-error It's technically readonly this.customEncoder.onPacket = (packet, meta) => { if (!(packet instanceof EncodedPacket)) { throw new TypeError('The first argument passed to onPacket must be an EncodedPacket.'); } if (meta !== undefined && (!meta || typeof meta !== 'object')) { throw new TypeError('The second argument passed to onPacket must be an object or undefined.'); } maybeEnsureIsKeyPacket(this.source._connectedTrack!, packet); this.encodingConfig.onEncodedPacket?.(packet, meta); this.lastMuxerPromise = this.muxer!.addEncodedVideoPacket(this.source._connectedTrack!, packet, meta) .catch((error) => { this.error ??= error; }); }; await this.customEncoder.init(); } else { if (typeof VideoEncoder === 'undefined') { throw new Error('VideoEncoder is not supported by this browser.'); } encoderConfig.alpha = 'discard'; // Since we handle alpha ourselves if (this.encodingConfig.alpha === 'keep') { // Encoding alpha requires using two parallel encoders, so we need to make sure they stay in sync // and that neither of them drops frames. Setting latencyMode to 'quality' achieves this, because // "User Agents MUST not drop frames to achieve the target bitrate and/or framerate." encoderConfig.latencyMode = 'quality'; } const hasOddDimension = encoderConfig.width % 2 === 1 || encoderConfig.height % 2 === 1; if ( hasOddDimension && (this.encodingConfig.codec === 'avc' || this.encodingConfig.codec === 'hevc') ) { // Throw a special error for this case as it gets hit often throw new Error( `The dimensions ${encoderConfig.width}x${encoderConfig.height} are not supported for codec` + ` '${this.encodingConfig.codec}'; both width and height must be even numbers. Make sure to` + ` round your dimensions to the nearest even number.`, ); } const support = await VideoEncoder.isConfigSupported(encoderConfig); if (!support.supported) { throw new Error( `This specific encoder configuration (${encoderConfig.codec}, ${encoderConfig.bitrate} bps,` + ` ${encoderConfig.width}x${encoderConfig.height}, hardware acceleration:` + ` ${encoderConfig.hardwareAcceleration ?? 'no-preference'}) is not supported by this browser.` + ` Consider using another codec or changing your video parameters.`, ); } /** Queue of color chunks waiting for their alpha counterpart. */ const colorChunkQueue: { chunk: EncodedVideoChunk; meta: EncodedVideoChunkMetadata | undefined; }[] = []; /** Each value is the number of encoded alpha chunks at which a null alpha chunk should be added. */ const nullAlphaChunkQueue: number[] = []; let encodedAlphaChunkCount = 0; let alphaEncoderQueue = 0; const addPacket = ( colorChunk: EncodedVideoChunk, alphaChunk: EncodedVideoChunk | null, meta: EncodedVideoChunkMetadata | undefined, ) => { const sideData: EncodedPacketSideData = {}; if (alphaChunk) { const alphaData = new Uint8Array(alphaChunk.byteLength); alphaChunk.copyTo(alphaData); sideData.alpha = alphaData; } let packet = EncodedPacket.fromEncodedChunk(colorChunk, sideData); // See if there's a relevant timing entry to refine the packet's timing data const preciseTimingIndex = binarySearchLessOrEqual( this.preciseTimings, colorChunk.timestamp, x => x.microsecondTimestamp, ); const entry = preciseTimingIndex !== -1 ? this.preciseTimings[preciseTimingIndex] : null; let actualType: PacketType | null = null; if (this.emittedEncoderPackets === 0 && packet.type === 'delta' && meta?.decoderConfig) { // https://github.com/Vanilagy/mediabunny/issues/365 // We expect the first packet to be a key packet. If it's not, let's actually verify that it's // not by getting the actual type. actualType = determineVideoPacketType( this.encodingConfig.codec, meta.decoderConfig, packet.data, ); } // Define the packet if ((entry && entry.microsecondTimestamp === colorChunk.timestamp) || actualType !== null) { packet = packet.clone({ timestamp: entry?.timestampIsValid ? entry.timestamp : undefined, duration: entry?.durationIsValid ? entry.duration : undefined, type: actualType ?? undefined, }); } maybeEnsureIsKeyPacket(this.source._connectedTrack!, packet); this.encodingConfig.onEncodedPacket?.(packet, meta); this.lastMuxerPromise = this.muxer!.addEncodedVideoPacket(this.source._connectedTrack!, packet, meta) .catch((error) => { this.error ??= error; }); this.emittedEncoderPackets++; }; const stack = new Error('Encoding error').stack; this.encoder = new VideoEncoder({ output: (chunk, meta) => { if (!this.alphaEncoder) { // We're done addPacket(chunk, null, meta); return; } const alphaFrame = this.alphaFrameQueue.shift(); assert(alphaFrame !== undefined); if (alphaFrame) { this.alphaEncoder.encode(alphaFrame, { // Crucial: The alpha frame is forced to be a key frame whenever the color frame // also is. Without this, playback can glitch and even crash in some browsers. // This is the reason why the two encoders are wired in series and not in parallel. keyFrame: chunk.type === 'key', }); alphaEncoderQueue++; alphaFrame.close(); colorChunkQueue.push({ chunk, meta }); } else { // There was no alpha component for this frame if (alphaEncoderQueue === 0) { // No pending alpha encodes either, so we're done addPacket(chunk, null, meta); } else { // There are still alpha encodes pending, so we can't add the packet immediately since // we'd end up with out-of-order packets. Instead, let's queue a null alpha chunk to be // added in the future, after the current encoder workload has completed: nullAlphaChunkQueue.push(encodedAlphaChunkCount + alphaEncoderQueue); colorChunkQueue.push({ chunk, meta }); } } }, error: (error) => { error.stack = stack; // Provide a more useful stack trace, the default one sucks this.error ??= error; }, }); this.encoder.configure(encoderConfig); if (this.encodingConfig.alpha === 'keep') { const stack = new Error('Encoding error').stack; // We need to encode alpha as well, which we do with a separate encoder this.alphaEncoder = new VideoEncoder({ // We ignore the alpha chunk's metadata // eslint-disable-next-line @typescript-eslint/no-unused-vars output: (chunk, meta) => { alphaEncoderQueue--; // There has to be a color chunk because the encoders are wired in series const colorChunk = colorChunkQueue.shift(); assert(colorChunk !== undefined); addPacket(colorChunk.chunk, chunk, colorChunk.meta); // See if there are any null alpha chunks queued up encodedAlphaChunkCount++; while ( nullAlphaChunkQueue.length > 0 && nullAlphaChunkQueue[0] === encodedAlphaChunkCount ) { nullAlphaChunkQueue.shift(); const colorChunk = colorChunkQueue.shift(); assert(colorChunk !== undefined); addPacket(colorChunk.chunk, null, colorChunk.meta); } }, error: (error) => { error.stack = stack; // Provide a more useful stack trace this.error ??= error; }, }); this.alphaEncoder.configure(encoderConfig); } } assert(this.source._connectedTrack); this.muxer = this.source._connectedTrack.output._muxer; this.encoderInitialized = true; })(); } async flushAndClose(forceClose: boolean) { if (!forceClose) { this.checkForEncoderError(); } // Final frame rate padding: fill remaining frames up to the last sample's original end timestamp if (!forceClose && this.frameRateLastSample) { const frameRate = this.encodingConfig.transform!.frameRate!; const alignedEnd = floorToDivisor(this.frameRateLastEndTimestamp!, frameRate); await this.padFrameRate(alignedEnd); } this.closed = true; this.frameRateLastSample?.close(); this.frameRateLastSample = null; if (this.customEncoder) { if (!forceClose) { void this.customEncoderCallSerializer.call(() => this.customEncoder!.flush()); } await this.customEncoderCallSerializer.call(() => this.customEncoder!.close()); } else if (this.encoder) { if (!forceClose) { // These are wired in series, therefore they must also be flushed in series await this.encoder.flush(); await this.alphaEncoder?.flush(); } if (this.encoder.state !== 'closed') { this.encoder.close(); } if (this.alphaEncoder && this.alphaEncoder.state !== 'closed') { this.alphaEncoder.close(); } this.alphaFrameQueue.forEach(x => x?.close()); this.splitter?.close(); } if (!forceClose) { this.checkForEncoderError(); } } getQueueSize() { if (this.customEncoder) { return this.customEncoderQueueSize; } else { // Because the color and alpha encoders are wired in series, there's no need to also include the alpha // encoder's queue size here return this.encoder?.encodeQueueSize ?? 0; } } checkForEncoderError() { if (this.error) { throw this.error; } } } let splitterGpuUnavailable = false; /** Utility class for splitting a composite frame into separate color and alpha components. */ export class ColorAlphaSplitter { static forceCpu = true; canvas: OffscreenCanvas | HTMLCanvasElement | null = null; private gl: WebGL2RenderingContext | null = null; private colorProgram: WebGLProgram | null = null; private alphaProgram: WebGLProgram | null = null; private vao: WebGLVertexArrayObject | null = null; private sourceTexture: WebGLTexture | null = null; private alphaResolutionLocation: WebGLUniformLocation | null = null; private worker: Worker | null = null; private pendingRequests = new Map< number, ReturnType> >(); private nextRequestId = 0; constructor(initialWidth: number, initialHeight: number) { const canMakeCanvas = typeof OffscreenCanvas !== 'undefined' // eslint-disable-next-line @typescript-eslint/no-deprecated || (typeof document !== 'undefined' && typeof document.createElement === 'function'); if (!ColorAlphaSplitter.forceCpu && canMakeCanvas && !splitterGpuUnavailable) { // Try the GPU path. If anything goes wrong, we silently fall back to the CPU path. try { if (typeof OffscreenCanvas !== 'undefined') { this.canvas = new OffscreenCanvas(initialWidth, initialHeight); } else { this.canvas = document.createElement('canvas'); this.canvas.width = initialWidth; this.canvas.height = initialHeight; } const gl = this.canvas.getContext('webgl2', { alpha: true, // Needed due to the YUV thing we do for alpha }) as unknown as WebGL2RenderingContext | null; // Casting because of some TypeScript weirdness if (!gl) { throw new Error('Couldn\'t acquire WebGL 2 context.'); } this.gl = gl; this.colorProgram = this.createColorProgram(); this.alphaProgram = this.createAlphaProgram(); this.vao = this.createVAO(); this.sourceTexture = this.createTexture(); this.alphaResolutionLocation = this.gl.getUniformLocation(this.alphaProgram, 'u_resolution')!; this.gl.useProgram(this.colorProgram); this.gl.uniform1i(this.gl.getUniformLocation(this.colorProgram, 'u_sourceTexture'), 0); this.gl.useProgram(this.alphaProgram); this.gl.uniform1i(this.gl.getUniformLocation(this.alphaProgram, 'u_sourceTexture'), 0); } catch (error) { this.gl = null; this.canvas = null; splitterGpuUnavailable = true; Logging._warn('Falling back to CPU for color/alpha splitting.', error); } } } async update(sourceFrame: VideoFrame) { if (this.gl) { return this.updateGpu(sourceFrame); } else { return this.updateCpu(sourceFrame); } } private updateGpu(sourceFrame: VideoFrame) { assert(this.gl); assert(this.canvas); if (sourceFrame.displayWidth !== this.canvas.width || sourceFrame.displayHeight !== this.canvas.height) { this.canvas.width = sourceFrame.displayWidth; this.canvas.height = sourceFrame.displayHeight; } this.gl.activeTexture(this.gl.TEXTURE0); this.gl.bindTexture(this.gl.TEXTURE_2D, this.sourceTexture); this.gl.texImage2D(this.gl.TEXTURE_2D, 0, this.gl.RGBA, this.gl.RGBA, this.gl.UNSIGNED_BYTE, sourceFrame); const colorFrame = this.runColorProgram(sourceFrame); const alphaFrame = this.runAlphaProgram(sourceFrame); sourceFrame.close(); return { colorFrame, alphaFrame }; } private createVertexShader(): WebGLShader { assert(this.gl); return this.createShader(this.gl.VERTEX_SHADER, `#version 300 es in vec2 a_position; in vec2 a_texCoord; out vec2 v_texCoord; void main() { gl_Position = vec4(a_position, 0.0, 1.0); v_texCoord = a_texCoord; } `); } private createColorProgram(): WebGLProgram { assert(this.gl); const vertexShader = this.createVertexShader(); // This shader is simple, simply copy the color information while setting alpha to 1 const fragmentShader = this.createShader(this.gl.FRAGMENT_SHADER, `#version 300 es precision highp float; uniform sampler2D u_sourceTexture; in vec2 v_texCoord; out vec4 fragColor; void main() { vec4 source = texture(u_sourceTexture, v_texCoord); fragColor = vec4(source.rgb, 1.0); } `); const program = this.gl.createProgram(); this.gl.attachShader(program, vertexShader); this.gl.attachShader(program, fragmentShader); this.gl.linkProgram(program); return program; } private createAlphaProgram(): WebGLProgram { assert(this.gl); const vertexShader = this.createVertexShader(); // This shader's more complex. The main reason is that this shader writes data in I420 (yuv420) pixel format // instead of regular RGBA. In other words, we use the shader to write out I420 data into an RGBA canvas, which // we then later read out with JavaScript. The reason being that browsers weirdly encode canvases and mess up // the color spaces, and the only way to have full control over the color space is by outputting YUV data // directly (avoiding the RGB conversion). Doing this conversion in JS is painfully slow, so let's utlize the // GPU since we're already calling it anyway. const fragmentShader = this.createShader(this.gl.FRAGMENT_SHADER, `#version 300 es precision highp float; uniform sampler2D u_sourceTexture; uniform vec2 u_resolution; // The width and height of the canvas in vec2 v_texCoord; out vec4 fragColor; // This function determines the value for a single byte in the YUV stream float getByteValue(float byteOffset) { float width = u_resolution.x; float height = u_resolution.y; float yPlaneSize = width * height; if (byteOffset < yPlaneSize) { // This byte is in the luma plane. Find the corresponding pixel coordinates to sample from float y = floor(byteOffset / width); float x = mod(byteOffset, width); // Add 0.5 to sample the center of the texel vec2 sampleCoord = (vec2(x, y) + 0.5) / u_resolution; // The luma value is the alpha from the source texture return texture(u_sourceTexture, sampleCoord).a; } else { // Write a fixed value for chroma and beyond return 128.0 / 255.0; } } void main() { // Each fragment writes 4 bytes (R, G, B, A) float pixelIndex = floor(gl_FragCoord.y) * u_resolution.x + floor(gl_FragCoord.x); float baseByteOffset = pixelIndex * 4.0; vec4 result; for (int i = 0; i < 4; i++) { float currentByteOffset = baseByteOffset + float(i); result[i] = getByteValue(currentByteOffset); } fragColor = result; } `); const program = this.gl.createProgram(); this.gl.attachShader(program, vertexShader); this.gl.attachShader(program, fragmentShader); this.gl.linkProgram(program); return program; } private createShader(type: number, source: string): WebGLShader { assert(this.gl); const shader = this.gl.createShader(type)!; this.gl.shaderSource(shader, source); this.gl.compileShader(shader); if (!this.gl.getShaderParameter(shader, this.gl.COMPILE_STATUS)) { Logging._error('Shader compile error:', this.gl.getShaderInfoLog(shader)); } return shader; } private createVAO(): WebGLVertexArrayObject { assert(this.gl); assert(this.colorProgram); const vao = this.gl.createVertexArray(); this.gl.bindVertexArray(vao); const vertices = new Float32Array([ -1, -1, 0, 1, 1, -1, 1, 1, -1, 1, 0, 0, 1, 1, 1, 0, ]); const buffer = this.gl.createBuffer(); this.gl.bindBuffer(this.gl.ARRAY_BUFFER, buffer); this.gl.bufferData(this.gl.ARRAY_BUFFER, vertices, this.gl.STATIC_DRAW); const positionLocation = this.gl.getAttribLocation(this.colorProgram, 'a_position'); const texCoordLocation = this.gl.getAttribLocation(this.colorProgram, 'a_texCoord'); this.gl.enableVertexAttribArray(positionLocation); this.gl.vertexAttribPointer(positionLocation, 2, this.gl.FLOAT, false, 16, 0); this.gl.enableVertexAttribArray(texCoordLocation); this.gl.vertexAttribPointer(texCoordLocation, 2, this.gl.FLOAT, false, 16, 8); return vao; } private createTexture(): WebGLTexture { assert(this.gl); const texture = this.gl.createTexture(); this.gl.bindTexture(this.gl.TEXTURE_2D, texture); this.gl.texParameteri(this.gl.TEXTURE_2D, this.gl.TEXTURE_WRAP_S, this.gl.CLAMP_TO_EDGE); this.gl.texParameteri(this.gl.TEXTURE_2D, this.gl.TEXTURE_WRAP_T, this.gl.CLAMP_TO_EDGE); this.gl.texParameteri(this.gl.TEXTURE_2D, this.gl.TEXTURE_MIN_FILTER, this.gl.LINEAR); this.gl.texParameteri(this.gl.TEXTURE_2D, this.gl.TEXTURE_MAG_FILTER, this.gl.LINEAR); return texture; } private runColorProgram(sourceFrame: VideoFrame) { assert(this.gl); assert(this.canvas); this.gl.useProgram(this.colorProgram); this.gl.viewport(0, 0, this.canvas.width, this.canvas.height); this.gl.clear(this.gl.COLOR_BUFFER_BIT); this.gl.bindVertexArray(this.vao); this.gl.drawArrays(this.gl.TRIANGLE_STRIP, 0, 4); return new VideoFrame(this.canvas, { timestamp: sourceFrame.timestamp, duration: sourceFrame.duration ?? undefined, alpha: 'discard', }); } private runAlphaProgram(sourceFrame: VideoFrame) { assert(this.gl); assert(this.canvas); this.gl.useProgram(this.alphaProgram); this.gl.uniform2f(this.alphaResolutionLocation, this.canvas.width, this.canvas.height); this.gl.viewport(0, 0, this.canvas.width, this.canvas.height); this.gl.clear(this.gl.COLOR_BUFFER_BIT); this.gl.bindVertexArray(this.vao); this.gl.drawArrays(this.gl.TRIANGLE_STRIP, 0, 4); const { width, height } = this.canvas; const chromaSamples = Math.ceil(width / 2) * Math.ceil(height / 2); const yuvSize = width * height + chromaSamples * 2; const requiredHeight = Math.ceil(yuvSize / (width * 4)); let yuv = new Uint8Array(4 * width * requiredHeight); this.gl.readPixels(0, 0, width, requiredHeight, this.gl.RGBA, this.gl.UNSIGNED_BYTE, yuv); yuv = yuv.subarray(0, yuvSize); assert(yuv[width * height] === 128); // Where chroma data starts assert(yuv[yuv.length - 1] === 128); // Assert the YUV data has been fully written // Defining this separately because TypeScript doesn't know `transfer` and I can't be bothered to do declaration // merging right now const init = { format: 'I420' as const, codedWidth: width, codedHeight: height, timestamp: sourceFrame.timestamp, duration: sourceFrame.duration ?? undefined, transfer: [yuv.buffer], }; return new VideoFrame(yuv, init); } private updateCpu(sourceFrame: VideoFrame): Promise<{ colorFrame: VideoFrame; alphaFrame: VideoFrame }> { if (!this.worker) { const blob = new Blob( [`(${colorAlphaSplitterWorkerCode.toString()})()`], { type: 'application/javascript' }, ); const url = URL.createObjectURL(blob); this.worker = new Worker(url); URL.revokeObjectURL(url); this.worker.addEventListener('message', (event: MessageEvent) => { const data = event.data; const pending = this.pendingRequests.get(data.id); if (!pending) { return; } this.pendingRequests.delete(data.id); if ('error' in data) { pending.reject(new Error(data.error)); } else { pending.resolve({ colorFrame: data.colorFrame, alphaFrame: data.alphaFrame }); } }); this.worker.addEventListener('error', (event) => { const error = new Error(event.message || 'Color/alpha splitter worker error.'); for (const pending of this.pendingRequests.values()) { pending.reject(error); } this.pendingRequests.clear(); }); } const id = this.nextRequestId++; const pending = promiseWithResolvers<{ colorFrame: VideoFrame; alphaFrame: VideoFrame }>(); this.pendingRequests.set(id, pending); this.worker.postMessage({ id, sourceFrame }, { transfer: [sourceFrame] }); return pending.promise; } close() { this.gl?.getExtension('WEBGL_lose_context')?.loseContext(); this.gl = null; this.canvas = null; this.worker?.terminate(); this.worker = null; const error = new Error('Color/alpha splitter closed.'); for (const pending of this.pendingRequests.values()) { pending.reject(error); } this.pendingRequests.clear(); } } type ColorAlphaSplitterWorkerRequest = { id: number; sourceFrame: VideoFrame; }; type ColorAlphaSplitterWorkerResponse = | { id: number; colorFrame: VideoFrame; alphaFrame: VideoFrame } | { id: number; error: string }; const colorAlphaSplitterWorkerCode = () => { // Reused across frames as long as the size matches, since consecutive frames usually share dimensions. let cpuSourceBuffer: Uint8Array | null = null; // Serialize execution internally so concurrent requests don't race on the shared cpuSourceBuffer. let chain: Promise = Promise.resolve(); self.addEventListener('message', (event: MessageEvent) => { const { id, sourceFrame } = event.data; chain = chain.then(async () => { try { const { colorFrame, alphaFrame } = await split(sourceFrame); self.postMessage({ id, colorFrame, alphaFrame }, { transfer: [colorFrame, alphaFrame] }); } catch (error) { self.postMessage({ id, error: (error as Error).message }); } finally { sourceFrame.close(); } }); }); const split = async (sourceFrame: VideoFrame) => { const format = sourceFrame.format as VideoSamplePixelFormat | null; if (!format) { throw new Error('CPU color/alpha splitting requires a known VideoFrame format.'); } const width = sourceFrame.codedWidth; const height = sourceFrame.codedHeight; const sourceSize = sourceFrame.allocationSize(); if (!cpuSourceBuffer || cpuSourceBuffer.byteLength !== sourceSize) { cpuSourceBuffer = new Uint8Array(sourceSize); } await sourceFrame.copyTo(cpuSourceBuffer); if (format === 'RGBA' || format === 'BGRA') { return splitInterleavedRgba(cpuSourceBuffer, width, height, format, sourceFrame); } else if ( format === 'I420A' || format === 'I420AP10' || format === 'I420AP12' || format === 'I422A' || format === 'I422AP10' || format === 'I422AP12' || format === 'I444A' || format === 'I444AP10' || format === 'I444AP12' ) { return splitPlanarYuvA(cpuSourceBuffer, width, height, format, sourceFrame); } throw new Error(`CPU color/alpha splitting does not support format '${format}'.`); }; const splitInterleavedRgba = ( source: Uint8Array, width: number, height: number, format: 'RGBA' | 'BGRA', sourceFrame: VideoFrame, ) => { const pixelCount = width * height; const chromaW = Math.ceil(width / 2); const chromaH = Math.ceil(height / 2); const alphaSize = pixelCount + chromaW * chromaH * 2; // Encode alpha as I420: Y = source A bytes, UV = 128 const alphaBuffer = new Uint8Array(alphaSize); for (let i = 0, j = 3; i < pixelCount; i++, j += 4) { alphaBuffer[i] = source[j]!; } alphaBuffer.fill(128, pixelCount); // Hand the source buffer straight to VideoFrame as RGBX/BGRX so the A bytes are ignored const colorFrame = new VideoFrame(source, { format: format === 'RGBA' ? 'RGBX' : 'BGRX', codedWidth: width, codedHeight: height, timestamp: sourceFrame.timestamp, duration: sourceFrame.duration ?? undefined, // No transfer! }); const alphaInit = { format: 'I420' as const, codedWidth: width, codedHeight: height, timestamp: sourceFrame.timestamp, duration: sourceFrame.duration ?? undefined, transfer: [alphaBuffer.buffer], }; const alphaFrame = new VideoFrame(alphaBuffer, alphaInit); return { colorFrame, alphaFrame }; }; const splitPlanarYuvA = ( source: Uint8Array, width: number, height: number, format: | 'I420A' | 'I420AP10' | 'I420AP12' | 'I422A' | 'I422AP10' | 'I422AP12' | 'I444A' | 'I444AP10' | 'I444AP12', sourceFrame: VideoFrame, ) => { const is10 = format.includes('P10'); const is12 = format.includes('P12'); const bytesPerSample = (is10 || is12) ? 2 : 1; let chromaW: number; let chromaH: number; if (format.startsWith('I420')) { chromaW = Math.ceil(width / 2); chromaH = Math.ceil(height / 2); } else if (format.startsWith('I422')) { chromaW = Math.ceil(width / 2); chromaH = height; } else { chromaW = width; chromaH = height; } const ySamples = width * height; const uvSamples = chromaW * chromaH; const yBytes = ySamples * bytesPerSample; const uvBytes = uvSamples * bytesPerSample; const aBytes = ySamples * bytesPerSample; const colorBytes = yBytes + uvBytes * 2; const colorFormat = format.replace('A', '') as VideoPixelFormat; const alphaChromaW = Math.ceil(width / 2); const alphaChromaH = Math.ceil(height / 2); const alphaUvSamples = alphaChromaW * alphaChromaH; const alphaUvBytes = alphaUvSamples * bytesPerSample; const alphaSize = aBytes + 2 * alphaUvBytes; const alphaBuffer = new Uint8Array(alphaSize); const aPlaneStart = colorBytes; alphaBuffer.set(source.subarray(aPlaneStart, aPlaneStart + aBytes), 0); // Fill UV planes with the neutral chroma value const uvOffset = aBytes; const neutralChroma = is10 ? 512 : (is12 ? 2048 : 128); if (bytesPerSample === 1) { alphaBuffer.fill(neutralChroma, uvOffset); } else { const uvView = new Uint16Array(alphaBuffer.buffer, uvOffset, 2 * alphaUvSamples); uvView.fill(neutralChroma); } const alphaFormat = (is10 ? 'I420P10' : (is12 ? 'I420P12' : 'I420')) as VideoPixelFormat; // Color frame is simply a prefix of the combined bytes const colorFrame = new VideoFrame(source.subarray(0, colorBytes), { format: colorFormat, codedWidth: width, codedHeight: height, timestamp: sourceFrame.timestamp, duration: sourceFrame.duration ?? undefined, }); const alphaInit = { format: alphaFormat, codedWidth: width, codedHeight: height, timestamp: sourceFrame.timestamp, duration: sourceFrame.duration ?? undefined, transfer: [alphaBuffer.buffer], }; const alphaFrame = new VideoFrame(alphaBuffer, alphaInit); return { colorFrame, alphaFrame }; }; }; /** * This source can be used to add raw, unencoded video samples (frames) to an output video track. These frames will * automatically be encoded and then piped into the output. * @group Media sources * @public */ export class VideoSampleSource extends VideoSource { /** @internal */ private _encoder: VideoEncoderWrapper; /** * Creates a new {@link VideoSampleSource} whose samples are encoded according to the specified * {@link VideoEncodingConfig}. */ constructor(encodingConfig: VideoEncodingConfig) { validateVideoEncodingConfig(encodingConfig); super(encodingConfig.codec); this._encoder = new VideoEncoderWrapper(this, encodingConfig); } /** * Encodes a video sample (frame) and then adds it to the output. * * @returns A Promise that resolves once the output is ready to receive more samples. You should await this Promise * to respect writer and encoder backpressure. */ add(videoSample: VideoSample, encodeOptions?: VideoEncoderEncodeOptions) { if (!(videoSample instanceof VideoSample)) { throw new TypeError('videoSample must be a VideoSample.'); } return this._encoder.add(videoSample, false, encodeOptions); } /** @internal */ override _flushAndClose(forceClose: boolean) { return this._encoder.flushAndClose(forceClose); } } /** * This source can be used to add video frames to the output track from a fixed canvas element. Since canvases are often * used for rendering, this source provides a convenient wrapper around {@link VideoSampleSource}. * @group Media sources * @public */ export class CanvasSource extends VideoSource { /** @internal */ private _encoder: VideoEncoderWrapper; /** @internal */ private _canvas: HTMLCanvasElement | OffscreenCanvas; /** * Creates a new {@link CanvasSource} from a canvas element or `OffscreenCanvas` whose samples are encoded * according to the specified {@link VideoEncodingConfig}. */ constructor(canvas: HTMLCanvasElement | OffscreenCanvas, encodingConfig: VideoEncodingConfig) { if ( !(typeof HTMLCanvasElement !== 'undefined' && canvas instanceof HTMLCanvasElement) && !(typeof OffscreenCanvas !== 'undefined' && canvas instanceof OffscreenCanvas) ) { throw new TypeError('canvas must be an HTMLCanvasElement or OffscreenCanvas.'); } validateVideoEncodingConfig(encodingConfig); super(encodingConfig.codec); this._encoder = new VideoEncoderWrapper(this, encodingConfig); this._canvas = canvas; } /** * Captures the current canvas state as a video sample (frame), encodes it and adds it to the output. * * @param timestamp - The timestamp of the sample, in seconds. * @param duration - The duration of the sample, in seconds. * * @returns A Promise that resolves once the output is ready to receive more samples. You should await this Promise * to respect writer and encoder backpressure. */ add(timestamp: number, duration = 0, encodeOptions?: VideoEncoderEncodeOptions) { if (!Number.isFinite(timestamp) || timestamp < 0) { throw new TypeError('timestamp must be a non-negative number.'); } if (!Number.isFinite(duration) || duration < 0) { throw new TypeError('duration must be a non-negative number.'); } const sample = new VideoSample(this._canvas, { timestamp, duration }); return this._encoder.add(sample, true, encodeOptions); } /** @internal */ override _flushAndClose(forceClose: boolean) { return this._encoder.flushAndClose(forceClose); } } /** * Options for {@link MediaStreamVideoTrackSource}. * @group Media sources * @public */ export type MediaStreamVideoTrackSourceOptions = { /** * The frame rate at which the underlying video track is sampled. Defaults to the frame rate specified in the * track's [`MediaTrackSettings`](https://developer.mozilla.org/en-US/docs/Web/API/MediaTrackSettings). Set to * `null` to only add a frame whenever the underlying track pushes one - this minimizes frame count but can * lead to wildly irregular FPS. */ frameRate?: number | null; /** * Controls the basis (zero point) for video frame timestamps. * * When set to `'synced-zero'`, timestamps will be relative to the first chunk of media from a `MediaStreamTrack` * added to the {@link Output}. * * When set to `'zero'`, timestamps will be relative to the first video frame emitted by this source. * * When set to `'unix'`, timestamps will be relative to the Unix epoch, so clearly associated with a distinct point * in time. Here, pausing via {@link MediaStreamVideoTrackSource.pause} will also create gaps in timestamps. Be sure * to pair this mode with {@link BaseTrackMetadata.isRelativeToUnixEpoch}. * * Defaults to `'synced-zero'`. */ timestampBase?: 'synced-zero' | 'zero' | 'unix'; }; /** * Video source that encodes the frames of a * [`MediaStreamVideoTrack`](https://developer.mozilla.org/en-US/docs/Web/API/MediaStreamTrack) and pipes them into the * output. This is useful for capturing live or real-time data such as webcams or screen captures. Frames will * automatically start being captured once the connected {@link Output} is started, and will keep being captured until * the {@link Output} is finalized or this source is closed. * @group Media sources * @public */ export class MediaStreamVideoTrackSource extends VideoSource { /** @internal */ private _options: MediaStreamVideoTrackSourceOptions; /** @internal */ private _encoder: VideoEncoderWrapper; /** @internal */ private _abortController: AbortController | null = null; /** @internal */ private _track: MediaStreamVideoTrack; /** @internal */ private _workerTrackId: number | null = null; /** @internal */ private _workerListener: ((event: MessageEvent) => void) | null = null; /** @internal */ private _promiseWithResolvers = promiseWithResolvers(); /** @internal */ private _errorPromiseAccessed = false; /** @internal */ private _paused = false; /** @internal */ private _lastVideoFrame: VideoFrame | null = null; /** @internal */ private _timerHandle: UnthrottledTimerHandle | null = null; /** @internal */ private _videoElement: HTMLVideoElement | null = null; /** A promise that rejects upon any error within this source. This promise never resolves. */ get errorPromise() { this._errorPromiseAccessed = true; return this._promiseWithResolvers.promise; } /** Whether this source is currently paused as a result of calling `.pause()`. */ get paused() { return this._paused; } /** * Creates a new {@link MediaStreamVideoTrackSource} from a * [`MediaStreamVideoTrack`](https://developer.mozilla.org/en-US/docs/Web/API/MediaStreamTrack), which will pull * video samples from the stream in real time and encode them according to {@link VideoEncodingConfig}. */ constructor( track: MediaStreamVideoTrack, encodingConfig: VideoEncodingConfig, options: MediaStreamVideoTrackSourceOptions = {}, ) { if (!(track instanceof MediaStreamTrack) || track.kind !== 'video') { throw new TypeError('track must be a video MediaStreamTrack.'); } validateVideoEncodingConfig(encodingConfig); if (typeof options !== 'object' || !options) { throw new TypeError('options must be an object.'); } if (options.frameRate != null && (typeof options.frameRate !== 'number' || options.frameRate <= 0)) { throw new TypeError('options.frameRate, when provided, must be either a positive number or null.'); } if ( options.timestampBase !== undefined && options.timestampBase !== 'synced-zero' && options.timestampBase !== 'zero' && options.timestampBase !== 'unix' ) { throw new TypeError( 'options.timestampBase, when provided, must be one of \'synced-zero\', \'zero\', or \'unix\'.', ); } encodingConfig = { ...encodingConfig, latencyMode: 'realtime', }; super(encodingConfig.codec); this._options = options; this._encoder = new VideoEncoderWrapper(this, encodingConfig); this._track = track; } /** @internal */ override async _start() { if (!this._errorPromiseAccessed) { Logging._warn( 'Make sure not to ignore the `errorPromise` field on MediaStreamVideoTrackSource, so that any internal' + ' errors get bubbled up properly.', ); } const frameRate = this._options.frameRate !== undefined ? this._options.frameRate : (this._track.getSettings().frameRate ?? null); this._abortController = new AbortController(); let firstVideoFrameTimestamp: number | null = null; let lastFrameTime: number | null = null; let frameCount = 0; let errored = false; let lastSampleTimestamp: number | null = null; let timestampOffset = 0; const tick = () => { assert(frameRate !== null); if (!this._lastVideoFrame) { return; } assert(lastFrameTime !== null); assert(firstVideoFrameTimestamp !== null); const now = performance.now(); // Add as many frames as warranted by the elapsed time. // > instead of >= intentionally because tick() is called before the _lastVideoFrame is changed while (now - lastFrameTime > 1000 / frameRate) { lastFrameTime += 1000 / frameRate; const timestamp = firstVideoFrameTimestamp + frameCount / frameRate; const frame = new VideoFrame(this._videoElement ?? this._lastVideoFrame, { timestamp: 1e6 * timestamp, duration: 1e6 / frameRate, }); addVideoFrame(frame, now); } }; if (frameRate !== null) { this._timerHandle = setIntervalUnthrottled(tick, 4); // Run it at 250 Hz } const onVideoFrame = (videoFrame: VideoFrame) => { if (frameRate === null) { addVideoFrame(videoFrame); } else { const now = performance.now(); if (!this._lastVideoFrame) { addVideoFrame(videoFrame.clone(), now); lastFrameTime = now; this._lastVideoFrame = videoFrame; } else { tick(); this._lastVideoFrame?.close(); this._lastVideoFrame = videoFrame; } } }; const addVideoFrame = (videoFrame: VideoFrame, now = performance.now()) => { if (errored) { videoFrame.close(); return; } frameCount++; const currentTimestamp = videoFrame.timestamp / 1e6; if (this._paused) { const frameSeen = firstVideoFrameTimestamp !== null; if (frameSeen) { if (lastSampleTimestamp !== null && this._options.timestampBase !== 'unix') { // In addition to dropping this frame, let's also keep track of the time we have lost due to the // pause. Doing it like this instead of simply keeping track of the paused time is better since // it retains the frame rate of the underlying source. const timeDelta = currentTimestamp - lastSampleTimestamp; timestampOffset -= timeDelta; } lastSampleTimestamp = currentTimestamp; } videoFrame.close(); return; } if (firstVideoFrameTimestamp === null) { firstVideoFrameTimestamp = currentTimestamp; let target: number; const timestampBase = this._options.timestampBase ?? 'synced-zero'; if (timestampBase === 'unix') { target = Date.now() / 1000; } else if (timestampBase === 'zero') { target = 0; } else { const output = this._connectedTrack!.output; if (output._firstMediaStreamTimestamp === null) { output._firstMediaStreamTimestamp = now / 1000; target = 0; } else { target = now / 1000 - output._firstMediaStreamTimestamp; } } timestampOffset = target - firstVideoFrameTimestamp; } lastSampleTimestamp = currentTimestamp; if (this._encoder.getQueueSize() >= 8) { // Drop frames if the encoder is overloaded videoFrame.close(); return; } const sample = new VideoSample(videoFrame, { timestamp: currentTimestamp + timestampOffset, }); void this._encoder.add(sample, true) .catch((error) => { errored = true; this._abortController?.abort(); this._promiseWithResolvers.reject(error); if (this._workerTrackId !== null) { // Tell the worker to stop the track sendMessageToMediaStreamTrackProcessorWorker({ type: 'stopTrack', trackId: this._workerTrackId, }); } }); }; if (typeof MediaStreamTrackProcessor !== 'undefined') { // We can do it here directly, perfect const processor = new MediaStreamTrackProcessor({ track: this._track }); const consumer = new WritableStream({ write: onVideoFrame }); processor.readable.pipeTo(consumer, { signal: this._abortController.signal, }).catch((error) => { // Handle AbortError silently if (error instanceof DOMException && error.name === 'AbortError') return; this._promiseWithResolvers.reject(error); }); } else { // It might still be supported in a worker, so let's check that const supportedInWorker = await mediaStreamTrackProcessorIsSupportedInWorker(); if (supportedInWorker) { this._workerTrackId = nextMediaStreamTrackProcessorWorkerId++; sendMessageToMediaStreamTrackProcessorWorker({ type: 'videoTrack', trackId: this._workerTrackId, track: this._track, }); this._workerListener = (event: MessageEvent) => { const message = event.data as MediaStreamTrackProcessorWorkerMessage; if (message.type === 'videoFrame' && message.trackId === this._workerTrackId) { onVideoFrame(message.videoFrame); } else if (message.type === 'error' && message.trackId === this._workerTrackId) { this._promiseWithResolvers.reject(message.error); } }; mediaStreamTrackProcessorWorker!.addEventListener('message', this._workerListener); } else if (frameRate !== null) { // No MediaStreamTrackProcessor support at all (e.g. Firefox), but we have a frame rate, so we can // manually sample from a hidden