Video & subtitles · Subtitle Toolkit
Beyond SRT and VTT: where ASS, TTML, DFXP and SCC subtitles are used
· Background
subtitles file-formats srt webvtt
SRT and WebVTT dominate the web, but broadcast, anime fandom and streaming delivery use other formats. This post maps the landscape, explains what each format is for and why a focused toolkit supports only two of them.
A folder full of extensions your player does not recognise — the format diversity in professional caption work
A localisation handoff can contain files that all put words on screen and share almost no internal structure. Some are plain cue lists, some are XML documents, some carry detailed styling and some encode a legacy broadcast caption stream. Treating every unfamiliar extension as “another SRT” risks deleting the feature that made the sender choose it.
The first triage question is not which converter opens the file, but what delivery system it belongs to. Preserve the original, identify the requested destination and use a parser that understands that format. Subtitle Toolkit explicitly reads and writes only SRT and WebVTT, which is a boundary rather than a failed promise of universality.
ASS and SSA — styled, positioned subtitles from the anime fansub tradition
SSA and its successor ASS are script-style subtitle formats with sections for metadata, reusable styles and timed dialogue events. They can express positioning and appearance beyond ordinary SRT. They are strongly associated with richly styled community subtitle workflows, including anime fansubbing, but those associations do not turn every ASS file into a fansub or define its licensing.
Brace-delimited override commands from SubStation formats sometimes survive a crude conversion as visible text. ToolAcre’s clean-up can remove blocks such as a positioning override from cue text, but it does not parse an ASS script, translate its styles or preserve karaoke and drawing behaviour. Convert with an ASS-aware tool before bringing a plain-text derivative here.
TTML and DFXP — XML-based timed text for broadcast and streaming delivery
TTML represents timed text as XML, allowing structure, namespaces, timing and presentation information to travel together. DFXP is a name encountered in the same timed-text family and in delivery workflows built around it. These documents are not blank-line cue blocks, so an SRT parser cannot safely recover them by searching for arrows.
Broadcast and streaming systems may choose a TTML-family profile with requirements narrower than the broad format. This article does not claim one profile, standard version or platform mapping for every .ttml or .dfxp file. The delivery specification is the authority, and conversion should preserve a master until the destination accepts and displays the result.
SCC and CEA-608 — the broadcast closed-caption legacy still required by some distributors
SCC files represent a broadcast-caption lineage associated with CEA-608 data rather than a human-friendly list of SRT blocks. Their contents and constraints reflect caption commands and transmission models that predate web sidecar text. A distributor asking for SCC is asking for that delivery representation, not for an SRT file with a different suffix.
The repository tool does not read SCC and does not encode CEA-608. It also does not verify broadcaster conformance. Keep the source caption project and use software designed for the requested broadcast workflow; export to SRT or WebVTT only when a separate destination actually calls for one of those files.
SBV, LRC and others — platform-specific and niche formats
SBV is encountered in platform caption workflows, while LRC is commonly used for timed lyrics. Other extensions encode still more assumptions about timing, styling or a particular application. Their superficial similarity—a time plus a line of text—is not enough to make them interchangeable, because punctuation, ordering and metadata determine how a parser reads the file.
This article does not provide an exhaustive registry or assert that one service always accepts one format. Platform support changes and professional deliveries are contractual. Read the current destination documentation and retain the original file, especially when a conversion necessarily discards layout or styling.
Why a toolkit chooses SRT and WebVTT only — the web's exchange formats, and the honesty of documenting that limit
A focused browser toolkit chooses SRT and WebVTT because its transform model is cue text, start and end milliseconds, plus an optional setting string. That model maps directly to the common block structure of those two formats. It does not contain an XML tree, reusable style sheet, broadcast command stream or embedded video track, so pretending to support those would silently flatten information.
The limitation is stated in the converter record: ASS and SSA, TTML, SAMI and captions embedded in a container are not handled. Even within WebVTT, NOTE, STYLE and REGION blocks are skipped. Honest scope lets a user stop before damage rather than discovering after export that a specialised feature vanished.
Takeaway: know the map, then pick the tool — how the Subtitle Toolkit covers the SRT and WebVTT part of the landscape
Know the map, preserve the master and pick the tool that understands the requested destination. Subtitle Toolkit covers the SRT and WebVTT part: tolerant cue parsing, canonical serialization, timing changes, text clean-up and structural validation. It is not a universal subtitle interchange engine and does not turn an unsupported format into a safe input by renaming it.
If a specialist application has already produced a reviewed SRT or WebVTT derivative, load that derivative here and inspect the cue count and issues before conversion. Keep the ASS, TTML, SCC or other master beside it so any lost styling or delivery metadata can be traced instead of guessed back into existence.