English

Video & subtitles · Subtitle Toolkit

WebVTT explained: from WebSRT to the W3C caption format for HTML5 video

· Background

subtitles webvtt file-formats browser-processing

A plain cue file gaining a WebVTT header, settings and browser track connection
Original ToolAcre vector illustration

WebVTT was created so browsers could show captions natively. This post covers its origins as WebSRT, how it differs from SRT, what the specification adds and where it is used today.

Why the browser needed its own caption format — the arrival of native video and the absence of a standard track format

Native video gave a web page a media element, but timed text still needed an interoperable file representation that browsers could fetch and parse. An informal SRT file was not enough for a browser feature that needed defined syntax, cue behaviour and hooks for presentation. WebVTT supplies that web-facing contract while keeping the recognisable cue-and-timestamp shape.

That browser context explains both its familiarity and its extra structure. A minimal VTT looks close to SRT, but it is not SRT with another extension. The signature, timestamp punctuation and cue grammar are part of deciding whether the browser has received WebVTT at all.

WebSRT and the rename to WebVTT — the early WHATWG work and the move to the W3C

The format was discussed under the earlier WebSRT name before becoming WebVTT, or Web Video Text Tracks. That lineage communicates the design path: begin with a widely understood subtitle block shape, then define the additional behaviour needed by web media. This article avoids dates, named editors and a detailed standards-body chronology because those facts are not verified by the repository materials used for the implementation claims.

For a front-end developer, the durable lesson is that WebVTT is specified input to browser media features rather than a player-specific convention. The implementation should target its current syntax and test in the consuming player instead of depending on historical labels.

What the specification adds over SRT — the WEBVTT header, cue identifiers, cue settings and styling hooks

WebVTT adds a required WEBVTT signature and permits cue identifiers that need not be numeric. Cue settings can follow the end time to express alignment and placement, while NOTE, STYLE and REGION blocks provide metadata and presentation structures beyond a plain SRT cue list. Browser styling hooks can then address displayed cues without burning text into the video.

ToolAcre implements a focused subset. It preserves settings found after the end timestamp, but drops cue identifiers and skips NOTE, STYLE and REGION blocks on load. That is sufficient for ordinary conversion and deliberately not a lossless editor for the full WebVTT model.

The dot, the header and UTF-8 — the small differences that stop an SRT file being a valid VTT file

Three small details separate a minimal VTT from SRT. The first line must carry the WEBVTT signature, timestamps use a full stop before milliseconds and the text must be UTF-8. ToolAcre detects VTT by testing for the signature while tolerating an optional byte-order mark and leading whitespace; text without that signature defaults to SRT.

The serializer writes the signature and dot punctuation rather than preserving whatever the source happened to use. File input is read as UTF-8. A legacy-encoded SRT can therefore show replacement characters before conversion, and changing its container syntax cannot recover bytes already decoded incorrectly.

Where WebVTT is used — the HTML5 track element, HLS captions and many video platforms

WebVTT is the sidecar format associated with the HTML video track element and appears in web delivery systems that consume timed text. Some streaming and platform workflows also use WebVTT, but acceptance details belong to the destination and should be checked there rather than inferred from the extension. A valid VTT file can still be served with the wrong media type or blocked by cross-origin delivery rules.

The toolkit produces a file representation; it does not configure an HTML page, server headers or a streaming manifest. Treat conversion as one stage in delivery and test the actual player with the same URL, headers and origin relationship the audience will use.

What this does not cover — the full styling and region model, which most caption files never use

The full styling and region model is outside this converter. Complex cue selectors, region layout and presentation policies deserve a WebVTT-aware authoring tool and browser testing. The parser skips STYLE and REGION blocks, so sending such a master through this path removes them from either output rather than preserving them invisibly.

The tool also carries a cue setting into SRT output even though SRT has no defined mechanism for acting on it. That documented limit is a reason to retain the original VTT master whenever placement is important, not a reason to assume every destination will ignore the extra text in the same way.

Takeaway: a real standard for the web — how the Subtitle Toolkit produces WebVTT that follows it

WebVTT turns a familiar timed-text pattern into defined web input. The signature identifies it, the timestamp grammar is canonical and the cue model can carry more than plain sequential text. Subtitle Toolkit produces the minimal canonical file from SRT by parsing times into whole milliseconds, then writing a WEBVTT header, dot separators and cues without identifier lines.

Compare the source and output cue counts and read all parse issues before attaching the result to a player. If the source relies on comments, regions, style blocks or identifiers, choose a workflow that preserves those features; a focused converter should state that boundary rather than imply complete round-trip fidelity.