English

Video & subtitles · Subtitle Toolkit

How a browser parses an SRT file: blocks, indexes, timecodes and text

· How it works

subtitles srt file-formats

Two SRT cue blocks separated by a blank line, each with an index line, a timecode line and text lines
Original ToolAcre vector illustration

SRT looks trivial until you meet real files. This post walks through how a parser splits blocks, reads indexes and timecodes, handles multi-line text and recovers from the malformed blocks that real-world files contain.

The file 'looks fine' but half the cues are missing — how a lenient-looking format hides strict expectations

SRT has no specification body, no MIME registration and no validator that ships with players. What exists instead is a shape that most software agrees on: a number, a timecode line, one or more lines of text, then a blank line. Because the shape is conventional rather than specified, two files can both look correct in a text editor while only one of them loads, and the failure is usually silent. A player that cannot read a cue tends to skip it rather than report it, so a file with a broken block plays with a gap instead of an error.

A parser therefore has two jobs that pull in opposite directions. It has to accept the variations that real files contain, because files are produced by transcription services, hand editing and format converters that each make different assumptions. It also has to refuse readings that would put a cue at the wrong time, because a silently wrong timestamp is worse than a reported failure.

Splitting into blocks — blank lines as separators and the trouble with stray whitespace and CRLF

Splitting happens on blank lines, not on the index numbers. The parser normalises line endings first, replacing both the CRLF pair and a lone CR with a single newline, because a file authored on Windows and edited on Unix can contain both. It then splits on a run of two or more newlines, trims each resulting block and discards the empty ones. That ordering matters: splitting before normalising would leave a stray carriage return at the end of a timecode line, and the timecode would then fail to match.

A byte order mark is stripped before any of this. A UTF-8 BOM at the start of a file is three bytes that a naive parser sees as part of the first index number, which is enough to make the first cue unreadable while every later cue parses. Trailing spaces on an otherwise blank separator line are handled by the trim, so a file whose blank lines contain a space still splits correctly.

The index line — why numbers are often wrong, duplicated or missing and why parsers should not trust them

The index number is read and then ignored. Real files number cues from zero, restart numbering after a merge, duplicate a number after a manual edit, or omit the line entirely when a converter wrote the file. Trusting those numbers means inheriting every one of those faults, so the parser assigns its own sequential number instead, counting the cues it has successfully built so far.

That choice also explains why the parser never requires the index line to be present. It locates the timecode line by searching the block for the first line containing the arrow, rather than assuming the timecode is the second line. A block with no index line parses normally, and a block with two stray lines before the timecode still parses, because position is not what identifies the timecode.

The timecode line — HH:MM:SS,mmm --> HH:MM:SS,mmm, tolerated variations and the ones that break players

The timecode line is matched against a single regular expression, and the tolerances in it are deliberate. Hours are optional, because WebVTT permits a two-field reading and converters emit it. Either a comma or a full stop is accepted as the millisecond separator regardless of which format the file claims to be, because mixed separators are common enough that rejecting them would fail more good files than bad ones. Fractional digits are padded on the right, so a cue ending in a single digit is read as hundreds of milliseconds rather than units.

Two readings are refused. A minutes or seconds field above fifty-nine is rejected rather than carried, because ninety seconds is not a clock reading and usually indicates a corrupt or misconverted file; normalising it silently would move the cue. A line whose start or end fails to parse produces a recorded issue naming the offending text and the expected shape, and the block is skipped rather than guessed at.

Text lines — multi-line cues, formatting tags and where a block really ends

Everything after the timecode line is the cue text, joined back together with newlines. There is no line limit and no attempt to reflow, so a three-line cue survives as three lines. This is the reason the blank line is load-bearing: it is the only thing that tells the parser the text has ended, which is why a cue whose own text contains a blank line will be read as two blocks and the second half will be reported as having no timecode.

Cue settings are separated from the end timestamp by a run of two or more spaces. WebVTT allows positioning directives such as alignment and line placement to follow the end time on the same line, so the parser splits them off before the timestamp is parsed and keeps them alongside the cue. A single space is not a separator, which keeps a merely untidy timecode line from losing its end time.

Worked example: parsing a five-cue file with two deliberate faults — what a robust parser recovers and what it flags

Take a five-block file in which block three has had its timecode line damaged to read 00:01:75,000 --> 00:01:78,000, and block four has lost its timecode line entirely during a copy and paste. The parser reads blocks one and two normally and numbers them one and two. Block three matches the shape of a timecode but carries a seconds field of seventy-five, so it is refused and recorded as a bad timestamp naming the line it could not read.

Block four contains no arrow at all, so it is recorded as having no timestamp, quoting the first forty characters of the block so the line can be found in the original file. Block five parses and becomes cue three, not cue five, because numbering counts successful cues. The result is three usable cues and two specific, located complaints, rather than an exception on the first fault and no information about the second.

What this does not cover — ASS/SSA styling, positioning codes and non-subtitle text dumped into SRT

This describes SRT and the parts of WebVTT that share its cue shape. It does not cover ASS and SSA, which carry a script header, style definitions and per-event style references, and which cannot be read by splitting on blank lines. Karaoke timing, drawing commands and the inline override tags those formats use are outside what a cue-and-timecode parser models.

It also does not repair text. A transcript pasted into a file with no timecodes produces a list of blocks that have no timestamp, which is reported accurately but cannot be turned into subtitles without timing information that is not present. Encoding faults are a separate concern: a file decoded with the wrong character set parses into perfectly valid cues whose text is wrong, and no amount of structural checking will detect that.

Takeaway: parse leniently, write strictly — how the Subtitle Toolkit reads messy SRT and writes back a clean one

The working rule is to parse leniently and write strictly. On the way in, accept optional hours, either separator, missing index lines, mixed line endings and a leading byte order mark, and record every fault as a located issue instead of throwing on the first one, so a file can be fixed in a single pass. On the way out, emit one canonical shape.

That is what the Subtitle Toolkit does when it converts. Cues are renumbered from one and kept contiguous, timestamps are re-emitted with a comma for SRT and a full stop for WebVTT, and the file that comes back is the shape players expect regardless of how irregular the input was. Paste a file that a player rejected into the converter and read the reported issues first; they name the cue and quote the line, which is usually enough to find the fault in the original.