Video & subtitles · Subtitle Toolkit
Why small subtitle timing errors are so noticeable to viewers
· Why it matters
subtitles timecodes accessibility
Viewers spot out-of-sync captions quickly, even when the error is a fraction of a second. This post explains, qualitatively, why the eye and ear are so sensitive to it, how sync errors arise in a workflow and how to measure and fix them.
The comments say the captions feel wrong but nobody can say why — the perceptual problem with sync errors
The reports are consistently vague. Viewers say the captions feel off, or distracting, or that they stopped using them, and they rarely say by how much or in which direction. That vagueness is a property of the fault rather than of the audience: the mismatch is perceived long before it becomes measurable to someone who is not looking for it.
It also means the complaint arrives without the information needed to fix it. The useful first step is not to adjust anything but to turn the impression into two numbers, a direction and a magnitude.
Why we notice — reading a line before or after hearing it breaks the link between speech and text
A caption is read while the same words are being heard, so the two streams are being matched continuously rather than checked occasionally. When text arrives before the speech it is read first, which removes the moment it was meant to support; when it arrives after, it confirms something already understood. Either way the reader has done work that the timing made pointless.
This is why the tolerance is asymmetric in practice. Slightly late captions tend to read as lag, while slightly early ones can spoil a line before it is delivered, which is more damaging in comedy and drama than the equivalent delay would be.
Where sync errors enter a workflow — trimming the intro after captioning, re-encoding at a different frame rate, or exporting with a different start time
Most sync errors are introduced after captioning rather than during it. Trimming an intro, a slate or a countdown moves every subsequent frame while leaving the caption timestamps where they were. Re-encoding at a different frame rate rescales the material against timestamps that were absolute. Exporting with a non-zero start timecode offsets the programme relative to a file that begins at zero.
In each case the caption file is unchanged and still internally consistent, which is why validation passes and the file looks correct in isolation. The fault is in the relationship between two artefacts, not inside either one.
Offset versus drift — telling the two apart in a few minutes of viewing
Distinguishing offset from drift takes a few minutes and settles which correction applies. Watch a line near the start and a line near the end. If both are wrong by about the same amount the error is a constant offset and a single shift fixes it. If the ending is worse than the opening the error is growing, which is drift, and shifting will make the ending worse while appearing to fix the opening.
The tool records this as a common mistake precisely because it is the intuitive wrong move: lining up the first line of a drifting file is satisfying and actively harmful, and both operations have to be undone before scaling.
Measuring the error — using the player's clock and a clear first line rather than guessing
Measure against the player clock rather than by impression. Choose a line with a sharp onset, a first word after silence rather than a phrase inside continuous speech, and note the time the caption appears and the time the word is heard. Repeat near the end of the programme. Two measurements on clear onsets are worth more than a dozen judgements on ambiguous ones.
Record the direction explicitly, because sign errors are common. Captions appearing before the dialogue need a positive offset to delay them; a negative offset moves them earlier still and doubles the fault.
Worked example: a trimmed intro that shifted everything by 1.2 seconds — diagnosing and correcting it
Take a programme where a one and a fifth second intro was trimmed after captioning. Every caption now appears one and a fifth seconds after its line. The measurement at the start says twelve hundred milliseconds late; the measurement near the end says the same. Equal errors mean an offset, so one shift corrects the file.
Enter it as 1200, not as 1.2. The field takes milliseconds with no seconds or timecode input, and typing 1 shifts the file by a single millisecond: the operation succeeds, the file genuinely changes, and the sync looks exactly as wrong as before. If the shift is negative and any cue would pass zero, the clamp warning matters, because those cues are pinned at the start permanently and shifting forward afterwards does not restore them.
What this does not cover — captions that are wrong in content or split badly across lines
This covers when captions appear, not what they say or how they are broken across lines. A caption that is perfectly synchronised and badly segmented is still hard to read, and a mistranscribed word is wrong at any offset.
It also does not cover captions that drift because the video itself was variable rate or was assembled from sources at different rates. That is a jump rather than a smooth ramp, and neither a single shift nor a single scale corrects it.
Takeaway: measure, then shift — how the Subtitle Toolkit's retime feature fixes a constant offset in one pass
Measure, then shift once. The sequence that wastes time is adjusting by feel, rewatching, adjusting again, because each pass changes the thing being measured and clamped cues accumulate at zero along the way.
Two timed observations give a direction and a magnitude, and a magnitude entered in milliseconds is one operation. Confirm on the same two lines afterwards rather than on the opening alone, since the opening is where a wrong correction looks most convincing.