English

Video & subtitles · Subtitle Toolkit

Why stray tags in a subtitle file can break a platform caption upload

· Why it matters

subtitles text-processing webvtt

One caption file accepted by three platform parsers with different amounts of clutter surviving into the output
Original ToolAcre vector illustration

Platforms parse caption uploads strictly and often silently. This post explains the common reasons a subtitle file is rejected or displays oddly, why formatting tags and codes are the usual culprits and how a clean, standard file avoids the problem.

The upload was accepted and the captions show literal tags — the quiet failure of an unclean file

The worst outcome is not rejection. A rejected upload tells you something is wrong while you are still looking at the form. The common outcome is acceptance followed by captions that display their own markup to viewers, because the upload check and the rendering path are different pieces of software asking different questions.

The upload check usually asks whether the file parses into cues at all. Rendering asks what to do with the contents of each cue, and a parser that ignores an unfamiliar tag draws it as text. Both steps did what they were designed to do.

What platforms actually accept — plain SRT and WebVTT with predictable structure, and little tolerance for extras

What platforms accept is narrower than what tools emit. In practice that means plain SRT or WebVTT with a predictable shape: blocks separated by blank lines, a timecode line, text lines, and little else. Tolerance for extras is low and, more importantly, undocumented, so the safe assumption is that anything beyond that shape is a risk rather than a feature.

This is not conservatism for its own sake. A platform that accepted arbitrary markup would have to decide how to render it consistently across web, mobile and television clients, which is a much larger commitment than accepting a caption file.

The usual culprits — HTML-like tags, ASS-style codes, BOMs, mixed line endings and blank cues

The recurring culprits are few. Angle-bracket tags for italics, speaker voices and cue classes survive conversion from formats that support them. Brace-delimited override codes arrive from SubStation formats and mean nothing outside them. A byte order mark at the start of the file attaches itself to the first index number. Mixed line endings from cross-platform editing break block splitting. Cues whose text became empty after markup was removed remain as timestamped blanks.

None of these are exotic. They are the normal output of a conversion chain where each step was individually reasonable, which is why they appear in files that look fine in an editor.

Why one platform tolerates what another rejects — differing parsers behind similar upload forms

Different platforms tolerate different subsets, and that is what makes the fault confusing to diagnose. The same file can upload cleanly and render correctly on one service, upload and render markup on a second, and be refused by a third, with no error message that names the actual cause. The file did not change between attempts; three parsers disagreed.

Treating one platform as the reference is therefore the wrong instinct. A file that works on the most permissive service tells you nothing about the others, and the strictest parser is the one that defines whether a file is portable.

Worked example: one file, three uploads — how the same clutter shows up differently and disappears after cleaning

Take one export carrying a top-of-frame override code on sixty cues, emphasis tags on forty, and a byte order mark. Uploaded to three services it might be accepted everywhere. On the first the tags are honoured and the override code is drawn literally. On the second both appear as text. On the third the mark costs the first cue, which simply never appears, and nobody notices because the opening line is usually a title.

Cleaning once removes all three classes at the source. The angle-bracket tags and brace codes are substituted out, the whitespace left behind is collapsed, cues emptied by the removals are dropped rather than emitted as timestamped blanks, and the file is rewritten with cues renumbered contiguously from one. The same output then goes to all three services.

What this does not cover — platform-specific style guides, character limits per line and burned-in captions

This covers structural portability, not editorial conformance. Platform style guides on line length, maximum characters, positioning of speaker labels and handling of sound effects are separate requirements that a clean file does not satisfy automatically. A structurally perfect caption file can still breach a style guide.

Burned-in captions are a different mechanism entirely. Text rendered into the video frames is not a caption file, cannot be toggled off, and is unaffected by anything described here.

Takeaway: clean once, upload anywhere — how the Subtitle Toolkit's clean and convert features produce a platform-friendly file

Clean once and upload the same file everywhere. The alternative, maintaining a separate export per platform, multiplies the number of files that can drift out of sync with the master and does not remove the underlying clutter from any of them.

Run the clean and convert together, then read the reported issues before uploading rather than after. The issues name cues by number, which is the difference between knowing a file has a problem and knowing which line to look at.