English

Video & subtitles · Subtitle Toolkit

Captions vs subtitles: why accessibility needs more than a transcript

· Why it matters

subtitles accessibility timecodes text-processing

A speech cue joined by speaker and sound indicators to form a more complete caption track
Original ToolAcre vector illustration

Subtitles translate speech; captions make a video usable without sound. This post explains what accessible captions include, why the distinction matters for deaf and hard-of-hearing viewers and how clean, well-timed files support it.

The video has subtitles, so why is it not accessible? — the gap between translation and access

A video can contain an accurate written version of every spoken sentence and still be unusable without sound. A conventional subtitle track may assume that the viewer can hear who is speaking, when music starts or why a room suddenly reacts. Accessibility requires the timed text to carry the information that the soundtrack otherwise supplies, not merely a transcript divided into cues.

That distinction changes the review question. Instead of asking only whether the words match, a university media team has to ask whether a viewer relying on the track can follow the scene, identify an off-screen voice and understand a meaningful non-speech sound. File conversion cannot supply those editorial decisions, but it can preserve and check the timed text once they have been made.

Subtitles: speech for people who can hear — the assumption baked into most subtitle files

Subtitles commonly concentrate on speech, especially when their primary job is to translate dialogue for an audience that can still hear tone, music and sound effects. The soundtrack continues to identify a familiar voice and signals whether a crash, alarm or laugh matters. A viewer using subtitles as a language aid receives those clues through sound, so the text does not have to repeat all of them.

This is a useful description of an editorial convention, not a rule enforced by SRT or WebVTT. Both formats carry timed text and can hold either subtitles or captions; the extension does not reveal whether the content is accessible. Renaming or converting a file therefore changes its representation, not the amount of information the writer included.

Captions: sound cues, speaker identification and music — what a caption track adds and why

A caption track adds information needed when the soundtrack is unavailable: speaker identification where the speaker is not obvious, descriptions of consequential sounds and indications of music when it contributes meaning. Those additions should be concise enough to read with the dialogue and specific enough to explain the event rather than filling every quiet moment with labels.

The Subtitle Toolkit deliberately does not generate any of this language. Its transform package states that transcription and translation are out of scope, and the same boundary applies to caption authorship. Removing formatting tags can simplify existing cue text, but it cannot decide that a door slam identifies an arrival or that a speaker label is needed.

Timing and readability — why late, overlapping or overly long cues are an accessibility failure, not just an annoyance

Complete text delivered at the wrong moment is still an access failure. A late cue withholds the line while the action moves on; an early cue can reveal a response before it occurs. Overlapping cues may compete for the same screen space, while a zero-duration cue does not display at all. ToolAcre validation names negative durations, zero durations, cues out of order and overlaps by cue number so those structural faults can be reviewed directly.

Readability remains a human judgement. The validator does not enforce line length, reading speed or a style guide, and the clean-up pass does not rewrite long prose. Use the report to remove impossible timings, then watch the complete programme with sound unavailable and check whether each cue can be read in context.

Formats and platforms — how SRT and WebVTT carry captions and where each is expected

SRT and WebVTT can both carry caption text, but delivery systems expect different file shapes. SRT conventionally numbers cues and writes a comma before milliseconds. WebVTT begins with a WEBVTT header, writes a full stop and may carry cue settings. ToolAcre parses either into the same integer-millisecond cue model and writes the selected target format from that model.

A browser track needs WebVTT rather than a renamed SRT file. Other delivery systems may request SRT. Conversion changes the header, timestamp punctuation and numbering policy while preserving cue text and timing; it does not certify that a particular platform accepts the editorial style or that the track includes every accessibility cue.

Worked example: upgrading a subtitle file to a caption file — adding speaker labels and sound descriptions, then checking timing

Start with a subtitle cue such as “I left it on the desk.” If the speaker is off screen, add a concise speaker identification supported by the scene. If a fire alarm interrupts the next line and drives the action, give that sound its own useful description. Repeat this editorial pass through the programme before changing formats, because no parser can infer missing acoustic context from the remaining dialogue.

Then run the authored file through the toolkit. Remove only markup you have decided is disposable, remembering that the cleaner strips angle-bracket tags and brace codes but deliberately preserves character entities. Validate the timings, remove genuinely empty cues, and check the opening, middle and ending against the video after any shift or scale.

What this does not cover — audio description, sign language and transcript pages, which are separate access needs

Captions are one access channel, not a complete accessibility programme. Audio description conveys visual information through sound; sign-language presentation communicates through a signed language; and a transcript page supports searching or reading outside playback. None is created by converting an SRT file, and none should be presented as interchangeable with a caption track.

The toolkit also does not assess the meaning or completeness of caption descriptions. It can report a cue that overlaps its predecessor, but not an important sound that an editor forgot to describe. Keep an editorial review and an audience-appropriate accessibility review alongside the mechanical file checks.

Takeaway: accessible captions are timed, complete and clean — how the Subtitle Toolkit helps with the timing, format and clean-up parts

Accessible captions are authored for use without sound, timed to the event, readable in context and delivered in the format the destination expects. Subtitle Toolkit helps with the mechanical part: tolerant parsing, strict output, conversion, timing changes, formatting removal and cue-level validation. Those operations protect good caption writing; they do not replace it.

Open the converter with a copy of the completed caption file, read every reported issue and export the required SRT or WebVTT version. Watch representative cues with the sound muted before delivery. The useful result is not simply a valid download, but a track that still communicates when audio contributes nothing.