Wszystkie wpisy
Guides2026/08/21

Captions vs. Subtitles: What's the Difference?

Captions are for viewers who can't hear the audio and include non-speech sounds; subtitles are for viewers who can't understand the language and translate dialogue only. A clear comparison, with the accessibility rules that decide which you need.

The words get used interchangeably, but captions and subtitles are not the same thing, and picking the wrong one can put you on the wrong side of an accessibility law.

The short version: captions assume the viewer cannot hear the audio, so they include non-speech sound — music, a door slamming, who is speaking. Subtitles assume the viewer can hear but does not understand the language, so they carry the spoken dialogue only. Everything else follows from that one distinction.

Captions vs. subtitles at a glance

CaptionsSubtitles
Assumes the viewerCannot hear the audioCan hear, doesn't understand the language
CoversDialogue + non-speech sound (music, SFX, speaker IDs)Dialogue only
Primary audienceDeaf and hard-of-hearing viewersViewers watching in a second language
Same language as audio?Usually yesUsually no (translated)
Legally required?Often — see belowRarely; a localisation choice
File formatsSRT, VTT, SCC, TTMLSRT, VTT, ASS/SSA

Both are timed text laid over video, and both ship in the same subtitle file formats — an .srt or .vtt can hold either. The difference is in the content of the cues, not the container. (For the container side of the story, see subtitle and transcript file formats explained.)

What captions include that subtitles don't

Because a captioned viewer can't hear anything, captions have to describe the soundtrack, not just the words:

  • Non-speech sounds in brackets: [phone ringing], [tense music], [laughter].
  • Speaker identification, so it's clear who is talking when the face is off-screen.
  • Tone and delivery where it changes the meaning — [sarcastic], [whispering].

Subtitles drop all of that. A subtitle track for a film assumes you can hear the explosion; it just tells you what the characters are saying.

Open vs. closed captions

A second distinction sits inside "captions":

  • Closed captions (CC) are a separate track the viewer can switch on or off. YouTube captions, the CC button on a streaming player, and a sidecar .srt file are all closed captions.
  • Open captions are burned into the video pixels permanently. They always show, on any player, with no button — common on social video where autoplay is muted and there's no caption UI.

Closed is the default for accessibility because it gives the viewer control. Open captions win when you can't trust the platform to render a caption track — a TikTok or an Instagram Reel, for instance.

SDH: the hybrid in the middle

You'll also see SDH — "subtitles for the deaf and hard of hearing." SDH is subtitle-style text (often in the subtitle workflow and formats) that also carries the non-speech information captions provide. It exists because Blu-ray and many streaming pipelines support subtitle tracks more readily than broadcast-style captions, so SDH delivers caption-grade information through the subtitle channel. If you're captioning for accessibility but your platform only takes subtitles, SDH is the answer.

Which do I need?

Work backwards from your viewer:

  • Publishing video to a general audience (YouTube, a course, a company site)? You need captions — same language as the audio, with non-speech cues. This is also what accessibility law expects.
  • Making your video watchable in another language? You need subtitles — the dialogue, translated.
  • Both is common: same-language captions and one or more translated subtitle tracks. Streaming services do exactly this.
  • Posting to muted-autoplay social feeds? Use open captions so the text survives regardless of the player.

A note on the law

In the United States, video accessibility is not optional for most organisations. The Americans with Disabilities Act (ADA) has been applied to websites and their video, Section 508 covers federal agencies and their contractors, and the FCC requires captioning for much broadcast and re-published online programming. The common thread is captions — accurate, same-language, including non-speech sound — not translated subtitles. If your reason for adding timed text is compliance, you almost certainly need captions. This is general information, not legal advice; check the rules that apply to your organisation.

How to make either one

Both start from the same place: an accurate, timestamped transcript of the audio.

  1. Transcribe the audio or video. Upload the file and let it transcribe — ScribeToAny detects the spoken language automatically (about 98 are supported) and returns timestamped segments. For captions with multiple people on screen, turn on speaker labels.
  2. Review against the audio. Automatic transcription is a first draft. Clicking a segment replays that moment, which is the fastest way to fix a misheard name or term before it ships on screen.
  3. Add non-speech cues if you're captioning. This is the manual step that turns a subtitle track into a caption track — bracket the sounds a viewer who can't hear would miss.
  4. Export the right file. SRT or VTT for a closed track most platforms accept; retime individual cues in the SRT editor if a line lingers or clips.
  5. For another language, translate the finished transcript. On the Max plan, ScribeToAny translates segment by segment into any of 56 languages while keeping the cue timings identical, so a translated SRT still lines up with the picture.

If you want the text burned in for social, generate the track first with the subtitle generator, then use it as open captions.


Captions are for access; subtitles are for language. Get the transcript right once and you can produce both from it. Start with the video-to-text tool and add the non-speech cues by hand where accessibility demands them.