All posts
Guides2026/08/21

Closed Captions vs. Subtitles: What's the Difference?

Closed captions add speaker IDs and sounds for viewers who can't hear; subtitles carry only the dialogue. A side-by-side example, SDH, and which you need.

People use the words interchangeably. Even YouTube's player button says "Subtitles/closed captions". But closed captions and subtitles are not the same thing, and picking the wrong one can put you on the wrong side of an accessibility law.

The short version: captions assume the viewer cannot hear the audio, so they include non-speech sound — music, a door slamming, who is speaking. Subtitles assume the viewer can hear but does not understand the language, so they carry the spoken dialogue only. Everything else follows from that one distinction.

The same moment as subtitles, closed captions and SDH

Definitions are easier to grasp with an example, so here is one moment shown three ways. It's from the two-person interview we use in our product tutorial. The host introduces Maya, who runs a bakery, and she answers. The English text and timings are real transcription output. The Spanish lines come from the tutorial's translated track. The one sound effect is added to show where it would go.

Subtitles (Spanish), for a viewer who can hear but doesn't speak English:

3
00:00:05,320 --> 00:00:06,320
esquina de Mill Street.

4
00:00:07,270 --> 00:00:08,470
Gracias por invitarme.

Closed captions (English), for a viewer who can't hear:

3
00:00:05,320 --> 00:00:06,320
HOST: …corner of Mill Street.

4
00:00:07,270 --> 00:00:08,470
[shop bell rings]
MAYA: Thanks for having me.

SDH (English), the same caption content delivered as a subtitle track: the text is the same as the closed captions. The difference is technical, not in the words. SDH travels in the subtitle channel of a player or disc (see SDH below).

Look at what the captions add: who is speaking (the camera may be on the host when Maya answers) and the sound a hearing viewer would notice. The subtitles drop both, because a viewer who can hear doesn't need them. Our transcript already had speaker labels, from the speaker detection step. The sound effect had to be written by a person, because transcription only captures words. See more ways to lay out the same output in our real transcript examples.

Closed captions vs. subtitles at a glance

CaptionsSubtitles
Assumes the viewerCannot hear the audioCan hear, doesn't understand the language
CoversDialogue + non-speech sound (music, SFX, speaker IDs)Dialogue only
Primary audienceDeaf and hard-of-hearing viewersViewers watching in a second language
Same language as audio?Usually yesUsually no (translated)
Legally required?Often — see belowRarely; a localisation choice
File formatsSRT, VTT, SCC, TTMLSRT, VTT, ASS/SSA

Both are timed text laid over video, and both ship in the same subtitle file formats — an .srt or .vtt can hold either. The difference is in the content of the cues, not the container. (For the container side of the story, see subtitle and transcript file formats explained.)

What captions include that subtitles don't

Because a captioned viewer can't hear anything, captions have to describe the soundtrack, not just the words:

  • Non-speech sounds in brackets: [phone ringing], [tense music], [laughter].
  • Speaker identification, so it's clear who is talking when the face is off-screen.
  • Tone and delivery where it changes the meaning — [sarcastic], [whispering].

Subtitles drop all of that. A subtitle track for a film assumes you can hear the explosion; it just tells you what the characters are saying.

Open vs. closed captions

A second distinction sits inside "captions":

  • Closed captions (CC) are a separate track the viewer can switch on or off. YouTube captions, the CC button on a streaming player, and a sidecar .srt file are all closed captions.
  • Open captions are burned into the video pixels permanently. They always show, on any player, with no button — common on social video where autoplay is muted and there's no caption UI.

Closed is the default for accessibility because it gives the viewer control. Open captions win when you can't trust the platform to render a caption track — a TikTok or an Instagram Reel, for instance.

SDH: the hybrid in the middle

You'll also see SDH — "subtitles for the deaf and hard of hearing." SDH is subtitle-style text (often in the subtitle workflow and formats) that also carries the non-speech information captions provide. It exists because Blu-ray and many streaming pipelines support subtitle tracks more readily than broadcast-style captions, so SDH delivers caption-grade information through the subtitle channel. If you're captioning for accessibility but your platform only takes subtitles, SDH is the answer.

Which do I need?

Work backwards from your viewer:

  • Publishing video to a general audience (YouTube, a course, a company site)? You need captions — same language as the audio, with non-speech cues. This is also what accessibility law expects.
  • Making your video watchable in another language? You need subtitles — the dialogue, translated.
  • Both is common: same-language captions and one or more translated subtitle tracks. Streaming services do exactly this.
  • Posting to muted-autoplay social feeds? Use open captions so the text survives regardless of the player.

A note on the law

In the United States, video accessibility is not optional for most organisations. The Americans with Disabilities Act (ADA) has been applied to websites and their video, Section 508 covers federal agencies and their contractors, and the FCC requires captioning for much broadcast and re-published online programming. The common thread is captions — accurate, same-language, including non-speech sound — not translated subtitles. If your reason for adding timed text is compliance, you almost certainly need captions.

Two standards are worth knowing by name:

  • WCAG 2.x, the web accessibility guidelines that most accessibility rules point to, requires captions for prerecorded video with audio (success criterion 1.2.2, Level A).
  • The FCC's caption quality standards for TV say captions must be accurate, synchronous, complete and properly placed. They don't apply to every web video. Still, they're a good checklist for any caption track: the right words, on screen with the speech, from start to finish, and not covering anything important.

This is general information, not legal advice; check the rules that apply to your organisation.

How to make either one

Both start from the same place: an accurate, timestamped transcript of the audio.

  1. Transcribe the audio or video. Upload the file and let it transcribe — ScribeToAny detects the spoken language automatically (about 98 are supported) and returns timestamped segments. For captions with multiple people on screen, turn on speaker labels.
  2. Review against the audio. Automatic transcription is a first draft. Clicking a segment replays that moment, which is the fastest way to fix a misheard name or term before it ships on screen.
  3. Add non-speech cues if you're captioning. This is the manual step that turns a subtitle track into a caption track — bracket the sounds a viewer who can't hear would miss.
  4. Export the right file. SRT or VTT for a closed track most platforms accept; retime individual cues in the SRT editor if a line lingers or clips.
  5. For another language, translate the finished transcript. The Translate action puts it into any of 134+ languages for free, on every plan, and keeps the cue timings identical, so a translated SRT still lines up with the picture.

If you want the text burned in for social, generate the track first with the subtitle generator, then use it as open captions.


Captions are for access; subtitles are for language. Get the transcript right once and you can produce both from it. Start with the video-to-text tool and add the non-speech cues by hand where accessibility demands them.