What Is an SRT File? The Format, an Example, and the Mistakes That Break It
An SRT file is a plain-text subtitle file of numbered, timed cues. The exact format, a real example, and what we found feeding 16 broken SRTs to ffmpeg.
An SRT file (.srt, short for SubRip Subtitle) is a plain-text subtitle
file. It holds a list of numbered cues. Each cue has a start time, an end
time and the words to show on screen between them. Almost every video player,
editor and video platform can read one. That's why it's the default answer to
"what subtitle file do I need?"
The name comes from SubRip, a free Windows program that pulled subtitles off DVDs and saved them as text. The program is long gone, but its output format became the lingua franca of subtitles. SRT never got a formal specification, though. Every tool reads it its own way, and that's where the trouble starts. Below is the format, a real example, and what happened when we fed 16 slightly broken SRT files to ffmpeg and to our own converter.
What an SRT file looks like
Here are the first three cues of a real file: the subtitles ScribeToAny made for a 64-second interview we use in our tutorial.
1
00:00:00,000 --> 00:00:01,680
Welcome back to Small Shop Stories.
2
00:00:02,160 --> 00:00:05,320
Today I'm talking to Maya, who runs a tiny bakery on the
3
00:00:05,320 --> 00:00:06,320
corner of Mill Street.
Each cue is four parts, always in this order:
- A number. The cue's position in the file, counting from 1.
- A timing line. The start and end time, separated by
-->(space, two hyphens, a greater-than sign, space). The format ishours:minutes:seconds,milliseconds, so00:00:05,320is 5.32 seconds. - The text. One or two lines is the norm. Most style guides cap a line at about 42 characters.
- A blank line. It marks the end of the cue.
Two details catch people out. The milliseconds come after a comma, not a dot; the dot is WebVTT's habit. And the file should be saved as UTF-8, or accented letters turn into garbage in some players.
What SRT can and can't do
SRT is deliberately simple. That's its strength and its limit.
- Basic styling works, mostly. Many players honour
<i>,<b>,<u>and<font color="…">. Support varies, so don't rely on colour. - Positioning isn't part of the format. Some files put
{\an8}at the start of a line to move it to the top of the screen. That's borrowed from the ASS format. Players that don't understand it show it as literal text. - There's no language field. The convention is to put the language in the
filename, such as
interview.en.srtandinterview.es.srt. Media servers and players use that to label the tracks. - Web browsers can't play it. The HTML5
<track>element only accepts WebVTT. We checked in Chrome 152: an SRT file in a<track>failed to load and showed no cues, and the same cues saved as.vttloaded fine. For web video, convert with SRT to VTT.
For how SRT compares with VTT, TXT, JSON, ASS and the rest, see subtitle and transcript file formats.
We fed 16 broken SRT files to ffmpeg, and to our own converter
With no official spec, "valid SRT" means "whatever the tool you're using accepts". To see what that means in practice, we took the three cues above and broke them in 16 different ways, the kinds of mistakes that come from hand editing, old Windows tools and copy-paste. Then we ran all 17 files (the valid one plus the 16 broken ones) through two readers.
- ffmpeg 8.1.1. A huge number of video tools use it under the hood. We converted each file to WebVTT.
- ScribeToAny's subtitle converter, the parser behind our free SRT tools.
Here's what happened. "Silent" means the conversion reported success and the output looked fine at a glance. At most, a warning was printed in the log.
| What's wrong with the file | ffmpeg 8.1.1 | ScribeToAny (before our fix) |
|---|---|---|
| Nothing (valid file) | âś… | âś… |
Dot instead of comma (00:00:01.680) | âś… | âś… |
| Byte-order mark at the start (Windows Notepad) | âś… | âś… |
| Windows line endings (CRLF) | âś… | âś… |
| Cue numbers out of order or skipped (5, 2, 9) | âś… | âś… |
One-digit hours (0:00:01,680) | âś… | âś… |
Extra coordinates after the times (X1:100 X2:600 …) | ✅ Ignored | ✅ Ignored |
| No cue numbers at all | ⚠️ Rejected when ffmpeg guesses the format; read fine when told it's SRT | ✅ |
No spaces around the arrow (-->) | ⚠️ Same as above | ✅ |
Short arrow (->) | ❌ Nothing read | ❌ Nothing read |
Two-digit milliseconds (,68) | ❌ Silent: read as 68 ms instead of 680, so times moved by up to 0.6 s | ✅ Read as 680 ms |
| End time before start time | ⚠️ Silent: cue shrunk to zero length, so it never shows | ⚠️ Kept as is; flagged by the editor's quality check |
| Overlapping cues | Kept as is | Kept as is |
<i> / <b> tags | âś… Kept | âś… Kept |
<font color> tag | Colour dropped | âś… Kept |
{\an8} position tag | Tag removed | ❌ Shown as text |
| No blank lines between cues | ✅ Three cues | ❌ Merged into one cue |
| Saved as Latin-1 / Windows-1252 ("café") | ❌ Silent: that cue dropped | ❌ Garbled ("caf�") |
What the test shows
The dangerous failures are the silent ones. A file that won't open gets
noticed. A file that opens with one cue missing, or with its timings shifted,
ships. The two-digit-millisecond case is the nastiest: ffmpeg read ,68 as 68
milliseconds rather than 680, so that time moved 0.6 seconds, and nothing
warned about it. Always write three digits: ,680 for 680 ms, ,068 for
68 ms.
Being "lenient" isn't the same as being right. ffmpeg rescued the file with no blank lines, but dropped the Latin-1 cue with only a log warning. Our converter read the Latin-1 cue, but showed the accents as garbage. Neither stopped with an error.
Three of these broke our own tool. Files with no blank lines between cues
collapsed into a single cue. {\an8} tags showed up as text in the converted
file. Legacy-encoded files came out garbled. We fixed all three the same day.
Our parser now finds cues by their timing lines instead of by blank lines. It
removes ASS position tags. And it detects the file's encoding (UTF-8, UTF-16 or
Windows-1252) instead of assuming UTF-8. We added each case to our regression
checks so they stay fixed.
What we didn't test: this was one ffmpeg build and one converter, with small synthetic files. Players such as VLC, editing apps and platforms like YouTube each have their own parser, and may behave differently on the same file.
How to check and fix an SRT file
If a platform rejects your SRT, or the subtitles drift, open the file in a plain-text editor and go down this list:
- Encoding. Re-save it as UTF-8. This one fix clears up most garbled accents.
- Timing lines. Each one should look exactly like
00:00:05,320 --> 00:00:06,320: two digits for hours, minutes and seconds, three for milliseconds, a comma, and-->with spaces. - Blank lines. There must be one empty line between cues and none inside a cue.
- Numbers. Number the cues 1, 2, 3… in order. Most readers don't care, but some strict validators do.
- Times that make sense. Each end time must come after its start time. Ideally, each cue ends before the next one starts.
- Tags. Remove
{\an8}-style tags unless you know the destination player supports them. Keep<i>for italics.
Faster: load the file into the SRT editor. It runs in your browser, so the file isn't uploaded anywhere. It reads every variation above except the short arrow, and its quality check flags cues that end before they start, overlap, or go by too fast to read. Then save a clean copy. (Too-fast cues are their own topic: see subtitle reading speed.)
How to make an SRT file
- From a video or audio file: upload it to the
subtitle generator.
ScribeToAny transcribes it in about 98 languages and exports a timed
.srt. - From a script you already have: text to SRT turns plain text into timed cues.
- From another subtitle format: convert with VTT to SRT or ASS to SRT.
- By hand: any plain-text editor works. Follow the four-part pattern above
and save as UTF-8 with the
.srtextension.
An SRT file is the simplest thing in video: numbered cues, a timestamp pair and a line of text. The trouble is that without a standard, every tool fills the gaps differently, often without saying so. Keep your files strictly to the pattern (three-digit milliseconds, blank lines, UTF-8) and they'll read the same everywhere. When one doesn't, check it in the SRT editor before it ships.