Convert VTT to Plain Text

WebVTT wraps its words in a header, optional cue identifiers, NOTE blocks and positioning settings — none of which you want when what you are after is the transcript. This drops all of it and hands back the text. It runs in your browser, so a caption file you were sent in confidence stays on your machine.

Runs entirely in your browser — your file is never uploaded.

Nothing loaded yet.

Need subtitles generated from the video itself? Upload it and get a transcript in ~98 languages. 50% off your first month with FIRST20.

Try it free

Why ScribeToAny

  • Upload MP3, MP4, M4A, MOV, AAC, WAV, OGG, OPUS, MPEG, WMA, WMV — audio is extracted from video automatically.
  • About 98 languages with automatic detection, powered by a Whisper-class engine on dedicated GPUs.
  • Eight export formats: TXT, SRT, VTT, TSV, CSV, JSON, PDF and DOCX — plus batch ZIP export.
  • Speaker recognition labels every segment. Click any sentence to replay the matching audio.

Frequently asked questions

What happens to NOTE and STYLE blocks?

They are skipped. The WebVTT spec forbids the cue arrow inside them, which is exactly how the parser tells them apart from real cues — nothing is guessed at.

Can I keep the timestamps?

Choose DOCX or PDF instead of TXT; both put a timestamp in front of each line. Plain text drops them by design, because the usual reason for wanting .txt is to paste the words somewhere else.

My captions came from YouTube and every line repeats. Why?

YouTube auto-captions roll each line into the next cue so the text scrolls on screen, so a straight export shows every line twice. Undoing that reliably needs the original audio — transcribing the video with ScribeToAny gives clean, non-overlapping segments instead.

Related tools

Start transcribing free

Free plan: 2 files per day, 100 minutes per month. No credit card required.

Start transcribing free