Can ChatGPT Transcribe Audio? We Uploaded a Real Recording to Find Out
We uploaded a one-minute, two-speaker MP3 to ChatGPT and asked for a transcript. What happened, the three ways ChatGPT can transcribe, and the limits.
Short answer: not from an uploaded file in a normal chat, at least in our test. On 4 October 2026 we uploaded a one-minute MP3 to ChatGPT's free web app and asked for a word-for-word transcript. Both the default mode and Think mode recognised the file and its length, then declined. ChatGPT can turn speech into text in three other ways:
- Dictation types what you say into the message box.
- Record mode transcribes live audio. It's for paid plans, in the macOS app only.
- OpenAI's API transcribes files, but you need to write code.
Here's the test, what each option actually does, and when a dedicated transcription tool is the better fit.
The test: one MP3, one prompt
We used the same recording as in our transcript examples guide. It's a 64-second interview between two people: a podcast host and a baker. We made it for our product tutorial from a written script, with two text-to-speech voices. Because we have the script, any transcript of it can be scored word for word.
We uploaded podcast-bakery-interview.mp3 (about 1 MB) to a new chat and sent
this prompt:
Transcribe this audio word for word. Label each speaker (Speaker 1,
Speaker 2) and give a timestamp at the start of each speaker turn.
Default mode (the page reported the model as gpt-5-6) answered after
about 8 seconds. It said it could see the file and that it was about a minute
long, but that it didn't currently have an audio transcription capability in
that chat session. It offered to format a transcript if we supplied one.
Think mode (gpt-5-6-t-mini) worked for 12 seconds. It measured the file
precisely as about 1 minute 3.6 seconds; the real length is 63.6 seconds. Then
it said it couldn't hear or decode the audio in that environment, and it
wouldn't produce a transcript "without risking making up text". Again it
offered a template with Speaker 1 and [00:00] placeholders.
Two things are worth noting:
- It didn't invent a transcript. That's the right behaviour. A confident fake transcript would have been worse than a refusal.
- It could read the file's length but not its words. That suggests the chat could inspect the file's data, but had no speech recognition to apply to it.
What we didn't test: this was the free plan, on the web app, on one day. ChatGPT's features change often and differ by plan. A paid plan or a different app may behave differently, so try your own file before you rely on either answer.
The ways ChatGPT can turn speech into text
ChatGPT isn't one product here. OpenAI offers several speech-to-text routes, and they do different jobs. From OpenAI's own help pages and API documentation:
| Option | What it transcribes | Who can use it | Speaker labels | Timestamps |
|---|---|---|---|---|
| File upload in a chat | Nothing, in our free-plan test | Everyone | n/a | n/a |
| Dictation (microphone icon) | What you say, typed into the message box for you to edit | ChatGPT users; the help article lists no plan limits | No | No |
| Record mode | A live recording of a meeting or voice note, then notes. Up to 4 hours per session | Plus, Pro, Business, Enterprise and Edu, in the macOS desktop app only | Yes: generic labels such as "Speaker 1", which you can rename | Not described in the help article |
OpenAI API (gpt-transcribe, gpt-4o-transcribe, whisper-1 and others) | An audio or video file you send from your own code | Developers with an API account (pay per use) | With gpt-4o-transcribe-diarize | With whisper-1 (word or segment level) |
A few details from those docs matter in practice:
- Record mode records; it doesn't import. The help article describes pressing Record and transcribing as you talk. It doesn't describe feeding it a file you already have. OpenAI also says it "works best in English today".
- The API caps files at 25 MB. It accepts mp3, mp4, mpeg, mpga, m4a, wav and webm. Longer recordings have to be compressed or split.
- In the API, timestamps and speaker labels come from different models.
Per the docs, timestamps are only available with
whisper-1, and speaker labels need the separate diarization model. A subtitle file with both takes more than one call.
When ChatGPT is the right tool, and when it isn't
Use ChatGPT's Record mode if you're on a paid plan, you use a Mac, and the audio is happening now: a meeting, a brainstorm, a voice note. You get a transcript and a summary in the same place.
Use dictation if you just want to talk instead of type.
Use something else if you have a recording already: an interview, a lecture, a podcast episode or a video file. You'll also want something else if you need any of these:
- timestamps on every line,
- an
.srtor.vttsubtitle file, - a language other than English that you need to rely on,
- a transcript you can check line by line against the audio before quoting it.
That's what dedicated transcription tools are built for. We ran the same 64-second file through ScribeToAny's engine. It returned the whole transcript with speaker labels and timestamps. 14 of 14 speaker turns were labelled correctly, and 4 of 189 words differed from the script. Two of those were real errors and two were "eighty" written as "80". The full scoring, mistakes included, is in our transcript examples.
The workflow that uses both
ChatGPT is very good at what comes after the transcript: summarising it, pulling out action items, writing chapter titles, drafting show notes. So split the job:
- Transcribe the recording with a tool built for it. Try the audio to text or video to text tools. Turn on speaker detection for interviews and meetings.
- Check the transcript against the audio, especially names, numbers and anything you'll quote. (See how accurate AI transcription is.)
- Paste it into ChatGPT and ask for the summary, action items or chapters. In ScribeToAny, the AI summary & chat button builds a ready-to-paste prompt with your transcript in it. It opens Claude by default, but the prompt works in ChatGPT too. ScribeToAny itself doesn't send your transcript to any AI service.
You get the accuracy and timing of a transcription engine, plus the writing help of a chat assistant, and you can check every step.
So, can ChatGPT transcribe audio? It can transcribe you, live, through dictation and Record mode, and developers can use OpenAI's API. But in our test, dropping an existing recording into a chat didn't produce a transcript. For a file you already have, transcribe it first, then bring the text to ChatGPT. Start with the interview transcription tool if there's more than one voice.