Transcribe an audio recording
Choose a file from a voice recorder or recording app and Whisper AI transcribes it inside your browser. There is no need to convert the format, and there is no time limit however long the recording is. The audio never reaches a server.
You can normally leave this on Auto. Change it only if the result clearly has the wrong number of speakers
How to transcribe a voice recorder file
- Connect the voice recorder to your computer over USB and copy the recording. For recordings made with a smartphone app, save the file to cloud storage from the share menu. For iPhone Voice Memos, see the steps on the M4A transcription page.
- If two or more people are speaking, turn on speaker diarization. If you know the language, choose it instead of auto-detect.
- Choose the recording. Finalized lines appear on screen as it goes, and when it finishes you can export the result as a TXT or SRT file.
File formats by recording device
Many voice recorders save MP3 or WAV, and some save WMA. Smartphone recording apps save M4A, MP3, AMR or other formats. All of these load as they are, so there is no need to convert them to MP3 first.
What affects the result most is how the recording was made. Three things matter: the speaker being close to the microphone, little background noise such as air conditioning, and voices not overlapping.
There is no time limit, even for long recordings. Finalized lines appear while processing, so you can check the content before it finishes. If you close the page partway, the text finalized up to that point stays in your history.
If something goes wrong
- If quiet or distant voices are missed, switch to Small or a larger model.
- In long recordings with silence, sentences nobody said may appear. Delete those lines on the results screen.
Frequently asked questions
Which recording formats are supported?
MP3, WAV, M4A, AAC, FLAC, OGG, Opus, AIFF, WMA, AMR and WebM are supported. If you pick a video file, only the audio track is extracted and processed.
How many hours of recording can be transcribed?
There is no time limit. Files can be up to 2 GB, or 500 MB on phones. Phones run in a reduced-memory mode to protect the device, so a computer is recommended for long recordings.
Can it separate multiple speakers in a recording?
Yes. With speaker diarization on, lines are split by speaker (up to eight). For meeting recordings, the Meeting transcription page, which has diarization turned on from the start, is the easiest to use.
Is my recording sent anywhere?
No. Transcription runs in your browser on your device, and neither the recording nor the resulting text is passed to an external server.
Getting the audio ready, and after transcribing
Help us improve the service
Sending us the circumstances of a run — successful or failed — helps us track down problems (a failure rate needs both). Knowing which devices and settings fail most often lets us revise the recommended defaults and fix the bugs behind them. If you don't mind, leaving this on is a real help.
What is sent
Your device specs (GPU model, memory, core count), browser and OS, the model and language you chose, how far the run got, the error details, the file's format with a rough size and length, and how long it took.
What is never sent
Audio, transcripts, file names, and IP addresses. The data goes only to this site's own server, never to a third-party service.