Unlimited transcription in your browser with Whisper AI
Choose an audio file and the Whisper model runs inside your browser to transcribe it. There is no need to set up Python, ffmpeg or GPU drivers. The audio never reaches a server, and there is no account or payment.
You can normally leave this on Auto. Change it only if the result clearly has the wrong number of speakers
How to use it
- Choose a model. The default is picked automatically from what your device can handle. On a PC with WebGPU, Large v3 Turbo offers a good balance of accuracy and speed.
- Choose the language. Auto-detect works, but specifying the actual language reduces recognition errors.
- Choose an audio or video file. Transcription starts as soon as the file is loaded, and when it finishes you can export the result as a TXT or SRT file.
About Whisper running in the browser
Whisper is a speech recognition model released by OpenAI. It is normally used by setting up a Python environment and installing ffmpeg and GPU drivers. This tool runs the same Whisper models converted to a format that works in the browser. A model is downloaded the first time, and after that the copy stored in your browser is used.
Six models are available: Tiny, Base, Small, Medium, Large v3 Turbo and Large v3. On a PC with WebGPU they run on the GPU; on phones and devices without WebGPU they run on the CPU. Larger models are more accurate but need a bigger download and more memory. Large v3 Turbo keeps the Large v3 encoder and uses a smaller decoder, so it runs fast with accuracy close to Large v3.
Whisper sometimes writes out sentences nobody said during silence or music. Repeated sentences are removed automatically, but a sentence that appears only once remains, so delete that line on the results screen.
If something goes wrong
- If processing stops partway, your device may be running out of memory. Switch to the next smaller model or split the audio into shorter parts.
- If proper nouns or technical terms are often missed, switch to Small or a larger model.
Frequently asked questions
Do I need Python or a GPU?
No. A recent browser such as Chrome or Edge is enough. On a PC with WebGPU it runs fast on the GPU, and on phones or devices without a usable GPU it runs on the CPU (it takes longer).
Which model should I choose?
The default is chosen automatically from what your device can handle. On a PC with WebGPU, Large v3 Turbo offers a good balance of accuracy and speed; on a phone, choose Tiny or Base.
Are the results the same as the Python version of Whisper?
The model is the same Whisper, but it is made lighter to run in the browser, so fine details of the output may not match the Python version.
Does it work offline?
Yes. Once the model has been downloaded, you can transcribe without a network connection.
Getting the audio ready, and after transcribing
Help us improve the service
Sending us the circumstances of a run — successful or failed — helps us track down problems (a failure rate needs both). Knowing which devices and settings fail most often lets us revise the recommended defaults and fix the bugs behind them. If you don't mind, leaving this on is a real help.
What is sent
Your device specs (GPU model, memory, core count), browser and OS, the model and language you chose, how far the run got, the error details, the file's format with a rough size and length, and how long it took.
What is never sent
Audio, transcripts, file names, and IP addresses. The data goes only to this site's own server, never to a third-party service.