How to convert speech to text
- Choose an audio or video file, or record from your microphone.
- Pick the spoken language, or let it detect the language for you.
- Wait while the transcript appears. The first time, the model is downloaded.
- Edit the text if needed, then copy it or download it as TXT, or as SRT or VTT subtitles.
Speech recognition by OpenAI Whisper (MIT License), running with Transformers.js.
Frequently asked questions
Is my audio uploaded?
No. The speech recognition model (OpenAI Whisper) is downloaded from this site once and runs in your browser, so your recording never leaves your device. That also means it works with confidential meetings and interviews.
Which model should I choose?
Fast (42 MB) works well for clear speech in English and other widely spoken languages. Accurate (77 MB) makes fewer mistakes, especially with accents, background noise and other languages, but takes about three times as long. Both are downloaded once and then cached.
How long does it take?
It depends on your device. On a recent laptop, the Fast model transcribes a minute of speech in roughly 10 to 30 seconds. Phones are slower. Keep the tab open until it finishes.
Which languages are supported?
Whisper understands about 100 languages and can also translate speech in any of them into English. Accuracy is best for English, Spanish, French, German, Portuguese, Italian, Japanese and Chinese. Less common languages, including Sinhala and Tamil, are recognized less reliably by these small models.
How do I make subtitles for a video?
Choose your video, turn on timestamps, and download the SRT or VTT file when it's done. Most video players, YouTube and editing apps can load these subtitle files.