Typing out a recording by hand takes about four times as long as the recording itself. AI speech recognition can do the first draft in minutes: interviews, lectures, meetings, voice notes and videos. Here's how to get an accurate transcript, for free, without uploading your audio.
How it works
The speech to text tool uses Whisper, a speech recognition model released by OpenAI in 2022 and trained on hundreds of thousands of hours of speech in many languages. The model is downloaded to your browser once and runs on your own device, so even confidential recordings stay private.
Step by step
- Choose the audio or video file. MP3, M4A, WAV and MP4 all work, as do voice notes from your phone.
- Choose the language, or let it detect the language automatically.
- Pick Fast for clear recordings or Accurate for accents, background noise or other languages.
- Wait while it listens. On a recent laptop, a minute of speech takes roughly 10 to 30 seconds.
- Read through the result, fix mistakes, then copy it or download it as text or subtitles.
Getting better accuracy
- Clean the audio first. Background noise causes most mistakes. Run noisy recordings through the noise remover first.
- Record closer to the speaker. For the next recording, put the phone near whoever is talking.
- Set the language. Automatic detection usually works, but choosing it avoids mistakes on short or mixed recordings.
- Expect to fix names. Names, brands and technical terms are the most common errors. Paste the text into the online notepad and use find and replace to fix each one everywhere at once.
Turning it into something useful
- Subtitles: download an SRT file for YouTube, or burn captions into a video with Add Subtitles to Video.
- Notes: paste the text into a document and add headings for each topic.
- Quotes: with timestamps on, you can find exactly where something was said.
Always ask permission before recording people, and check your local rules for recording calls and meetings.