What you need to know about converting audio to text
Converting audio to text means taking a recording — a voice memo, podcast, interview, or meeting — and turning it into written words. You have three main routes: use software on your own computer, use a web-based tool in your browser, or send the file to a service that does the work for you. Which one works best depends on how long your audio is, whether it's clear or noisy, and whether you want the text right away or can wait a few hours.
The fastest option for short clips is usually a web tool like Otter.com or Rev.com — you upload, wait a few minutes to a few hours, and download the text. If you have lots of audio to convert regularly, desktop software like Express Scribe or Audacity with a speech-to-text plugin costs less over time. For very long files or poor audio quality, a human transcription service gives you accuracy but takes days and costs money.
Key Takeaways
- Web-based tools like Otter.com and Rev.com work in any browser and need no installation, but may charge per minute of audio.
- Desktop software like Express Scribe runs on your computer and works offline, but requires a one-time purchase or subscription.
- Audio quality matters — clear speech with no background noise converts much more accurately than recordings made in a car or crowded room.
- Most tools work best with English; if your audio is in another language, check the tool's language list before uploading.
- Human transcription services give near-perfect accuracy but cost more and take several hours to a few days.
Using a web tool to convert audio in your browser
Web-based tools are the easiest starting point because you do not need to install anything. Otter.com, Rev.com, and Descript all let you upload an audio file, and the tool converts it to text automatically. Most offer a free tier that covers a small amount of audio per month — typically 600 minutes for Otter, or a few files for Rev — so you can test before paying.
To use Otter.com: go to otter.com, create a free account, click "New Note" or "Upload," select your audio file from your computer, and wait. Otter shows you the text as it processes, usually finishing within 30 minutes for files under an hour. You can edit the text directly in Otter, download it as a document, or copy and paste it elsewhere. The free tier gives you 600 minutes per month; after that, you pay per additional minute.
Rev.com works similarly but focuses on accuracy. Upload your file, wait (usually 24 hours for the free tier, faster if you pay), and download the text. Rev also offers human transcription if you want near-perfect accuracy — a human listens to your audio and types it out, which costs more but catches words the software misses.
Descript is built for video and audio editing but includes transcription. Upload your file, Descript transcribes it, and you can edit the text and the audio together — if you delete a word from the text, the audio deletes too. This is useful if you are making a podcast or video and want to edit by transcript.
Converting audio on your computer with desktop software
If you convert audio regularly or have very long files, desktop software is often cheaper than paying per minute on a web tool. Express Scribe is the most common choice for transcription work. It costs about $50 one-time or $10 per month, plays audio at adjustable speeds, and lets you use foot pedals to pause and rewind without touching the keyboard.
To use Express Scribe: download it from nch.com.au/scribe, install it on your Windows or Mac computer, open the program, and load your audio file. Express Scribe does not transcribe automatically — it is a player that makes it easier for you to listen and type. You listen to a section, pause with a foot pedal or keyboard shortcut, type what you heard, and repeat. This is manual transcription, not automatic, so it takes time but gives you full control and high accuracy.
If you want automatic transcription on your computer, Audacity (free, audacityteam.org) can record audio, and some plugins add speech-to-text. However, Audacity's built-in speech recognition is limited. A better option is to use Windows Speech Recognition (built into Windows 10 and 11) or Mac Dictation (built into macOS) — both can transcribe audio if you route the audio through your microphone input, though this is slower than web tools.
Preparing your audio file for the best results
The quality of your audio file directly affects how accurately the tool converts it to text. Clear speech, minimal background noise, and a steady volume all make a difference. If your audio is recorded in a quiet room with one person speaking clearly, conversion is usually 90% accurate or better. If it is recorded in a car, coffee shop, or with multiple people talking, accuracy drops.
Before uploading, listen to your file and note any problems. If the volume is very quiet, you can increase it in Audacity (free) or your phone's voice memo app before uploading. If there is heavy background noise — traffic, music, wind — consider re-recording if possible. Most tools let you trim the file to remove long silences at the start or end, which saves processing time and money.
File format matters less than you might think. Most tools accept MP3, WAV, M4A, and OGG files. If your audio is in an unusual format, convert it first using a free tool like CloudConvert (cloudconvert.com) or Audacity.
Understanding the cost and time trade-offs
Free web tools usually give you 600 to 1,200 minutes per month at no cost, which is enough for occasional use. If you need more, you pay per minute — typically $0.10 to $0.25 per minute of audio. A one-hour recording costs $6 to $15 on most platforms. Some tools charge a flat monthly fee instead: Otter's paid plan is about $10 per month for unlimited transcription.
Desktop software like Express Scribe costs $50 to $100 one-time or $10 per month, but you do the transcription yourself by listening and typing, so there is no per-minute charge. This is cheaper if you have many hours to transcribe, but it takes your time. Human transcription services charge $1 to $2 per minute and take 24 to 48 hours, so a one-hour recording costs $60 to $120 but is nearly perfect.
Speed also varies. Web tools with automatic transcription finish in minutes to a few hours. Human transcription takes a day or more. Desktop software depends on how fast you can listen and type — typically 4 to 6 hours of work per hour of audio if you are experienced.
Fixing errors after conversion
No transcription tool is 100% accurate. Names, technical terms, and words that sound similar often get wrong. After conversion, read through the text and fix obvious errors. Most web tools let you edit directly in their interface; others let you download the text and edit it in Word or Google Docs.
Common mistakes include homophones (to/too/two, their/there), proper names spelled phonetically (Sasha becomes "Sasha" or "Sacha"), and technical terms the tool does not recognize. If your audio includes industry jargon or names, some tools let you add a custom dictionary before uploading, which improves accuracy for those words.
If accuracy is critical — for legal documents, medical records, or formal transcripts — use a human transcription service or hire a professional transcriber. The cost is higher, but the accuracy is near-perfect.
Frequently Asked Questions
Can I convert audio to text on my phone?
Yes. Both iPhone and Android have built-in voice-to-text for typing, but they do not transcribe existing audio files. For that, use a web tool like Otter.com or Rev.com in your phone's browser, or download the Otter app from the App Store or Google Play. Upload your audio file the same way you would on a computer.
What if my audio has multiple speakers?
Most automatic tools transcribe everything into one block of text without labeling who is speaking. Some tools like Descript and Otter can identify different speakers if the voices are distinct, but accuracy varies. For interviews or meetings where you need clear speaker labels, human transcription is more reliable.
Is the text I upload to these services private?
Web-based tools store your files on their servers. Read the privacy policy of the tool you choose — most say they use your audio only to transcribe it and do not share it with third parties. If your audio contains sensitive information, use desktop software instead, which processes everything on your computer with no upload.
Can I convert a video file to text?
Yes. Tools like Descript, Rev.com, and Otter.com accept video files and transcribe the audio. Upload your MP4, MOV, or other video file the same way you would an audio file, and the tool extracts the audio and converts it to text.
What languages does audio-to-text support?
Most tools support English well. Many also support Spanish, French, German, Mandarin, and other languages, but accuracy is usually lower than for English. Check the tool's language list before uploading audio in a language other than English.