How-to
How to transcribe audio to text: every method, step by step
By the exma team · August 5, 2026 · 8 min read
TL;DR: The fastest path for most people: upload the file to an AI transcription service and get a speaker-labeled, timestamped transcript in minutes. Need text during the event? Use live transcription. Need certified accuracy? AI draft + human review. Comfortable with a terminal? Whisper is free. Typing it yourself costs 4–6 hours per hour of audio — worth it almost never.
Method 1: Upload to an AI transcription service (fastest for most people)
If the recording already exists, this is almost always the answer. Modern engines exceed 95% accuracy on reasonable audio and return an hour-long file in minutes — with punctuation, speaker labels, and timestamps you'd never add by hand. (How the engines pull that off is covered in speech to text, explained.)
- Pick a service and sign up. Free tiers are enough to test. For confidential or legal audio, pick a tool that encrypts data and never trains models on it — that's exma's default posture.
- Upload the file. MP3, WAV, M4A, MP4 and similar all work. Video uploads fine too — the audio track is extracted automatically, no conversion step.
- Wait a few minutes. The engine transcribes, separates speakers, and timestamps every word.
- Review with synced audio. Click any line to hear the source. Fix names and the occasional unclear passage — this is minutes of work, not hours, because you're checking, not typing.
- Export to TXT, DOCX, or PDF with speakers and timestamps intact.
Method 2: Transcribe live, while it happens
When you need the text during the event — a meeting you want searchable while it runs, a hearing that needs read-backs, live captions — use real-time transcription instead of recording now and uploading later. Words appear within a fraction of a second and the draft is complete the moment the event ends. exma captures from the microphone, the computer's system audio (any meeting platform, no bot joining the call), or both. The mechanics are in our real-time transcription guide.
Method 3: Human transcription service
Professional transcribers (Rev, GoTranscript, and similar) charge roughly $1.50–2.50 per audio minute and return files in a day or more. Where they earn it: genuinely bad audio — heavy crosstalk, distant microphones, thick accents — and transcripts that must carry a certification. For everything else, AI first: even human-certified workflows now start from an AI draft because correcting is faster than typing. The full trade-off is in human vs. AI legal transcription.
Method 4: Do it yourself with open-source Whisper
OpenAI's Whisper is free, open-source, accurate, and runs on your own machine — audio never leaves your computer. The cost is your time: install Python or a wrapper app, run the model, and accept that speaker labels, editing, exports, and support are all yours to build. Right for developers and the privacy-absolute; wrong for anyone who just needs the transcript today.
Method 5: Type it yourself
The honest math: 4–6 hours of typing per hour of audio, more with crosstalk or bad recordings, even with a foot pedal and playback tricks. The only cases where manual still wins: a two-minute clip you can knock out now, or audio so sensitive no third party — human or machine — may hear it (and even then, local Whisper usually beats typing).
The five methods, compared
| Method | Speed (per audio hour) | Cost | Speaker labels | Best for |
|---|---|---|---|---|
| AI upload | Minutes | Free tier → dollars | Automatic | Almost everything |
| Live transcription | Instant — done when the event ends | Free tier → dollars | Automatic | Meetings, hearings, captions |
| Human service | A day or more | ~$90–150 | Yes | Terrible audio, certification |
| Whisper (DIY) | Minutes–hours + setup | Free | Not built in | Developers, absolute privacy |
| Manual typing | 4–6 hours | Your time | You write them | Tiny clips; audio nobody else may hear |
Getting a better transcript from any method
- Fix the audio before the transcription. A phone in the middle of the table beats a laptop mic at one end; a quiet room beats software cleanup afterwards.
- Capture digitally when possible. Recording a remote call's system audio is dramatically cleaner than re-recording the speaker's output through a microphone.
- One voice at a time. Crosstalk is the single biggest accuracy killer for machines and humans alike.
- Confidential audio: check the data terms first. Retention, encryption, and model-training rights — the 12-point security checklist is the pre-flight list.
Frequently asked questions
How long does it take to transcribe one hour of audio?
AI: minutes. Human service: a day or more, less with rush fees. Yourself: four to six hours of work.
Can I transcribe audio to text for free?
Yes — free tiers of AI services (exma, Otter, Notta) for limited monthly volume, or Whisper unlimited if you can run it yourself.
Can I transcribe a video?
Upload it directly — MP4, MOV and similar work; the audio track is extracted automatically.
What's the most accurate method?
For most audio, a good AI service. For hard audio or certified transcripts, an AI draft reviewed by a human against the recording — accuracy plus a signature.
Transcribe your first recording free
Upload audio or video, or stream a live session — speaker labels, synced timestamps, and TXT/DOCX/PDF exports included.
Create your free workspace