Which audio format each transcription service accepts
They have a recording and a service that rejected it, or they are about to pay for transcription and want to submit something that will not bounce. Every service publishes a different list and none of them agree.
MP3 is the answer if you only want one. Every mainstream transcription service accepts it, and speech survives a low MP3 bitrate far better than music does. WAV is the second universal — accepted everywhere, but seven to ten times larger, and worth nothing extra on a recording that has already been through a lossy encoder. Before you convert anything, try the file you already have: MP3, M4A, WAV and MP4 are on almost every published list, so a voice memo or a Zoom recording usually goes straight through. The formats that genuinely get refused are Opus, AMR, WMA, Core Audio .caf and the proprietary dictation formats from Olympus and Philips recorders. For those, a 64 to 96 kbps MP3 is the conversion to make.
Try what you already have first
Half the people searching this have a file that would have been accepted. iPhone Voice Memos produce M4A, Android recorders produce M4A or MP3, Zoom writes M4A alongside the video, and Teams and Meet write MP4. All four are on every list in the table below, and converting one first costs you time plus a second generation of lossy encoding for nothing.
Work out what the file actually is first. The extension is not always the truth — a file named .mp3 can hold AAC and be refused on inspection. Dropping it into the converter on this site reads the real codec, sample rate, channel count and bitrate out of the header and shows them, without converting anything.
The refusals worth knowing about:
- Opus, usually arriving as
.opusor an.oggfrom WhatsApp, Telegram or Signal. Only some services list it. Getting a WhatsApp voice note into MP3 covers the whole route. - AMR and 3GA, from older phone call recorders and some Samsung handsets. Nothing accepts these.
- WMA, from Windows-era voice recorders and dictaphone software.
- .caf, which some iOS recording apps write instead of M4A.
- .dss and .ds2, the Olympus and Philips dictation formats. These need the manufacturer's own software or a service that specifically supports them; a general converter will not read them.
What each service accepts
Each of these comes from the service's own supported-file-types help page, read on 24 August 2026. They change, and no two agree, which is why the safe move is to match the intersection rather than the edges.
| Service | Audio formats it publishes | Best thing to send | Worth knowing |
|---|---|---|---|
| Rev | MP3, WAV, M4A, AAC, FLAC, OGG, WMA, plus MP4, MOV, AVI and most video | The original, unconverted | The widest list of the five. Human transcription is priced by the minute, so trimming dead air saves money as well as time. |
| Otter.ai | MP3, AAC, WAV, AIFF, M4A, WMA, plus MP4, MOV, WMV, AVI | M4A or MP3 | No FLAC and no Opus. A WhatsApp voice note has to be converted. |
| Word (Microsoft 365, Transcribe) | MP3, WAV, M4A, MP4 | MP3 | Lives in Word for the web at Home ▸ Dictate ▸ Transcribe, not in desktop Word. There is a per-file size cap and a monthly budget of transcribed minutes. |
| OpenAI's transcription API (Whisper and its successors) | flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, webm | MP3 at 64–96 kbps | A fixed per-request size cap, historically 25 MB, is what stops most long meetings. |
| Descript | MP3, WAV, M4A, AAC, FLAC, OGG, plus most video | WAV if you will edit, MP3 if you will not | It keeps the audio for editing afterwards, so quality matters more here than in a pure transcription tool. |
Format tables on transcription blogs go stale quietly, and some of the ones ranking for this query still list services that no longer exist. If a deadline or an invoice depends on it, open the service's own help page and check. The intersection above — MP3, WAV, M4A — has been stable for years and is the safest thing to aim at.
Why WAV is not automatically the better choice
The instinct is that uncompressed input must produce a better transcript. It does not, for two reasons.
The first is that speech recognition models do not listen the way you do. Nearly all of them resample the input to 16 kHz mono and convert it to a spectrogram before anything else happens. Everything above roughly 8 kHz is discarded by the model itself. A 320 kbps stereo MP3 and a 96 kbps mono one arrive at that stage looking very similar.
The second is that converting to WAV cannot bring back what a lossy encoder already removed. If your recording is an M4A, exporting it to WAV writes out the samples the AAC decoder produced and nothing else. The missing detail is still missing, in a file roughly ten times the size. WAV against MP3, with the arithmetic goes through what that trade actually costs.
WAV is the right answer in three cases: the recording has never been compressed and you are sending the original; the service is also an editor and you will work on the audio afterwards; or the file is short enough that size is irrelevant. Two minutes of WAV is about 20 MB and nobody cares. Two hours is about 1.2 GB and everybody does.
Mono, sample rate, and what the model actually hears
Mono is fine for speech-to-text, and often preferable. A single microphone recording a room writes two identical channels and stores the same audio twice, so downmixing halves the file with no effect on the transcript. If your recorder has a mono setting, turn it on before the interview rather than fixing it afterwards.
Sample rate is the same story. 16 kHz is enough, because that is what the model resamples to anyway; recording at 48 kHz triples the size of a WAV and changes nothing downstream. It matters only if you might also publish the audio.
Concretely, for one hour of speech:
| What you send | Approximate size for one hour |
|---|---|
| WAV, 44.1 kHz stereo, 16-bit | 605 MB |
| WAV, 16 kHz mono, 16-bit | 115 MB |
| MP3, 192 kbps stereo | 86 MB |
| MP3, 96 kbps | 43 MB |
| MP3, 64 kbps | 29 MB |
If each speaker was genuinely recorded onto their own channel — a two-input interface, a Zoom multi-track export, some podcast recorders — that channel separation is real information about who is talking. Downmixing discards it. Send the stereo file, or the separate tracks if the service takes them.
When it is refused for size rather than format
Most rejections on long recordings have nothing to do with the codec. Every service caps a single submission somewhere, and a two-hour meeting recorded as a stereo WAV hits that ceiling whatever the format list says.
The order to work in: convert to MP3 first, at 64 or 96 kbps for speech; trim the silence at the start and end second; and split the recording only if the first two are not enough. Going from a 44.1 kHz stereo WAV to a 64 kbps MP3 takes an hour of audio from about 605 MB to about 29 MB, which is a factor of twenty and usually the end of the problem. Fitting a long recording under a transcription size limit covers the splitting step, including why the pieces should overlap by a few seconds.
Do not go below 64 kbps hoping to save more. Speech starts to smear at 32 kbps and below, and a smeared consonant is exactly what a recogniser gets wrong. The bytes saved are few; the accuracy cost is not.
What is worth doing to the recording, and what is not
Worth doing: cut the dead air. Five minutes of chair-shuffling before the interview starts is five minutes of billed time on a per-minute service and five minutes of hallucinated filler on an automatic one. Cutting to the actual conversation is the single highest-value edit.
Not worth doing: normalising, boosting the bass, or running heavy noise reduction. Peak normalisation lifts the whole waveform and leaves the ratio between the voice and the room exactly as it was, so a recogniser sees no difference. Aggressive noise reduction is worse than useless: it eats the quiet consonants along with the hiss.
Worth checking: whether the audio is playing in only one channel. A recording with one silent side still transcribes, but downmixing it to mono afterwards drops the voice 6 dB for no reason. Fix the channel first.
Converting the file without sending it anywhere
Confidential recordings are the normal case here: interviews, client calls, medical dictation, legal work. The service that transcribes it will necessarily see it, and that is a decision you have already made. The converter in between does not have to.
The converter on this page is FFmpeg compiled to WebAssembly, running in the browser tab. The file is read by JavaScript on your own machine and never leaves it; there is no server on the other end to receive it. The engine is a 32 MB download on your first visit, cached for a year after that. Pick MP3 and 64 or 96 kbps for speech.
Two related routes, if they match what you have. An iPhone voice memo is M4A and usually goes through as it is, but converting one to MP3 is the fallback when a service refuses it. A Zoom recording ships as both an MP4 and an M4A, and pulling the audio out of the MP4 needs no encoder at all when the video's soundtrack is already AAC — the stream is copied across untouched, so it finishes in about a second however long the meeting was.
The converter on the home page handles this. Free, no upload, no sign-up.