Choosing a transcription model when accuracy is the easy part
Match a transcription model to your real audio by settling live versus recorded, speaker count, language and storage rules before you go anywhere near an error rate.
on this page · 0 / 0 checked
Every transcription vendor publishes one number, and it is always a word error rate. The numbers are real, they are getting better, and they are close enough together that the comparison feels settled the moment you see them side by side. So you pick the low one, wire it up, and a fortnight later you are hand-editing transcripts of client calls because the model merged two speakers into one, dropped the bit where someone said a price, and confidently wrote a sentence nobody uttered during a pause.
Nothing went wrong with the accuracy number. The choice was decided by four things the number does not describe: whether anyone is waiting on the words, how many voices are in the recording, which language and accent they are speaking, and where the audio is legally and contractually allowed to end up. Get those four right and almost any current model will do. Get them wrong and the best model on the leaderboard produces a transcript you cannot use. This guide is for a freelancer or small team transcribing their own calls, interviews and meetings, or wiring transcription into a small product. If you are building live captioning at scale or transcribing regulated clinical records, you need vendor evaluation and legal review, not a guide. Prices below are US list prices fetched on 4 September 2026.
Live and recorded are two different products at two different prices
Vendors sell these as separate models, not as a toggle. Google ships gemini-3.5-transcribe-live through the Live API for interactive voice, and a separate gemini-3.5-transcribe through the Interactions API for recorded audio, meetings and call logs [1]. OpenAI lists gpt-live-transcribe and gpt-transcribe as distinct line items [3]. Deepgram and AssemblyAI both split their price tables into streaming and pre-recorded columns [4][5].
The gap is mostly money. Google prices audio input at $0.005 per minute on the live model and $0.003 per minute on the recorded one, with text output billed on top at $0.004 and $0.002 per minute respectively [2]. OpenAI’s gap is wider: $0.017 per minute for gpt-live-transcribe against $0.0045 for gpt-transcribe [3]. AssemblyAI charges $0.45 per hour for Universal-3.5 Pro Realtime and $0.21 per hour for the same model asynchronously [5]. Deepgram’s Nova-3 monolingual runs $0.0048 per minute streaming and $0.0043 pre-recorded [4].
The gap is also accuracy, on the identical model family. Google reports an average word error rate of 4.0% in streaming mode and 2.6% for non-streaming use cases, and 5.50% against 5.04% on the FLEURS benchmark [1]. That is the same vendor, the same release, measured twice. Streaming has to commit to words before the sentence finishes, and it pays for that.
So the first question is not which model. It is whether a human is sitting there waiting. If nobody is waiting, you want the recorded endpoint, and you are choosing a cheaper and more accurate product than the one the launch post led with.
The number of people in the room is a limit, not a setting
Speaker labels are where transcription quietly stops working, and the constraint is usually buried a paragraph below the headline. Gemini 3.5 Transcribe attributes speech in pre-recorded audio with timestamps for up to three speakers, and Google states that support for three or more is experimental [1]. That is fine for a client call and a founder interview. A four-person standup, a panel, or a workshop with a room mic sits outside the supported range, and the failure mode is not an error message. It is a transcript that reads plausibly with the wrong names attached.
Other vendors treat it as a paid extra rather than a default. AssemblyAI bills speaker diarization at $0.02 per hour on top of the base rate, for both async and streaming [5]. That is a trivial amount of money and an easy thing to forget to switch on, which is worse.
It is also worth being precise about what diarization gives you. It separates voices into Speaker A and Speaker B. It does not know that Speaker A is your client. Mapping labels to names is your job, and on a recording where someone joins twenty minutes late the mapping usually shifts halfway through. Budget for that step rather than being surprised by it.
”Supports 85 languages” is a coverage claim, not an accuracy claim
Gemini 3.5 Transcribe automatically detects and transcribes over 85 languages [1]. That sentence tells you the model will return text. It does not tell you the text will be right.
The clearest statement of the problem comes from OpenAI’s own Whisper model card, which is unusually direct about it. The models “perform unevenly across languages,” with “lower accuracy on low-resource and/or low-discoverability languages,” and they “exhibit disparate performance on different accents and dialects of particular languages, which may include a higher word error rate across speakers of different genders, races, ages, or other demographic criteria” [6]. That is a vendor saying, in writing, that a single headline error rate averages over people it serves badly.
Vendor price lists agree, if you read them as evidence. Deepgram charges more for multilingual than monolingual on the same model generation: Nova-3 multilingual is $0.0052 per minute pre-recorded against $0.0043 for monolingual, and $0.0058 against $0.0048 streaming [4]. Multilingual is a different product with different economics, not a checkbox on the same one.
The practical move is narrow. Do not ask whether your language is supported. Take a recording of the specific person whose voice you will be transcribing most often, in the accent and the setting they actually use, and run it. A model that handles broadcast Spanish beautifully can fall apart on a fast regional speaker on a bad phone line, and no supported-languages list will warn you.
Compare cost per hour of audio, then notice that your own time dwarfs it
Vendors quote in incompatible units on purpose. Google and OpenAI publish per-minute prices, sometimes alongside a per-million-token price for the same model [2][3]. AssemblyAI publishes per hour, at $0.21 for Universal-3.5 Pro and $0.15 for Universal-2 [5]. Deepgram publishes per minute [4]. Convert everything to one unit before you compare anything, because a per-minute figure with three leading zeros is very hard to hold in your head next to a per-hour figure.
Two things distort the comparison once you have it. The first is promotional pricing. Deepgram lists Nova-3 monolingual streaming at $0.0048 per minute as a promotional rate against a regular rate of $0.0077, and Flux English at $0.0065 promotional against $0.0077 regular [4]. Budget on the regular rate, because that is the one you will eventually pay. The second is free credit, which is generous enough to hide the real cost for months. Deepgram gives new accounts $200 [4]. AssemblyAI’s free tier covers up to 185 hours of pre-recorded transcription and up to 333 hours of streaming [5]. Both are enough to run a small operation for a long time without ever seeing a bill, which is fine until you plan around it.
The larger distortion is that the model fee is rarely the expensive part. If you spend ten minutes tidying speaker labels and fixing names for every hour of audio, that ten minutes costs you far more than the transcription did at any realistic hourly rate. Which means the model that saves you two minutes of cleanup per hour is worth paying several times more for, and the cheapest model on the list is often the most expensive thing in the workflow.
Vendor fee plus your cleanup time. A per-minute price becomes a per-hour price when you multiply it by 60. Computed in the page; nothing is sent anywhere.
Where the recording is allowed to live decides more than which model reads it
Audio of other people is a different category of data from your own notes, and the terms differ sharply by tier. Google’s Gemini API terms state that for unpaid services, Google uses submitted content to provide, improve and develop Google products and services including machine learning technologies, and that human reviewers may read, annotate and process API input and output [7]. For paid services, Google does not use prompts or responses to improve its products, and logs data solely for detecting and preventing violations of the Prohibited Use Policy [7]. Those are two genuinely different arrangements, and the free tier is the one people prototype on and then forget to leave.
If the recording contains a client, a candidate or a patient, read that paragraph before you upload rather than after. Free-tier convenience is not worth a clause you would have to explain later.
The obligation also starts earlier than the model, at the moment you press record. California Penal Code section 632 makes it an offence to intentionally record a confidential communication without the consent of all parties, with a fine of up to $2,500 per violation for a first offence and up to $10,000 for a repeat [8]. Other jurisdictions differ, and a call with participants in several of them is governed by more than one rule. None of this is a transcription question, and no model setting addresses it. Say at the top of the call that you are recording, get the answer on the recording, and keep the file.
Test on your worst recording, not your cleanest
Build a test set before you shortlist anything. Ten minutes is enough, and it should be ten minutes you would rather not listen to: the call from a cafe, the one where two people talk over each other, the one full of product names and initialisms, the one where the client dialled in from a car. Run each candidate over the same ten minutes and read the output with the audio playing.
Read it, specifically, rather than skimming it. The errors that hurt are not the ones that look like errors. Whisper’s model card is blunt about this: the predictions “may include texts that are not actually spoken in the audio input (i.e. hallucination)” [6]. Invented text arrives in fluent, correctly punctuated sentences that sit comfortably in the paragraph around them. A skim will not catch a fabricated commitment; a read against the audio will.
Rerun the test when you change anything. A new model version, a new microphone, a new recurring participant with an accent the old model handled and the new one does not. Ten minutes of listening once a quarter is cheaper than any of the alternatives.
What still goes wrong
Invented text remains the sharpest edge and none of this removes it. Whisper’s card recommends against use in high-risk domains like decision-making contexts, where flaws in accuracy lead to pronounced flaws in outcomes [6]. That applies to any transcription model, not just Whisper. A transcript is a good record of roughly what was said. It is not evidence, it is not a contract, and it should not be the only artefact behind a decision that matters. If a number or a commitment in a transcript is load-bearing, listen to that section of the audio before you act on it.
Speaker attribution degrades in exactly the conditions where you most want it. Cross-talk, a shared room mic, someone joining late, a participant on speakerphone. Google’s own three-speaker ceiling on Gemini 3.5 Transcribe is a fair marker of where the technology currently sits [1], and vendors that advertise no ceiling are generally not claiming the labels will be right, only that they will exist.
The ground also moves under you. Gemini 3.5 Transcribe is in public preview [1], which means behaviour and availability can change without the courtesy owed to a stable product. Promotional pricing expires, and Deepgram’s own table already shows the landing point, listing Nova-3 monolingual streaming at a promotional $0.0048 per minute against a regular rate of $0.0077 [4]. Neither of those is a reason to avoid a good model. They are a reason to keep your transcription step behind one function you can change in an afternoon, rather than threaded through six scripts and a Zapier chain, so that when the price or the policy moves you are editing a line rather than rebuilding a pipeline.
- 01Google — Gemini 3.5 Transcribeblog.google
- 02Google — Gemini API pricingai.google.dev
- 03OpenAI — API pricingdevelopers.openai.com
- 04Deepgram — Pricingdeepgram.com
- 05AssemblyAI — Pricingassemblyai.com
- 06OpenAI — Whisper model cardgithub.com
- 07Google — Gemini API Additional Terms of Serviceai.google.dev
- 08California Penal Code § 632law.justia.com