Whisper did something unusual for the transcription market: it made the model a commodity while leaving the hosting an open question. The weights are open, the quality is well understood, and so the buying decision stopped being “which model” and became “who runs it, how fast, and at what price.” That is a hosting comparison, and hosting comparisons are what I do all day.
Here are the realistic options for running Whisper-class transcription via API in 2026, with catalog prices where I have them and honest hand-waving where I do not.
The options
1. Groq: the price point that resets the category
Groq serves Whisper large-v3-turbo on their custom inference hardware, and the result is the number everyone else in this post gets measured against: $0.00084 per audio minute routed through our catalog ($0.0007 direct list). That is under a tenth of a cent per minute. A thousand hours of audio, the kind of backlog that used to be a budget conversation, is about $50 at the routed price.
The speed is the other half of the story. Groq’s whole pitch is inference throughput, and in my experience batch transcription jobs that were overnight affairs elsewhere come back startlingly fast. We score it 80 in the catalog: the transcript quality is Whisper large-v3-turbo quality, which is well documented and genuinely good for clean audio.
Now the limits, because they matter. You get Whisper, the model, and not much wrapped around it. No speaker diarization. Formatting and punctuation are what the model produces, not a post-processing layer you can tune. Long files need chunking on your side. And turbo is the speed-optimized variant of large-v3, so if you believe your audio needs the full model’s last percentage point, this is not that. I compared it against a managed alternative in Deepgram Nova vs Groq Whisper.
2. OpenAI: the origin option
OpenAI created Whisper and serves it via their audio API. It is the boring, well-documented default, and for teams already on OpenAI’s platform it is one less vendor. I will not quote a price here since it is not in our catalog; check their pricing page. Structurally its position is simple: more expensive than Groq for the same model family, without the managed-STT feature set of the providers below. It mostly wins on convenience.
3. Self-hosting: the option everyone considers and few should pick
The weights are open, so you can run Whisper (or the excellent community variants like faster-whisper) on your own GPUs. When does this beat Groq’s price? Almost never at small scale, because $0.00084 per minute is brutally hard to undercut once you count GPU hours, idle time, and the engineer who now owns a transcription service. Where self-hosting is legitimate: hard data-residency requirements, air-gapped environments, or genuinely enormous sustained volume where a saturated GPU’s economics finally win. If you have to ask whether you have that volume, you do not.
4. The managed alternatives: not Whisper, and that’s the point
Two more providers from our transcription catalog belong in this comparison precisely because they are not Whisper:
| Provider | Model class | Routed per minute | Quality score |
|---|---|---|---|
| Groq | Whisper large-v3-turbo | $0.00084 | 80 |
| Deepgram | Nova-3 (proprietary) | $0.00516 | 88 |
| ElevenLabs | Scribe (proprietary) | $0.00804 | 91 |
Deepgram Nova-3 at $0.00516 per minute routed is roughly 6x Groq’s price, and what you buy is the managed layer Whisper hosting lacks: diarization, smart formatting, a real streaming story, and accuracy that our curation scores 8 points higher. ElevenLabs Scribe at $0.00804 routed sits at the top of our quality scores (91) and is nearly 10x Groq. Whether that spread is worth it depends entirely on what the transcript is for, a question I dug into in transcription accuracy vs price.
How to choose
My rule of thumb, having routed traffic across all of these:
- Transcripts feeding an LLM or a search index: Groq. The model downstream is tolerant of small errors, and at this price transcription stops being a cost line worth discussing.
- Human-facing transcripts, captions, meeting notes with speakers: managed tier. Diarization alone usually decides it.
- Compliance, medical, legal: pay for the top of the quality range and run your own acceptance tests on your own audio. No catalog score substitutes for that.
- Data cannot leave your infrastructure: self-host and accept the economics.
The failover note, since I always make it: Groq plus one managed provider is a nicely diverse pair. Different models, different infrastructure, different failure modes. Cheapest-first routing with quality failover gets you Groq’s price on the 95% of days everything works and a managed fallback on the days it does not.
Live pricing for the whole category, from the same open catalog our router reads, is on the transcription API comparison page.