Sauti / openai
OpenAI gpt-4o-mini-transcribe
Budget tier of the gpt-4o generation
WER
12.8%
95% CI 11.4%–14.3%
Rank
#8
of 12 measured
Latency
1.6s
mean, per clip
Cost
$0.18
per audio hour
10 other models are a statistical tie with this one — their confidence intervals overlap, so the ordering between them is not something this data can settle: nova-3, gpt-transcribe, Whisper large-v3 (self-hosted), universal-2, universal-3-5-pro, Whisper small (self-hosted), whisper-1, scribe_v2, scribe_v1, gpt-4o-transcribe.
What it is
The cheap lane of the previous OpenAI generation, at half the price of gpt-4o-transcribe.
Pick it when
It is the interesting result in the gpt-4o family: it placed ahead of the full-size gpt-4o-transcribe on our corpus at half the price. The two intervals overlap, so read that as "no measurable accuracy penalty for the cheaper model" rather than as the mini winning.
Think twice when
It was the lowest-placed of the three current OpenAI options on our corpus, and also the cheapest of them. That is a genuine trade rather than a mistake: gpt-transcribe placed higher for more per hour, and whether the difference is worth paying is a call the overlapping intervals cannot make for you.
Assessment written against the board of 2026-08-08. The figures above are read live from the current published snapshot (2026-08-08), so if those dates differ, trust the figures.
What the vendor publishes
openai publishes no stated word error rate figure for this model on FLEURS. That absence is itself the finding. Source (read 2026-08-15).Same launch post and same treatment as gpt-4o-transcribe: FLEURS named, superiority over Whisper v2/v3 asserted, no stated English figure.
How this was measured
250 clips (2.9 hours) of unscripted conversational English from People’s Speech (MLCommons, CC-BY), human-transcribed. Scored on a held-out split (N=77) never used to tune anything, through one fixed text normalizer, with bootstrap 95% confidence intervals. One corpus, one language, one modality: batch pre-recorded. It does not tell you how this model handles your accents, your domain vocabulary, or streaming.
Cost is openai’s own published pre-recorded pay-as-you-go list price, read from their pricing page on 2026-08-08. Streaming and committed-volume rates differ.