Sauti / openai

OpenAI gpt-4o-mini-transcribe

Budget tier of the gpt-4o generation

WER

12.8%

95% CI 11.4%–14.3%

Rank

#8

of 12 measured

Latency

1.6s

mean, per clip

Cost

$0.18

per audio hour

10 other models are a statistical tie with this one — their confidence intervals overlap, so the ordering between them is not something this data can settle: nova-3, gpt-transcribe, Whisper large-v3 (self-hosted), universal-2, universal-3-5-pro, Whisper small (self-hosted), whisper-1, scribe_v2, scribe_v1, gpt-4o-transcribe.

What it is

The cheap lane of the previous OpenAI generation, at half the price of gpt-4o-transcribe.

Pick it when

It is the interesting result in the gpt-4o family: it placed ahead of the full-size gpt-4o-transcribe on our corpus at half the price. The two intervals overlap, so read that as "no measurable accuracy penalty for the cheaper model" rather than as the mini winning.

Think twice when

It was the lowest-placed of the three current OpenAI options on our corpus, and also the cheapest of them. That is a genuine trade rather than a mistake: gpt-transcribe placed higher for more per hour, and whether the difference is worth paying is a call the overlapping intervals cannot make for you.

Assessment written against the board of 2026-08-08. The figures above are read live from the current published snapshot (2026-08-08), so if those dates differ, trust the figures.

What the vendor publishes

openai publishes no stated word error rate figure for this model on FLEURS. That absence is itself the finding. Source (read 2026-08-15).Same launch post and same treatment as gpt-4o-transcribe: FLEURS named, superiority over Whisper v2/v3 asserted, no stated English figure.

How this was measured

250 clips (2.9 hours) of unscripted conversational English from People’s Speech (MLCommons, CC-BY), human-transcribed. Scored on a held-out split (N=77) never used to tune anything, through one fixed text normalizer, with bootstrap 95% confidence intervals. One corpus, one language, one modality: batch pre-recorded. It does not tell you how this model handles your accents, your domain vocabulary, or streaming.

Cost is openai’s own published pre-recorded pay-as-you-go list price, read from their pricing page on 2026-08-08. Streaming and committed-volume rates differ.

See the full leaderboard →Methodology

Other models we measured