Sauti / openai

OpenAI gpt-transcribe

OpenAI's current flagship transcription model

WER

11.8%

95% CI 10.3%–13.6%

Rank

#3

of 12 measured

Latency

2.0s

mean, per clip

Cost

$0.27

per audio hour

Every other model on the board is a statistical tie with this one. Its confidence interval overlaps all 11 of them, so the rank above tells you where it landed, not that it is better or worse than anything else here.

What it is

The current generation of OpenAI transcription, replacing the gpt-4o-transcribe family. Unusually, it is both better placed and cheaper than the model it supersedes.

Pick it when

It is the OpenAI model to reach for. It placed highest of their four models on our corpus and costs materially less per hour than gpt-4o-transcribe. The accuracy gap sits inside the confidence intervals so we will not call it a win, but the price gap does not, and it points the same way.

Think twice when

Like every LLM-based transcriber here it is non-deterministic: repeat runs of the same clip differ slightly. If you need reproducible output for the same input, that is a real constraint.

Assessment written against the board of 2026-08-08. The figures above are read live from the current published snapshot (2026-08-08), so if those dates differ, trust the figures.

What the vendor publishes

openai publishes no stated word error rate figure for this model on FLEURS. That absence is itself the finding. Source (read 2026-08-15).Names Common Voice and FLEURS but publishes no absolute WER for any language, English included — only "lower word error rates than prior models" and a hallucination-rate comparison (~90% fewer than Whisper v2). No figure to check against ours.

How this was measured

250 clips (2.9 hours) of unscripted conversational English from People’s Speech (MLCommons, CC-BY), human-transcribed. Scored on a held-out split (N=77) never used to tune anything, through one fixed text normalizer, with bootstrap 95% confidence intervals. One corpus, one language, one modality: batch pre-recorded. It does not tell you how this model handles your accents, your domain vocabulary, or streaming.

Cost is openai’s own published pre-recorded pay-as-you-go list price, read from their pricing page on 2026-08-08. Streaming and committed-volume rates differ.Current flagship. Cheaper than the gpt-4o generation it replaces.

See the full leaderboard →Methodology

Other models we measured