Sauti / elevenlabs

ElevenLabs Scribe v2

Current Scribe generation

WER

13.6%

95% CI 12.3%–15.2%

Rank

#10

of 12 measured

Latency

3.2s

mean, per clip

Cost

$0.22

per audio hour

10 other models are a statistical tie with this one — their confidence intervals overlap, so the ordering between them is not something this data can settle: nova-3, gpt-transcribe, Whisper large-v3 (self-hosted), universal-2, universal-3-5-pro, Whisper small (self-hosted), gpt-4o-mini-transcribe, whisper-1, scribe_v1, gpt-4o-transcribe.

What it is

ElevenLabs' current speech recognition model, marketed at 90+ languages with keyterm prompting, entity detection and diarization as paid add-ons we did not enable.

Pick it when

For the features around the transcript rather than the transcript itself: diarization and keyterm prompting are the reason to be here, and the base rate is mid-pack.

Think twice when

If raw accuracy on unscripted conversational English is the goal. ElevenLabs markets Scribe as the most accurate model available; on our corpus it placed in the bottom third. That is one benchmark and not the multilingual case they lead with, but it is a direct check of a bold claim.

Assessment written against the board of 2026-08-08. The figures above are read live from the current published snapshot (2026-08-08), so if those dates differ, trust the figures.

What the vendor publishes

elevenlabs publishes no stated word error rate figure for this model on FLEURS. That absence is itself the finding. Source (read 2026-07-20).Claims top FLEURS accuracy across languages; capture the exact English FLEURS figure when confirmed.

How this was measured

250 clips (2.9 hours) of unscripted conversational English from People’s Speech (MLCommons, CC-BY), human-transcribed. Scored on a held-out split (N=77) never used to tune anything, through one fixed text normalizer, with bootstrap 95% confidence intervals. One corpus, one language, one modality: batch pre-recorded. It does not tell you how this model handles your accents, your domain vocabulary, or streaming.

Cost is elevenlabs’s own published pre-recorded pay-as-you-go list price, read from their pricing page on 2026-08-08. Streaming and committed-volume rates differ.Published as $0.22/hr. Entity detection (+$0.07/hr) and keyterm prompting (+$0.05/hr) are add-ons we do not enable.

See the full leaderboard →Methodology

Other models we measured