Sauti / assemblyai
AssemblyAI Universal-3.5 Pro
AssemblyAI's current flagship
WER
12.3%
95% CI 10.8%–14.0%
Rank
#6
of 12 measured
Latency
6.4s
mean, per clip
Cost
$0.21
per audio hour
Every other model on the board is a statistical tie with this one. Its confidence interval overlaps all 11 of them, so the rank above tells you where it landed, not that it is better or worse than anything else here.
What it is
The pre-recorded flagship, billed per hour of audio submitted. Requests send an ordered speech_models array so unsupported languages fall back to universal-2.
Pick it when
When you need the wider language coverage, or when you want a model that still returns a transcript if the declared language is wrong rather than returning nothing.
Think twice when
When price matters: universal-2 tied it on our corpus at appreciably lower cost. The confidence intervals overlap almost entirely, so we cannot claim the pro model is more accurate on this data.
Assessment written against the board of 2026-08-08. The figures above are read live from the current published snapshot (2026-08-08), so if those dates differ, trust the figures.
What the vendor publishes
assemblyai publishes 5.9% on vendor-internal; we measured 12.3% on People’s Speech. Those are different corpora, so this is a difference in test set as much as in the model. It is comparable only when the benchmark matches. Source (read 2026-07-20).AssemblyAI's own comparison; internal set, NOT our People's Speech — not apples-to-apples.
How this was measured
250 clips (2.9 hours) of unscripted conversational English from People’s Speech (MLCommons, CC-BY), human-transcribed. Scored on a held-out split (N=77) never used to tune anything, through one fixed text normalizer, with bootstrap 95% confidence intervals. One corpus, one language, one modality: batch pre-recorded. It does not tell you how this model handles your accents, your domain vocabulary, or streaming.
Cost is assemblyai’s own published pre-recorded pay-as-you-go list price, read from their pricing page on 2026-08-08. Streaming and committed-volume rates differ.Published as $0.21/hr of audio submitted.