Sauti / assemblyai

AssemblyAI Universal-2

Prior generation, still sold as its own tier

WER

12.1%

95% CI 10.7%–13.9%

Rank

#5

of 12 measured

Latency

5.8s

mean, per clip

Cost

$0.15

per audio hour

Every other model on the board is a statistical tie with this one. Its confidence interval overlaps all 11 of them, so the rank above tells you where it landed, not that it is better or worse than anything else here.

What it is

A separately-priced product in its own right, not merely the fallback for languages the flagship does not cover.

Pick it when

It is the value result in AssemblyAI's lineup. It placed a hair ahead of Universal-3.5 Pro on our corpus at roughly seven-tenths of the price, with intervals so overlapped that paying more for the flagship bought nothing measurable on this data.

Think twice when

One concrete gotcha we hit: forced to a language the audio is not in, it returns an EMPTY transcript with a completed status and no error, where the pro model transcribes anyway. If your pipeline declares languages it is not certain about, that failure is silent.

Assessment written against the board of 2026-08-08. The figures above are read live from the current published snapshot (2026-08-08), so if those dates differ, trust the figures.

What the vendor publishes

assemblyai publishes no stated word error rate figure for this model on any public benchmark. That absence is itself the finding. Source (read 2026-08-15).Relative only: "a relative WER reduction of 3% compared to Universal-1" and "surpasses the nearest external model by 15% relative". Per-dataset numbers exist in their Table 1, but no consolidated English WER — they argue explicitly that single-number WER comparisons mislead.

How this was measured

250 clips (2.9 hours) of unscripted conversational English from People’s Speech (MLCommons, CC-BY), human-transcribed. Scored on a held-out split (N=77) never used to tune anything, through one fixed text normalizer, with bootstrap 95% confidence intervals. One corpus, one language, one modality: batch pre-recorded. It does not tell you how this model handles your accents, your domain vocabulary, or streaming.

Cost is assemblyai’s own published pre-recorded pay-as-you-go list price, read from their pricing page on 2026-08-08. Streaming and committed-volume rates differ.Published as $0.15/hr. Also the ordered-fallback model for languages universal-3-5-pro does not cover.

See the full leaderboard →Methodology

Other models we measured