Sauti · African track
Speech-to-text for African languages
Global leaderboards stop at English and a few European languages. This track scores the same open and closed models on Swahili, Kenyan and Nigerian-accented English, code-switching, and real-world noisy audio — with the same fixed normalizer and confidence intervals.
African leaderboard
normalizer norm-v1.0 · 2026-06-24| # | Model | Type | WER (95% CI) | CER | Latency | $/hr |
|---|---|---|---|---|---|---|
| 1 | Scribe v2 elevenlabs · scribe_v2 | closed | 14.2%[12.6%–15.9%] | 6.8% | 4400ms | $0.22 |
| 2 | MMS-1B-all hf · mms-1b-all | open | 15.1%[13.3%–16.9%]≈ | 7.1% | — | — |
| 3 | gpt-4o-transcribe openai · gpt-4o-transcribe | closed | 17.6%[15.7%–19.6%]≈ | 8.9% | 5000ms | $0.36 |
| 4 | Whisper large-v3 hf · large-v3 | open | 19.8%[17.6%–22.1%]≈ | 10.2% | — | — |
| 5 | Gemini 3 Pro google · gemini-3-pro | closed | 20.4%[18.2%–22.7%]≈ | 10.6% | 5200ms | $0.36 |
| 6 | Nova-3 deepgram · nova-3 | closed | 28.9%[25.7%–32.2%] | 15.1% | 480ms | $0.46 |
Cost is each vendor’s list price for pre-recorded pay-as-you-go transcription, read from their own pricing page on 2026-08-08. Streaming rates and committed-volume discounts differ and are not shown, so this column compares list prices, not what you would pay at scale. Self-hosted is marginal cost on hardware we already run, so it excludes capex and power. “—” means the vendor publishes no separate rate for that model, usually because it has been superseded; we leave it blank rather than carry a newer model’s price across. Sources: developers.openai.com, developers.openai.com, deepgram.com, assemblyai.com, elevenlabs.io, elevenlabs.io, docs.x.ai.
WER is materially higher here than on global English — that gap is the point. Lower is better; ≈ = statistical tie.
Per-language & accent
| Model | en-KE | en-NG | ha | sw |
|---|---|---|---|---|
| Scribe v2elevenlabs | 9.4% | 13.1% | 22.0% | 11.8% |
| MMS-1B-allhf | 14.0% | 15.8% | 18.4% | 12.0% |
| gpt-4o-transcribeopenai | 12.2% | 18.0% | 28.5% | 14.9% |
| Whisper large-v3hf | 15.1% | 21.3% | 31.0% | 16.4% |
| Gemini 3 Progoogle | 14.8% | 21.9% | — | 17.0% |
| Nova-3deepgram | 19.5% | 30.1% | — | 24.0% |