When voice AI listens worse to those who already get listened to less
Level
Medium
Timing
Pre deployment
What this risk is
Speech recognition and natural language processing AI systems performing significantly worse for speakers of certain languages, dialects, accents, or language varieties — producing higher error rates, misrecognitions, and failures for non-dominant linguistic groups, creating barriers to AI-mediated services.
How it occurs · Mechanisms
The linguistic coverage of AI systems follows a power law: a small number of high-resource languages (English, Mandarin, Spanish) receive the vast majority of training data and engineering attention, while the majority of the world’s 7,000+ languages receive minimal or no coverage.
World’s languages vs. AI coverage:
- 7,000+ spoken languages globally
- Top AI speech systems cover ~100 languages at production quality
- Top 10 languages by AI training data represent <5% of world's languages
- 3.5 billion people primarily speak languages with limited AI support
Mitigations · Governance
- Representative training data — Deliberately collect training data from diverse language varieties, accents, and dialects
- Disaggregated evaluation — Evaluate speech systems across demographic groups and language varieties; don’t report only aggregate accuracy
- Accent-adaptive models — Deploy models that adapt to individual speaker characteristics
- Human fallback — For high-stakes applications, provide human interpretation alternatives when AI systems fail
- Language expansion programs — Invest in low-resource language data collection
—
Risk you cannot name is risk you cannot manage.
Map your AI portfolio against this taxonomy with Zertia.
