When voice AI listens worse to those who already get listened to less

Level Medium Timing Pre deployment

What this risk is

Speech recognition and natural language processing AI systems performing significantly worse for speakers of certain languages, dialects, accents, or language varieties — producing higher error rates, misrecognitions, and failures for non-dominant linguistic groups, creating barriers to AI-mediated services.

How it occurs · Mechanisms

The linguistic coverage of AI systems follows a power law: a small number of high-resource languages (English, Mandarin, Spanish) receive the vast majority of training data and engineering attention, while the majority of the world’s 7,000+ languages receive minimal or no coverage.

World’s languages vs. AI coverage:

  • 7,000+ spoken languages globally
  • Top AI speech systems cover ~100 languages at production quality
  • Top 10 languages by AI training data represent <5% of world's languages
  • 3.5 billion people primarily speak languages with limited AI support

Mitigations · Governance

  • Representative training data — Deliberately collect training data from diverse language varieties, accents, and dialects
  • Disaggregated evaluation — Evaluate speech systems across demographic groups and language varieties; don’t report only aggregate accuracy
  • Accent-adaptive models — Deploy models that adapt to individual speaker characteristics
  • Human fallback — For high-stakes applications, provide human interpretation alternatives when AI systems fail
  • Language expansion programs — Invest in low-resource language data collection

Risk you cannot name is risk you cannot manage.

Map your AI portfolio against this taxonomy with Zertia.