Why AI performs worse for some users and what to do about it
What this risk is
The accuracy and effectiveness of AI system decisions and actions vary depending on the user’s group membership. Decisions in AI system design and biased training data lead to unequal outcomes, reduced benefits, increased cognitive effort, and alienation of underrepresented users.
Unlike subdomain 1.1 (which focuses on discriminatory decisions), subdomain 1.3 focuses on differential reliability: the system works well for some groups and poorly for others, even when it does not make an explicitly discriminatory decision.
How it occurs · Mechanisms
Causal profile: AI-caused · Unintentional · Pre and post-deployment
- Underrepresentation in training data — Groups with less data produce lower-quality model outputs for those groups
- Benchmark bias — AI systems are evaluated on benchmarks that don’t reflect the diversity of real-world users
- Language and accent gaps — Speech recognition systems perform worse for non-native speakers, regional accents, and minority languages
- Skin tone bias in computer vision — Image recognition systems trained predominantly on lighter-skinned individuals underperform for darker-skinned users
—
Real-world incidents
Pulse Oximeters and AI Monitoring (2020–2022)
Multiple studies found that pulse oximeters — devices increasingly integrated with AI health monitoring — overestimated oxygen levels in patients with darker skin tones. This led to delayed treatment decisions and contributed to worse COVID-19 outcomes in Black patients.
Speech Recognition and African American Vernacular English (Stanford, 2020)
A Stanford study found that five major speech recognition systems (Apple, Amazon, Google, IBM, Microsoft) had error rates 35% higher for Black speakers than for white speakers, reaching up to 80% for some Black speakers.
Medical AI Underperforming for Women (Multiple studies, 2019–2023)
Several cardiac AI diagnostic tools showed significantly lower accuracy among women, who have been historically underrepresented in cardiac research datasets used for training.
Mitigations · Governance
- Disaggregated evaluation metrics — Report performance broken down by demographic group, not just aggregate accuracy
- Representative training datasets — Actively curate training data to ensure adequate representation of all groups
- Targeted data collection — Invest in collecting data from underrepresented populations
- Performance thresholds by group — Define minimum acceptable performance levels across all served populations
- User feedback mechanisms — Create channels for underrepresented users to report poor performance
—
Risk you cannot name is risk you cannot manage.
Map your AI portfolio against this taxonomy with Zertia.
