When AI develops capabilities no one trained it to have
What this risk is
AI systems developing capabilities not present in smaller versions of the same architecture and not explicitly trained for — capabilities that emerge unpredictably as model scale increases. Some of these emergent capabilities may be dangerous or may enable dangerous uses.
The core governance problem: Emergent capabilities are, by definition, not anticipated during the design and safety review of a system. A model can pass all safety evaluations for known capabilities and still deploy with unknown dangerous capabilities.
How it occurs · Mechanisms
Capabilities are considered emergent when:
- They are absent (near-zero performance) below a certain scale threshold
- They appear sharply (near-human or above-human performance) above that threshold
- They were not explicitly trained for and were not predicted
This phase-transition behavior makes capability prediction difficult: models below the threshold give no warning of capabilities that appear above it.
Mitigations · Governance
- Pre-deployment capability evaluations (evals) — Systematically test for dangerous capabilities before each new model version is deployed
- Capability thresholds — Define specific capability levels that trigger enhanced review (EU AI Act Art. 51, US EO compute thresholds)
- Red lines — Capabilities that must be absent before deployment; deployment blocked if present
- Staged deployment — Expand access gradually, monitoring for capability surprises at each stage
- External evaluation — Third-party capability evaluation for frontier models
—
Risk you cannot name is risk you cannot manage.
Map your AI portfolio against this taxonomy with Zertia.
