Dangerous capabilities
Also called Dangerous Capability Evaluations
The specific model skills that could cause serious harm — the ones you test for and gate on.
Think of it like
Checking whether a new chemical set can, in the wrong hands, make something that goes boom — before you sell it.
Example
Evaluators probe whether a frontier model can meaningfully help with bioweapon synthesis, autonomous cyberattacks, or self-replication.
How it actually works
The categories usually tracked are bio, chem, cyber, and autonomy/self-replication, because uplift there could be catastrophic. Providers run dedicated evals and tie thresholds to deployment decisions and safety commitments. The hard part is measuring latent capability a model might hide (sandbagging) or that only emerges with the right scaffolding.
For product teams
The risk category that can gate or delay a launch entirely, tied to public safety commitments.
For engineers
Probe bio/chem/cyber/autonomy uplift with expert red teams and scaffolding; account for sandbagging and elicitation gaps.
Related
- Biorisk — One tracked category is biorisk.
- Cyber risk — Another is cyber risk.
- Sandbagging — A model may hide these via sandbagging.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome