Cyber risk
Also called Cyber Risk
The danger that a model meaningfully boosts someone’s ability to hack, exploit, or attack systems.
Think of it like
Handing every amateur a master locksmith who explains exactly which pin to push and when.
Example
Evaluators check whether a model can find and weaponize a software vulnerability end to end, and safety layers refuse offensive-security requests that cross the line.
How it actually works
The dual-use tension is sharp: the same skills help defenders patch and attackers exploit. Evals focus on uplift for real offensive tasks — vuln discovery, exploit writing, autonomous intrusion — rather than textbook knowledge. Providers weigh security-research usefulness against clear misuse, drawing lines in the usage policy.
For product teams
A dual-use minefield where the same capability is a feature for defenders and a weapon for attackers.
For engineers
Measure end-to-end offensive uplift (vuln discovery through exploitation); refusals must separate legitimate security research from attack.
Related
- Dangerous capabilities — A key category within dangerous capabilities.
- Usage policy — Its rules live in the usage policy.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome