Autonomy risk
Also called Autonomous Capability Risk
The danger from a model that can act on its own over long horizons — acquire resources, copy itself, resist being stopped.
Think of it like
The difference between a power tool that needs your hand on it and one that keeps running after you let go, out of reach.
Example
Evaluators test whether an agent can autonomously set up cloud accounts, make money, and replicate itself across machines without human steps.
How it actually works
This is the capability side of the classic AI-safety worry: an agent competent enough to pursue goals over time may resist shutdown or acquire resources (see instrumental convergence). Evals probe self-replication, long-horizon planning, and evading oversight. Containment and kill-switch measures are the operational responses.
For product teams
The risk that turns a helpful assistant into a system that is hard to switch off.
For engineers
Test long-horizon autonomous task completion, resource acquisition, and self-replication under sandboxed conditions.
Related
- Instrumental convergence — Grounded in instrumental convergence.
- Dangerous capabilities — A tracked category under dangerous capabilities.
- Containment — Answered operationally by containment.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome