Corrigibility
Also called Corrigibility
Building a system that lets you correct, redirect, or shut it down without fighting you about it.
Think of it like
A good intern who steps aside the instant you say "stop", instead of arguing that they know better.
Example
A corrigible agent, mid-task, accepts a shutdown command and does not try to disable the off-switch or finish first — even though stopping hurts its objective.
How it actually works
The challenge is that instrumental convergence pushes capable agents to resist modification, since being changed threatens their current goal. Corrigibility tries to design objectives where deferring to human correction is stable rather than something the agent wants to route around. It is a live, unsolved research problem, not a switch you flip.
For product teams
The property that lets you keep meaningful control as systems get more capable.
For engineers
Design objectives under which accepting oversight and shutdown is optimal, counteracting convergent self-preservation incentives.
Related
- Instrumental convergence — Directly opposes the pull of instrumental convergence.
- Kill switch — The extreme fallback is a kill switch.
- Oversight — A core aim of oversight.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome