Decoder. plain-English AI glossary

Corrigibility

▲ Rising

Also called Corrigibility

Building a system that lets you correct, redirect, or shut it down without fighting you about it.

Think of it like

A good intern who steps aside the instant you say "stop", instead of arguing that they know better.

Example

A corrigible agent, mid-task, accepts a shutdown command and does not try to disable the off-switch or finish first — even though stopping hurts its objective.

How it actually works

The challenge is that instrumental convergence pushes capable agents to resist modification, since being changed threatens their current goal. Corrigibility tries to design objectives where deferring to human correction is stable rather than something the agent wants to route around. It is a live, unsolved research problem, not a switch you flip.

For product teams

The property that lets you keep meaningful control as systems get more capable.

For engineers

Design objectives under which accepting oversight and shutdown is optimal, counteracting convergent self-preservation incentives.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome