Decoder. plain-English AI glossary

Scalable oversight

▲ Rising

Also called Scalable Oversight

Supervising AI that is too capable or fast for a human to check directly — often by using AI to help.

Think of it like

A head editor who cannot read every article, so they train and trust junior editors, spot-checking to keep quality honest.

Example

To grade a model’s hard math proof, a human is assisted by a second model that critiques the first, so the human only has to judge the debate.

How it actually works

The core problem: as tasks exceed human ability to evaluate, naive human feedback stops being a reliable signal. Approaches include AI-assisted critique, debate between models, and recursive decomposition, all aiming to amplify limited human judgment. It connects to weak-to-strong generalization — can a weaker supervisor still steer a stronger model?

For product teams

The bet that lets us keep steering systems that outrun our ability to check them by hand.

For engineers

Amplify human evaluation via AI critique, debate, or recursive decomposition when direct assessment is infeasible.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome