Scalable oversight
Also called Scalable Oversight
Supervising AI that is too capable or fast for a human to check directly — often by using AI to help.
Think of it like
A head editor who cannot read every article, so they train and trust junior editors, spot-checking to keep quality honest.
Example
To grade a model’s hard math proof, a human is assisted by a second model that critiques the first, so the human only has to judge the debate.
How it actually works
The core problem: as tasks exceed human ability to evaluate, naive human feedback stops being a reliable signal. Approaches include AI-assisted critique, debate between models, and recursive decomposition, all aiming to amplify limited human judgment. It connects to weak-to-strong generalization — can a weaker supervisor still steer a stronger model?
For product teams
The bet that lets us keep steering systems that outrun our ability to check them by hand.
For engineers
Amplify human evaluation via AI critique, debate, or recursive decomposition when direct assessment is infeasible.
Related
- Oversight — The broader practice it extends is oversight.
- Weak-to-strong generalization — Tightly linked to weak-to-strong generalization.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome