Oversight
Also called Human Oversight
Keeping humans meaningfully in the loop to catch, judge, and correct what a model does.
Think of it like
A flight instructor with their own set of controls, ready to take over the moment things go sideways.
Example
Before an AI agent can wire money or delete a database, its plan is routed to a person who has to approve it.
How it actually works
Oversight covers everything from a human approving individual actions to reviewing samples and setting policy. The catch is that it only works while humans can actually understand and evaluate what the model is doing — which gets harder as systems get more capable. That limit is exactly what scalable oversight tries to push past.
For product teams
The practical control surface — where a human can still say no before harm lands.
For engineers
Insert human judgment at high-risk decision points; effectiveness is bounded by human ability to evaluate model outputs.
Related
- Scalable oversight — Its harder, capability-scaling form is scalable oversight.
- HITL — Overlaps with keeping a human in the loop.
- Corrigibility — Depends on the system being corrigible.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome