Weak-to-strong generalization
Also called Weak-to-Strong Generalization
Whether a weaker teacher can steer a much stronger student to be even better than the teacher itself.
Think of it like
A grade-school coach who somehow trains an athlete that ends up outperforming anything the coach could ever do.
Example
Researchers fine-tune a large model using labels from a smaller, weaker model and find it generalizes past the weak labels’ mistakes.
How it actually works
This is a proxy for the real future problem: humans will be the "weak" supervisors of superhuman models. Early results suggest strong models can partly recover their full capability from imperfect weak supervision, but not fully, and the gap is the research target. It is a concrete experimental handle on scalable oversight.
For product teams
A hopeful sign that limited human feedback can still guide systems smarter than us — with caveats.
For engineers
Use weak-labeler supervision on a stronger model and measure recovered performance; the elicited-vs-ceiling gap is the metric.
Related
- Scalable oversight — A concrete testbed for scalable oversight.
- A live thread in alignment research.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome