HellaSwag
A benchmark where models predict which of four video clip endings makes sense—measuring common-sense reasoning.
Think of it like
Like watching a video with the last few seconds cut off and guessing what happens next.
Example
Video shows someone vacuuming. Options: (A) they turn it off, (B) they eat it, (C) they paint with it, (D) they drive it. Correct answer is (A).
How it actually works
HellaSwag is designed to be easy for humans (~95% accuracy) but hard for models (~85% for frontier models). The hard wrong answers require understanding physical plausibility and social norms. It's multiple-choice so no credit for partial reasoning. Useful for distinguishing models in a narrow capability band.
For product teams
Use HellaSwag to differentiate models when they're close on other benchmarks.
For engineers
Don't over-index on HellaSwag alone; it's high-level reasoning, not factual knowledge or coding.
Related
- Benchmark — Category.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome