Decoder. plain-English AI glossary

HellaSwag

● Core

A benchmark where models predict which of four video clip endings makes sense—measuring common-sense reasoning.

Think of it like

Like watching a video with the last few seconds cut off and guessing what happens next.

Example

Video shows someone vacuuming. Options: (A) they turn it off, (B) they eat it, (C) they paint with it, (D) they drive it. Correct answer is (A).

How it actually works

HellaSwag is designed to be easy for humans (~95% accuracy) but hard for models (~85% for frontier models). The hard wrong answers require understanding physical plausibility and social norms. It's multiple-choice so no credit for partial reasoning. Useful for distinguishing models in a narrow capability band.

For product teams

Use HellaSwag to differentiate models when they're close on other benchmarks.

For engineers

Don't over-index on HellaSwag alone; it's high-level reasoning, not factual knowledge or coding.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome