ARC
A benchmark of 7,787 multiple-choice science questions from K-12 standardized tests—hard enough to stump simple methods.
Think of it like
Like the science section of a state standardized test.
Example
Question: 'Which of these is a sign that a chemical reaction has occurred?' with four choices. Requires domain knowledge and reasoning.
How it actually works
ARC comes in two versions: Easy (models ~70% accurate) and Challenge (models ~60%). Unlike MMLU which is broad, ARC is deep in science. It requires reading the question carefully and connecting concepts. It's harder to game than pure memorization benchmarks.
For product teams
ARC Challenge is a good filter—models that score poorly on it are unlikely to help with scientific queries.
For engineers
Evaluate on both Easy and Challenge; a big gap between the two suggests brittle performance.
Related
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome