Decoder. plain-English AI glossary

ARC

● Core

A benchmark of 7,787 multiple-choice science questions from K-12 standardized tests—hard enough to stump simple methods.

Think of it like

Like the science section of a state standardized test.

Example

Question: 'Which of these is a sign that a chemical reaction has occurred?' with four choices. Requires domain knowledge and reasoning.

How it actually works

ARC comes in two versions: Easy (models ~70% accurate) and Challenge (models ~60%). Unlike MMLU which is broad, ARC is deep in science. It requires reading the question carefully and connecting concepts. It's harder to game than pure memorization benchmarks.

For product teams

ARC Challenge is a good filter—models that score poorly on it are unlikely to help with scientific queries.

For engineers

Evaluate on both Easy and Challenge; a big gap between the two suggests brittle performance.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome