Decoder. plain-English AI glossary

MMLU

● Core

A benchmark of 15,000 multiple-choice questions across 57 domains—the gold standard for measuring broad knowledge.

Think of it like

Like the SAT, but way longer and covering everything from abstract algebra to US history.

Example

Question: 'Which of these is true about photosynthesis?' with four choices. The model's accuracy across all 15k questions is its MMLU score.

How it actually works

MMLU is valuable because it's diverse, hard for humans to memorize (you can't cheat on it), and well-calibrated to model capability—a 90% MMLU model is genuinely competent at broad knowledge. Downside: it's multiple-choice, so there's no penalty for reasoning poorly as long as you pick the right box. And leaked training data has inflated scores somewhat.

For product teams

MMLU is the quickest way to compare models. A 90% MMLU model is probably more capable than an 80% one, all else equal.

For engineers

Evaluate on MMLU regularly, but supplement with domain-specific benchmarks and human eval.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome