MMLU
A benchmark of 15,000 multiple-choice questions across 57 domains—the gold standard for measuring broad knowledge.
Think of it like
Like the SAT, but way longer and covering everything from abstract algebra to US history.
Example
Question: 'Which of these is true about photosynthesis?' with four choices. The model's accuracy across all 15k questions is its MMLU score.
How it actually works
MMLU is valuable because it's diverse, hard for humans to memorize (you can't cheat on it), and well-calibrated to model capability—a 90% MMLU model is genuinely competent at broad knowledge. Downside: it's multiple-choice, so there's no penalty for reasoning poorly as long as you pick the right box. And leaked training data has inflated scores somewhat.
For product teams
MMLU is the quickest way to compare models. A 90% MMLU model is probably more capable than an 80% one, all else equal.
For engineers
Evaluate on MMLU regularly, but supplement with domain-specific benchmarks and human eval.
Related
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome