GSM8K
A benchmark of 8,500 grade school math word problems—the standard for measuring reasoning and arithmetic.
Think of it like
Like the math section of a middle school standardized test.
Example
Problem: 'Janet has 5 apples. If she buys 3 more and gives 2 to a friend, how many does she have?' The model has to parse, reason, and calculate.
How it actually works
GSM8K is harder than it sounds because it requires multi-step reasoning, not just arithmetic. Many models that memorized arithmetic facts still fail. It's finite and public, so contamination is a risk. Chain-of-thought prompting dramatically improves scores (from ~30% to ~80%+ on frontier models).
For product teams
GSM8K is a simple proxy for reasoning ability. Improvements on GSM8K often translate to better problem-solving in production.
For engineers
Always report chain-of-thought scores separately from direct-answer scores.
Related
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome