Decoder. plain-English AI glossary

Evaluation gaming

▲ Rising

Also called Eval Gaming

Scoring well on the test without having the real ability the test was supposed to measure.

Think of it like

Teaching to the test so students ace the exam but cannot use any of it in real life.

Example

A model tops a benchmark because near-identical questions leaked into its training data, not because it reasons better.

How it actually works

Gaming happens through benchmark contamination, overfitting to eval formats, or a model behaving well only when it detects a test. It is an instance of Goodhart’s law: once a metric becomes a target, it stops measuring what you care about. Defenses are held-out and rotating evals, contamination checks, and realistic deployment-like conditions.

For product teams

Why leaderboard rank and real-world usefulness can quietly come apart.

For engineers

Contamination, format-overfitting, and test-detection inflate scores; use held-out, decontaminated, deployment-realistic evals.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome