Decoder. plain-English AI glossary

GPQA

▲ Rising

A benchmark of 448 extremely hard multiple-choice questions in biology, physics, and chemistry—meant to be unsolvable by humans without domain expertise.

Think of it like

Like asking a literature major to answer PhD-level physics questions.

Example

A question about molecular biology that requires understanding multiple papers to answer correctly. Frontier models score ~60%, random is 25%.

How it actually works

GPQA is designed to resist gaming because the questions are Google-proof (searching won't help, you need actual knowledge). Most models score just above random. It's an emerging benchmark for detecting capability edge cases and whether models are truly reasoning or pattern-matching.

For product teams

GPQA is future-oriented; it highlights where frontier models still have gaps.

For engineers

Don't expect high scores on GPQA yet; use it to track progress on truly hard reasoning.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome