GPQA
A benchmark of 448 extremely hard multiple-choice questions in biology, physics, and chemistry—meant to be unsolvable by humans without domain expertise.
Think of it like
Like asking a literature major to answer PhD-level physics questions.
Example
A question about molecular biology that requires understanding multiple papers to answer correctly. Frontier models score ~60%, random is 25%.
How it actually works
GPQA is designed to resist gaming because the questions are Google-proof (searching won't help, you need actual knowledge). Most models score just above random. It's an emerging benchmark for detecting capability edge cases and whether models are truly reasoning or pattern-matching.
For product teams
GPQA is future-oriented; it highlights where frontier models still have gaps.
For engineers
Don't expect high scores on GPQA yet; use it to track progress on truly hard reasoning.
Related
- Benchmark — Category.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome