Decoder. plain-English AI glossary

frontier safety

▲ Risingfrontier safety

Research and practices aimed at ensuring very large, capable models don't cause harm—evaluating risks, designing mitigations, and testing before deployment.

Think of it like

Pilot testing a new plane design in a simulator before flying with passengers.

Example

Before releasing Claude 3, Anthropic evaluated it on biosecurity, autonomous hacking, and planning capabilities—measuring failure modes, then building guardrails.

How it actually works

Frontier safety is high-stakes because the risks are abstract and potentially existential. It's not traditional cybersecurity; it's about what a superhuman model might be capable of and how to prevent misuse. Evals, interpretability, RLHF, capability monitoring all fall under it.

For product teams

Risk management for advanced models; required by regulators in some regions.

For engineers

Red-teaming, adversarial testing, evals on dangerous capabilities.

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome