Decoder. plain-English AI glossary

AI safety

● Core

The field of making sure AI systems don't cause harm — from today's chatbot glitches to tomorrow's autonomous risks.

Think of it like

Automotive safety engineering — not just building a fast car, but making sure it has brakes, seatbelts, and crash testing.

Example

Testing a model for bias, building kill switches for agents, and researching alignment are all AI safety work.

How it actually works

AI safety spans a spectrum: near-term (bias, toxicity, privacy, misuse) and long-term (loss of control, deceptive alignment, power concentration). The near-term work is engineering — evals, guardrails, audits. The long-term work is more research-flavored and contested.

For product teams

The discipline of preventing AI harm — from bias audits to existential risk research.

For engineers

Near-term: bias, toxicity, misuse mitigation. Long-term: alignment robustness, control, corrigibility.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome