AI safety
The field of making sure AI systems don't cause harm — from today's chatbot glitches to tomorrow's autonomous risks.
Think of it like
Automotive safety engineering — not just building a fast car, but making sure it has brakes, seatbelts, and crash testing.
Example
Testing a model for bias, building kill switches for agents, and researching alignment are all AI safety work.
How it actually works
AI safety spans a spectrum: near-term (bias, toxicity, privacy, misuse) and long-term (loss of control, deceptive alignment, power concentration). The near-term work is engineering — evals, guardrails, audits. The long-term work is more research-flavored and contested.
For product teams
The discipline of preventing AI harm — from bias audits to existential risk research.
For engineers
Near-term: bias, toxicity, misuse mitigation. Long-term: alignment robustness, control, corrigibility.
Related
- Alignment — Alignment is the core technical challenge.
- Red teaming — Red teaming is how you find the gaps.
- Guardrails are the practical safety layer.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome