AI Safety
Also called AI Safety
The field working to make AI systems behave as intended and avoid causing harm.
Think of it like
Seatbelts, crumple zones, and driving tests for cars — the whole apparatus that lets a powerful machine be used safely.
Example
A lab’s safety work spans refusals for dangerous requests, red teaming before launch, and research into whether a model’s goals match its designers’.
How it actually works
AI safety is a broad umbrella, from the practical to the speculative. On the near end: content filters, robustness to jailbreaks, bias, misuse. On the far end: alignment of highly capable systems whose objectives might diverge from ours. The unifying question is making sure systems do what we actually want as they grow more capable. People disagree sharply on which risks deserve the most attention, and that debate is part of the field.
For product teams
The discipline that decides whether an AI product is trustworthy enough to put in front of users at all.
For engineers
Spans robustness, misuse prevention, and alignment; the practical near-term work and the longer-term capability-scaling concerns share the goal of intended behavior.
Related
- Alignment — The core technical sub-problem.
- Existential Risk — The far-tail concern it debates.
- Interpretability — Understanding models to make them safer.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome