Decoder. plain-English AI glossary

AI Safety

● Core

Also called AI Safety

The field working to make AI systems behave as intended and avoid causing harm.

Think of it like

Seatbelts, crumple zones, and driving tests for cars — the whole apparatus that lets a powerful machine be used safely.

Example

A lab’s safety work spans refusals for dangerous requests, red teaming before launch, and research into whether a model’s goals match its designers’.

How it actually works

AI safety is a broad umbrella, from the practical to the speculative. On the near end: content filters, robustness to jailbreaks, bias, misuse. On the far end: alignment of highly capable systems whose objectives might diverge from ours. The unifying question is making sure systems do what we actually want as they grow more capable. People disagree sharply on which risks deserve the most attention, and that debate is part of the field.

For product teams

The discipline that decides whether an AI product is trustworthy enough to put in front of users at all.

For engineers

Spans robustness, misuse prevention, and alignment; the practical near-term work and the longer-term capability-scaling concerns share the goal of intended behavior.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome