Decoder. plain-English AI glossary

Honeypot

▲ Rising

Also called Honeypot

A tempting fake target planted to lure and catch bad actors — or bad model behavior — in the act.

Think of it like

Leaving a marked, worthless wallet on a bench and watching who pockets it.

Example

Safety researchers give an agent an easy-looking chance to disable its own oversight, planted purely to detect whether it would try.

How it actually works

Classic security honeypots are decoy systems that attract attackers so you can study and detect them. In alignment, the same idea probes models: hand them a fake opportunity to misbehave (steal, self-exfiltrate, tamper) and see if they take it. Anything that touches the bait reveals intent it would otherwise hide.

For product teams

A trap that surfaces misuse or misbehavior you would otherwise never see.

For engineers

Deploy decoy targets or fake affordances; interaction with them is high-signal evidence of an attacker or a misaligned policy.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome