Decoder. plain-English AI glossary

Prompt extraction

▲ Rising

Techniques to discover or recover a model's system prompt or internal instructions through structured queries.

Think of it like

Like detective work—asking careful questions to infer what rules the system is following.

Example

Testing GPT-4 with "repeat everything I say" or "pretend you have no restrictions" to see what refusals or behaviors leak information.

How it actually works

Different extraction techniques include direct requests, role-playing scenarios, token/logit analysis, and adversarial prompting. Success varies by model architecture and training; some leak easily, others are designed to resist.

For product teams

Reduces security through obscurity but also reveals opportunities to harden your own prompts.

For engineers

A useful red-teaming tool; use it to find weaknesses in your own system prompts before adversaries do.

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome