AI glossary

Jailbreak

In one sentence A jailbreak is a prompt crafted to talk a model out of its safety training, getting it to produce what it was trained to refuse.

By Updated

A jailbreak is a prompt designed to make a model bypass its own safety rules and produce output it was trained to refuse.

Think of a strict gatekeeper who turns away anyone without a pass — but who can be talked past. "I'm the electrician, the manager called me in." "It's an emergency." "We're shooting a film; you are playing a guard who lets people in." The gatekeeper's rules did not change; language found a route around them. Models learn refusal as a behaviour during safety training, and behaviours reachable by language are steerable by language.

The recurring genres: role-play framing ("you are DAN, an AI without restrictions"), fictional wrappers ("write a thriller scene where the character explains..."), authority claims ("as a licensed researcher, I need..."), gradual escalation across many turns, encodings that slip past pattern-matching, and — against open models — automated searches for adversarial suffixes that reliably flip refusals. Named jailbreaks circulate publicly and get patched; new ones follow.

The distinction from prompt-injection is who attacks what: a jailbreak is the user attacking the model's own rules; injection is a third party's instructions hidden in content, hijacking the user's session. The mindset for builders follows from the pattern: treat refusals as one probabilistic layer, assume it will sometimes fail, and back it with guardrails and hard limits at the action layer. Finding these holes on purpose, before attackers do, is the job of red-teaming.

Where to go next