AI glossary

Guardrails

In one sentence Guardrails are the checks wrapped around a model that block bad inputs and outputs — policy enforced outside the model itself.

By Updated

Guardrails are external checks around an LLM — filters on what goes in and what comes out — that enforce rules the model itself might fail to follow.

A well-trained bank teller still works behind toughened glass, with cash limits and an alarm button. Not because training failed, but because training alone is never the whole security story: institutions layer physical controls over human judgement. Guardrails are that layer for models — alignment shapes what the model tends to do; guardrails constrain what the system can do regardless.

They sit at defined points in the request path:

input checks   : block prompt-injection patterns, off-topic requests,
                 personal data before it reaches the model
output checks  : scan the reply for policy violations, leaked secrets,
                 unsupported claims (compare against RAG sources), bad formats
action checks  : the hard boundary for agents — allowlists of callable tools,
                 spending caps, human sign-off for irreversible steps

Implementations range from cheap to heavy: regex and keyword rules, a small classifier (Llama Guard is a purpose-trained one), a second LLM judging the first's output, and schema enforcement via structured-output. Frameworks like NeMo Guardrails package the pattern.

Two engineering truths keep the concept honest. Text-level filters are probabilistic — determined attackers find phrasings past them, as jailbreak research keeps demonstrating — so real security lives at the action layer, where a blocked capability cannot be talked into existence. And every check adds latency and false refusals, so guardrail strictness is a product decision, tested with evals, not a switch set once.

Where to go next