AI & Data › LLM & AI Engineering
Guardrails
Checks on model inputs and outputs for safety and correctness.
Also known as: guardrails, LLM guardrails, safety layer
Guardrails are the checks around a language model that constrain what goes in and what comes out: filtering unsafe or off-topic input, validating output format and content, blocking disallowed actions, and routing uncertain cases to people. They are a layered system, because no single check catches everything.
input checks → model → output checks (schema, content, policy) → action gate → user
Guardrails can be rule-based (patterns, allowlists, schema checks), model-based (a classifier or second model judging content), or procedural (required approvals). Rule-based checks are predictable and cheap; model-based checks catch more nuance but add cost and their own errors.
The classic mistakes:
- A single filter as the whole defence. Any one check has blind spots. Layer checks at input, output and action boundaries.
- Guardrails that rely on the model behaving. Asking the model to police itself is the weakest layer. Enforce critical rules in code.
- Over-blocking. Strict filters reject legitimate requests and frustrate users. Measure false positives as carefully as misses.
- Unmeasured coverage. Teams often do not know what their guardrails miss. Build an adversarial test set and track results.
- Ignoring output encoding. Model output rendered as markup or executed as code is a separate risk. Encode it for its destination (see output encoding).
Practice: define the policies, enforce the critical ones deterministically, add model-based checks where rules fall short, and red-team the whole stack regularly.
Review the guardrail logs on a schedule. Blocked requests show what users actually attempt, and allowed requests that were later reported show what the checks missed. Both feed directly into the next round of rules and tests.