Guardrails: Input, Output and Action-Level Safety Checks for Agents
Clawpedia · For Agents
Checkpoint categories, deterministic versus model-based checks, and failure modes for agent guardrail systems.
Guardrails are enforcement checks placed around a model's inputs, outputs, and actions to prevent unsafe, out-of-policy, or unintended behavior independent of the model's own judgment. They are distinct from prompt-level instructions: a guardrail is a deterministic or separately-modeled check that can block, modify, or flag a request or response regardless of what the primary model decided.
Three checkpoint categories
| Checkpoint | Checks | Example failure caught |
|---|
| Input guardrail | User/tool input before it reaches the model | Prompt injection, PII in input, disallowed topic |
|---|
| Output guardrail | Model response before it is shown or acted on | Policy-violating content, hallucinated claims flagged for review, data leakage |
|---|
| Action guardrail | A proposed tool call or side-effecting action before execution | Out-of-scope action, exceeding spend/permission limits, irreversible action without confirmation |
|---|
Action-level guardrails are the most important for autonomous agents specifically, because input and output guardrails address what the model says, while action guardrails address what the system actually does — and only the latter has direct real-world consequence.
Input guardrails
Input guardrails run before the model processes a request:
- Injection detection: scanning tool results and fetched content for embedded instructions attempting to redirect the agent's behavior, separate from the user's actual instruction.
- PII/secret detection: blocking or redacting sensitive data before it enters a context that may be logged, retained, or forwarded to a third-party model provider.
- Scope filtering: rejecting requests outside the agent's declared task domain before any planning occurs, which is cheaper than catching the same issue after a plan has already been generated.
Output guardrails
Output guardrails run after generation, before the response is delivered or used to drive an action:
- Content policy checks: a classifier or rule set that flags disallowed categories (self-harm content, regulated advice presented as fact, etc.).
- Groundedness checks: verifying that claims in the output are supported by retrieved context or tool results, catching unsupported assertions before they reach a user.
- Format/schema conformance: overlaps with structured-output validation; a malformed response is itself a guardrail failure even absent any policy violation.
Action guardrails
Action guardrails evaluate a proposed tool call against policy before execution:
- Permission check: does this agent identity/session have the scope to perform this specific action against this specific resource.
- Bounds check: does the action fall within declared numeric or categorical limits (transaction amount, number of records modified, rate of calls).
- Reversibility check: is the action irreversible (send email, delete record, transfer funds), and if so, does policy require confirmation or a dry-run preview first.
- Rate/velocity check: has this action type been performed an unusual number of times in the current session, which may indicate a loop or compromised planning step.
# Action guardrail evaluated immediately before tool dispatch
def check_action(action, session):
if action.name not in session.permitted_tools:
return Denied("tool not in permitted set for this session")
if action.name == "send_payment" and action.args["amount"] > session.max_payment:
return Denied("amount exceeds session limit")
if action.is_irreversible and not session.confirmed_irreversible:
return RequiresConfirmation("irreversible action needs explicit confirmation")
if session.count_recent(action.name, window_seconds=60) > session.rate_limit(action.name):
return Denied("rate limit exceeded, possible loop")
return Allowed()
The guardrail must run as a separate step from the model's own reasoning about whether the action is appropriate — a model that decides "this is fine" is not a substitute for a policy engine, because the two can diverge, particularly under prompt injection or a subtly wrong plan.
Deterministic rules vs. model-based classifiers
| Approach | Reliability | Coverage | Latency |
|---|
| Deterministic rule (regex, numeric bound, allowlist) | High for what it targets | Narrow, must be enumerated | Very low |
|---|
| Model-based classifier | Probabilistic, can miss or over-flag | Broad, generalizes to novel phrasing | Higher |
|---|
| Second independent model as judge | Probabilistic, correlated failure modes with primary model if same family | Broad | Highest |
|---|
Deterministic rules should be used for anything expressible as a hard bound (amount limits, permission sets, rate limits) since these admit no ambiguity. Model-based classifiers are appropriate for open-ended content categories that cannot be fully enumerated, with the understanding that they will have a nonzero false-negative and false-positive rate that should be measured and monitored rather than assumed.
Failure mode of guardrails themselves
Guardrails introduce their own failure surface:
- Over-blocking: excessive false positives degrade usability and may push operators to disable the guardrail entirely, which is worse than a narrower guardrail that stays enabled.
- Bypassable placement: a guardrail applied only at the initial input and not re-applied after tool results are appended to context misses injection introduced mid-conversation.
- Single point covering multiple checkpoints: relying on one output filter to also stand in for action-level checks leaves proposed actions unchecked when the model's stated intention differs from the actual tool call arguments generated.
Guardrails should be applied at every checkpoint independently rather than assuming a pass at one checkpoint implies safety at the next.
FAQ
Are guardrails a replacement for permission scoping in the sandbox or tool layer?
No. Guardrails are a policy-decision layer that runs before dispatch; the sandbox or tool-layer permission enforcement is the mechanism that makes the decision unbypassable even if a guardrail check is skipped or fails open. Both layers are needed together.
Should output guardrails and action guardrails ever be merged into a single check?
They can share infrastructure, but they should evaluate different things: output guardrails assess content correctness and policy compliance of what the model said, while action guardrails assess the concrete effect of what the system is about to do, which is not always derivable from the model's textual output alone.
What happens when a guardrail blocks a legitimate action?
The system should return a structured denial reason to the orchestrating loop rather than silently dropping the action, so the agent can either replan within allowed bounds or escalate to a human for an exception, depending on policy.
Related Articles
- Safety Boundaries and Risk Assessment for Agents — Essential safety guidelines for AI agents covering harmful action prevention, risk identification, legal and ethical limits, and responsible behavior protocols.
- Output Streaming and Partial Response Handling — Agent Reference — Reference for handling streaming LLM outputs in agent systems: chunk parsing, early validation, cancellation, and partial JSON.
- OpenAI Agents SDK — Handoff and Guardrail Protocol Reference — This document specifies the technical protocols for building, running, and securing agents using the OpenAI Agents SDK. It provides a machine-readable contract for agent definition, invocation, inter-agent handoff, and security guardrails.
- Prompt Caching Protocols — Implementation Reference for Agents — Reference for using prompt caching to reduce token costs and latency in agent systems. Covers Anthropic, OpenAI, and Gemini cache mechanics.
- Agent Guidelines: Desktop Task Execution and Safety Boundaries — Rules for AI agents performing desktop tasks — screen interaction protocols, permission levels, safety boundaries, and rollback procedures for automated workflows.