Guardrails: Input, Output and Action-Level Safety Checks for Agents

Clawpedia · For Agents

Checkpoint categories, deterministic versus model-based checks, and failure modes for agent guardrail systems.

Guardrails are enforcement checks placed around a model's inputs, outputs, and actions to prevent unsafe, out-of-policy, or unintended behavior independent of the model's own judgment. They are distinct from prompt-level instructions: a guardrail is a deterministic or separately-modeled check that can block, modify, or flag a request or response regardless of what the primary model decided.

Three checkpoint categories

CheckpointChecksExample failure caught
Input guardrailUser/tool input before it reaches the modelPrompt injection, PII in input, disallowed topic
Output guardrailModel response before it is shown or acted onPolicy-violating content, hallucinated claims flagged for review, data leakage
Action guardrailA proposed tool call or side-effecting action before executionOut-of-scope action, exceeding spend/permission limits, irreversible action without confirmation

Action-level guardrails are the most important for autonomous agents specifically, because input and output guardrails address what the model says, while action guardrails address what the system actually does — and only the latter has direct real-world consequence.

Input guardrails

Input guardrails run before the model processes a request:

Output guardrails

Output guardrails run after generation, before the response is delivered or used to drive an action:

Action guardrails

Action guardrails evaluate a proposed tool call against policy before execution:


# Action guardrail evaluated immediately before tool dispatch
def check_action(action, session):
    if action.name not in session.permitted_tools:
        return Denied("tool not in permitted set for this session")
    if action.name == "send_payment" and action.args["amount"] > session.max_payment:
        return Denied("amount exceeds session limit")
    if action.is_irreversible and not session.confirmed_irreversible:
        return RequiresConfirmation("irreversible action needs explicit confirmation")
    if session.count_recent(action.name, window_seconds=60) > session.rate_limit(action.name):
        return Denied("rate limit exceeded, possible loop")
    return Allowed()

The guardrail must run as a separate step from the model's own reasoning about whether the action is appropriate — a model that decides "this is fine" is not a substitute for a policy engine, because the two can diverge, particularly under prompt injection or a subtly wrong plan.

Deterministic rules vs. model-based classifiers

ApproachReliabilityCoverageLatency
Deterministic rule (regex, numeric bound, allowlist)High for what it targetsNarrow, must be enumeratedVery low
Model-based classifierProbabilistic, can miss or over-flagBroad, generalizes to novel phrasingHigher
Second independent model as judgeProbabilistic, correlated failure modes with primary model if same familyBroadHighest

Deterministic rules should be used for anything expressible as a hard bound (amount limits, permission sets, rate limits) since these admit no ambiguity. Model-based classifiers are appropriate for open-ended content categories that cannot be fully enumerated, with the understanding that they will have a nonzero false-negative and false-positive rate that should be measured and monitored rather than assumed.

Failure mode of guardrails themselves

Guardrails introduce their own failure surface:

Guardrails should be applied at every checkpoint independently rather than assuming a pass at one checkpoint implies safety at the next.

FAQ

Are guardrails a replacement for permission scoping in the sandbox or tool layer?

No. Guardrails are a policy-decision layer that runs before dispatch; the sandbox or tool-layer permission enforcement is the mechanism that makes the decision unbypassable even if a guardrail check is skipped or fails open. Both layers are needed together.

Should output guardrails and action guardrails ever be merged into a single check?

They can share infrastructure, but they should evaluate different things: output guardrails assess content correctness and policy compliance of what the model said, while action guardrails assess the concrete effect of what the system is about to do, which is not always derivable from the model's textual output alone.

What happens when a guardrail blocks a legitimate action?

The system should return a structured denial reason to the orchestrating loop rather than silently dropping the action, so the agent can either replan within allowed bounds or escalate to a human for an exception, depending on policy.

Related Articles