AI agent guardrails: fail-closed checks before the next step

Where to put runtime guardrails on agents, what deterministic rules catch, and how shadow mode fits a rollout.

Last verified:

AI agent guardrails are the checks that decide whether an agent may continue after a model turn or tool proposal. Without them, a schema-valid JSON object can still trigger a bad refund, a wrong CRM write, or an over-broad tool call.

ValGuard is a runtime validation and policy-enforcement layer for LLM calls and agent steps. Deterministic rules run on every completion before the next model, tool, or customer-facing action. It includes a lightweight step graph so rules apply at handoffs. It is not a general workflow engine, connector platform, or durable-execution system.

This guide is the pillar for "how do I put guardrails on AI agents?" Related: LLM output validation, MCP security, stack map.

What "guardrails" usually means

Teams overload the word. Separate four jobs:

JobTypical toolsHot-path?
Dialog / topic railsNeMo GuardrailsSometimes
In-process validatorsGuardrails AI, custom PythonYes
Cloud content filtersBedrock / AzureYes
Runtime policy gateValGuard, some gateway pluginsYes

If your risk is a wrong tool argument that spends money, you need a fail-closed gate with a rule ID. A topic filter alone is not enough. See vs cloud guardrails and vs Guardrails AI.

Where to put the gate

flowchart LR
  A[Agent / canvas] --> V[Validation gate]
  V --> M[Model]
  M --> V
  V -->|pass| T[Tool or business API]
  V -->|block| H[Escalate]

Put the gate on every completion that can authorize the next side effect. Skip only paths that never leave the sandbox.

Mechanism: OpenAI-compatible proxy

export VG_PROXY=https://api.valguard.ai
export VG_API_KEY=vg_live_...
export VG_AGENT=refund-guard

curl -s "$VG_PROXY/v1/chat/completions" \
  -H "Authorization: Bearer $VG_API_KEY" \
  -H "X-VG-Agent: $VG_AGENT" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"Refund 500 on order 12"}]}'

Wire frameworks from integrations: n8n, LangGraph, OpenAI Agents SDK.

Shadow mode before enforce

Run rules on live traffic without blocking. Tune false positives. Then enforce. Shadow mode is on every plan, including Free. Hosted shadow still sends requests to ValGuard. It is not offline evaluation. Trust.

What deterministic rules catch

  • Schema and enums the next node reads
  • Cross-field math (line totals, refund caps against trusted data)
  • PII and leak patterns you configured
  • Tool name allowlists and argument shapes when you gate tools

What they do not catch

  • Fluent false claims with no rule signature
  • Side effects that never pass through an LLM call
  • Taste and tone (keep a judge offline if you need one: vs LLM-as-judge)

Failure path

OutcomeAgent should
PassContinue to tool / next step
Re-askAccept ValGuard's retry once; treat final body as source of truth
BlockEscalate; attach rule IDs to the ticket
Unreachable proxyFail the request; any direct-provider fallback bypasses rules

Latency

Engine packs: microseconds. Proxy path: about 0.36 ms p50 with a mocked upstream. Model: hundreds of milliseconds. Block/re-ask buffers the stream. Methodology.

Build vs buy

If one Python service owns the path forever, a library may be enough. If policy must span services, buy a shared gate. See build vs buy.

Related

Next step

Quickstart. Start one agent in shadow mode on a high-risk handoff.