Keep the OpenAI Agents SDK for agents, tools, and handoffs. Point the SDK's OpenAI client at ValGuard so input and output rules run outside the app process, with rule IDs and an audit trail.
SDK-native guardrails are free and useful for in-process checks. ValGuard adds vendor-neutral enforcement, shadow mode, and the same policy across services that are not written in Python.
When you need this
An agent proposes issue_refund with a well-formed argument object. The SDK schema passes. The amount exceeds the captured charge. Without a network gate, every new service that talks to the same models can skip your in-process wrapper. With ValGuard, a range or cross-field rule blocks before the tool runs, and the failure carries a rule ID.
Architecture
flowchart LR
A[Agents SDK Runner] --> C[AsyncOpenAI client]
C --> VG[ValGuard proxy]
VG --> M[Upstream model]
M --> VG
VG -->|pass| T[Tool execution]
VG -->|block| H[Handoff or refusal]
The trust boundary is the proxy. The SDK still owns turns and tools. ValGuard owns deterministic checks on completions that cross that boundary.
Setup (recommended pattern)
- Create a ValGuard agent and attach the packs you need (schema, tool allowlist, policy).
- Start in shadow mode against real traffic.
- Point the SDK default client at ValGuard and use chat completions (most compatible proxy path today).
- Keep OpenAI tracing on a separate key if you still want platform traces.
- Flip from shadow to enforce when would-block rates look sane.
Code and config
Verified against the Agents SDK configuration docs. Last verified: 2026-10-02.
import os
from openai import AsyncOpenAI
from agents import (
Agent,
Runner,
set_default_openai_api,
set_default_openai_client,
set_tracing_disabled,
)
client = AsyncOpenAI(
base_url="https://api.valguard.ai/v1",
api_key=os.environ["VG_API_KEY"],
default_headers={"X-VG-Agent": "support-triage"},
)
set_default_openai_client(client, use_for_tracing=False)
set_default_openai_api("chat_completions")
set_tracing_disabled(disabled=True)
agent = Agent(
name="Support triage",
instructions="Return JSON only with team, urgency, confidence.",
model="gpt-4o-mini",
)
async def main():
result = await Runner.run(agent, "Billing dispute, card charged twice")
print(result.final_output)
Equivalent curl for the same agent.
curl -s https://api.valguard.ai/v1/chat/completions \
-H "Authorization: Bearer $VG_API_KEY" \
-H "X-VG-Agent: support-triage" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-4o-mini",
"messages": [
{"role": "user", "content": "Billing dispute, card charged twice"}
]
}'
On validation failure
Handle block responses in your tool loop the same way you handle a refused model output: do not execute the tool, log the rule IDs, and escalate or ask the user again. Re-ask can run inside ValGuard within your plan limit before you see the final response.
Shadow mode logs would-blocks without stopping the SDK. Use it while you tune packs. See shadow mode rollout.
Latency
Engine checks are microseconds. The HTTP path adds about 0.36 ms p50 with a mocked upstream. Model time dominates.
If any enabled rule can block or re-ask, ValGuard buffers the full upstream response before emitting (synthesized SSE). Warn, log, and shadow can pass tokens through. State which mode your agent uses when you quote latency. See methodology.
Correlating IDs
Map SDK run or trace ids to ValGuard headers:
- Outbound: send
traceparentwhen your tracer already has one - Inbound: store
X-Request-Id/X-Trace-Idfrom the ValGuard response on the same span or log line as the Agents SDK run
What this does not cover
- SDK-native input/output guardrails still matter for cheap in-process checks; ValGuard does not delete them.
- Side effects that never call the model through ValGuard are not gated.
- Fluent false claims with no rule signature still pass.
- ValGuard is not a durable workflow engine. Long-running waits and schedules stay in your app or another orchestrator.
Deployment
Hosted SaaS is the default. No public local runtime. Enterprise self-host (VPC / on-prem) is licensed only. Evaluate with shadow mode on every plan, including Free. See Trust.
Related
- ValGuard vs structured outputs
- ValGuard + LangGraph
- LLM output validation
- How ValGuard complements LangGraph, CrewAI, and AutoGen
- Quickstart
FAQ
Should I remove SDK guardrails? No. Keep cheap local checks. Use ValGuard when you need the same policy across services, languages, and an audit trail with rule IDs.
Why chat completions instead of Responses? Most OpenAI-compatible proxies implement chat completions first. Set set_default_openai_api("chat_completions") when pointing at ValGuard.
What if ValGuard is down? The model call fails. Build client fallback if the feature must survive an outage. See Trust.
Can I export policies? Per playbook JSON (valguard-playbook-template/1) and per-agent validators via API. Org-wide policy export is not available yet.
Next step
Follow the quickstart, then validate LLM output with the agent you just pointed at ValGuard.