LlamaIndex pipelines call an LLM for synthesis, routing, and tool planning. Swap the OpenAI-compatible client to ValGuard so those completions pass deterministic rules before downstream tools run.
When you need this
A RAG answer cites the right shape JSON but invents a policy exception. Schema checks pass; business rules should not. ValGuard enforces packs on the completion before the query engine returns text to the user.
Setup
- Create a ValGuard agent for the synthesis step (and separate agents if routing uses another model).
- In LlamaIndex, set the OpenAI-compatible LLM
api_basetohttps://api.valguard.ai/v1and passX-VG-Agent. - Start in shadow mode on staging indexes.
- Enforce when would-block rates fit review capacity.
Code
import os
from llama_index.llms.openai import OpenAI
llm = OpenAI(
model="openai/gpt-4o-mini",
api_base="https://api.valguard.ai/v1",
api_key=os.environ["VG_API_KEY"],
default_headers={"X-VG-Agent": "rag-synthesis"},
)
Verify parameter names against your installed LlamaIndex version before production.
Honest limits
- Retrieval quality and chunking are still your job. ValGuard does not prove citations are faithful without grounding rules you configure.
- Non-LLM postprocessors are outside the proxy path.
- See benchmark methodology for engine vs HTTP vs model latency.