ValGuard + LlamaIndex: gate LLM steps in pipelines

Configure the OpenAI-compatible LLM client in LlamaIndex with ValGuard as base URL so retrieval-augmented answers pass policy before tools run.

Last verified:

LlamaIndex pipelines call an LLM for synthesis, routing, and tool planning. Swap the OpenAI-compatible client to ValGuard so those completions pass deterministic rules before downstream tools run.

When you need this

A RAG answer cites the right shape JSON but invents a policy exception. Schema checks pass; business rules should not. ValGuard enforces packs on the completion before the query engine returns text to the user.

Setup

  1. Create a ValGuard agent for the synthesis step (and separate agents if routing uses another model).
  2. In LlamaIndex, set the OpenAI-compatible LLM api_base to https://api.valguard.ai/v1 and pass X-VG-Agent.
  3. Start in shadow mode on staging indexes.
  4. Enforce when would-block rates fit review capacity.

Code

import os
from llama_index.llms.openai import OpenAI

llm = OpenAI(
    model="openai/gpt-4o-mini",
    api_base="https://api.valguard.ai/v1",
    api_key=os.environ["VG_API_KEY"],
    default_headers={"X-VG-Agent": "rag-synthesis"},
)

Verify parameter names against your installed LlamaIndex version before production.

Honest limits

  • Retrieval quality and chunking are still your job. ValGuard does not prove citations are faithful without grounding rules you configure.
  • Non-LLM postprocessors are outside the proxy path.
  • See benchmark methodology for engine vs HTTP vs model latency.

Related

Next step

Quickstart and validate LLM output.