LiteLLM routes, budgets, and fallbacks across vendors. Put ValGuard closest to the model by setting api_base on the alias you want guarded so every completion passes your rule packs.
When you need this
LiteLLM retries on any error. A ValGuard block looks like a failed completion unless your fallback rule keys off status codes. Separate provider failures from policy blocks in router config.
Setup
- Create ValGuard agents and packs.
- Add a LiteLLM
model_listentry withapi_base: https://api.valguard.ai/v1andextra_headersforX-VG-Agent. - Point application code at the guarded alias only for flows that need enforcement.
- Shadow mode first on traffic mirrors.
Code
model_list:
- model_name: gpt-4o-mini-validated
litellm_params:
model: openai/gpt-4o-mini
api_base: https://api.valguard.ai/v1
api_key: vg_live_...
extra_headers:
X-VG-Agent: support-triage
import litellm
resp = litellm.completion(
model="gpt-4o-mini-validated",
messages=[{"role": "user", "content": "Hi"}],
)
Honest limits
- LiteLLM fallback logic runs around ValGuard responses. Document which HTTP codes mean policy block vs provider outage.
- LiteLLM guardrail plugins are separate from ValGuard. Do not assume both run unless you wire both.
- Model time still dominates latency.
Related
Next step
Quickstart and methodology.