Vapi assistants call a custom LLM URL for reasoning while Vapi handles telephony and speech. Point that URL at ValGuard so voice workflows reuse the same validators as chat — with realistic latency expectations.
When you need this
A phone agent confirms a medical appointment the clinic never approved. The transcript sounds polite. ValGuard blocks approval language until structured status allows it — but block rules buffer the LLM reply, which adds dead air. Budget UX accordingly.
Setup
- Create a ValGuard agent for the voice assistant role.
- In Vapi, set the custom LLM endpoint to
https://api.valguard.ai/v1/chat/completions(verify exact Vapi field names in their docs). - Send ValGuard API key and
X-VG-Agentper assistant. - Start in shadow mode on test numbers.
Code
Treat Vapi as an OpenAI-compatible client:
# Equivalent to what Vapi sends when configured with ValGuard as custom LLM
curl -s https://api.valguard.ai/v1/chat/completions \
-H "Authorization: Bearer $VG_API_KEY" \
-H "X-VG-Agent: vapi-support" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"Caller asks to cancel policy 8821"}]}'
Honest limits
- Speech-to-text and text-to-speech quality are outside ValGuard.
- Real-time voice needs block rules only where buffering is acceptable, or use warn/log plus human review queues.
- Telephony compliance (recording consent, PCI over voice) is still your program.
Related
Next step
Quickstart and shadow mode.