An LLM-as-judge scores tone, taste, and fuzzy quality. ValGuard applies deterministic rules with rule IDs before side effects. Keep the judge for language. Put rules in front of money, PHI, and tools.
ValGuard is a runtime validation and policy-enforcement layer for LLM calls and agent steps. It is not an evaluation platform. Evals measure models offline; validation enforces a contract at runtime. See evals vs validation.
Quick answer
Choose a judge for soft questions such as voice and tone. Choose ValGuard for hard checks such as totals and enums. Use both in a sandwich: model, optional judge fields, deterministic gate, then action.
Verdict
Judges are the right tool for nuance; they are the wrong sole authority before irreversible actions.
What each is built for
LLM-as-judge is a second model (or the same model with a rubric prompt) that emits a score or rationale. It tracks when prompts, judge models, or rubrics change. Reproducibility is hard unless you freeze all three.
ValGuard runs named rules on every completion: schema, enums, arithmetic, policy packs, grounding overlap. Same input yields the same pass/block. Shadow mode measures would-blocks before you enforce.
Comparison table
| Capability | LLM-as-judge | ValGuard | Who wins |
|---|---|---|---|
| Tone / taste / fuzzy quality | Strong | Weak by design | Judge |
| Exact arithmetic / enums | Unreliable | Deterministic | ValGuard |
| Latency | Extra model call (100s of ms) | Engine µs + ~0.41 ms HTTP p50 | ValGuard for the gate |
| Cost | Tokens every time | No judge tokens for the rule | ValGuard |
| Reproducibility | Rubric + model drift | Rule ID stable | ValGuard |
| Audit story | Score + prose | Pass/block per rule ID | ValGuard for compliance |
| Catching fluent false claims | Sometimes | Only with a rule signature | Neither alone |
Where ValGuard is stronger
- Same verdict tomorrow after you bump the model.
- Microsecond packs for hot-path checks; see benchmarks.
- Block, re-ask, warn, log with explicit on-fail behavior.
- Shadow mode on Free and every paid plan.
Where LLM-as-judge is stronger
- Voice, empathy, and open-ended policy gray areas.
- Ranking many candidates when no crisp rule exists.
- Research and offline eval loops (paired with datasets).
- Producing a
risk_scorefield that rules can threshold.
Cost and latency
A judge adds a full model call. ValGuard's gate does not. Engine packs are microseconds; HTTP about 0.36 ms p50 with a mocked upstream. Quote all three layers. Streaming: block/re-ask buffer; warn/log/shadow can pass through. Methodology.
Failure example
Without a gate: a judge scores a prior-auth summary 0.82 "looks fine." The customer-facing text says the MRI is approved. No structural check ran. The wrong message ships.
In a voice call center, a judge rates empathy 0.9 while the wrap-up JSON sets resolution: "refund_approved" for an out-of-policy case. The CRM webhook fires on tone, not on policy.
With ValGuard: keep the judge score as a field. A policy pack blocks approval language until a human or structured status allows it. The judge never authorizes the send alone.
Code
Judge payload as data ValGuard can check:
{
"wrap_up_code": "billing_followup",
"summary": "Customer asked about a duplicate charge.",
"risk_score": 0.42,
"tone_ok": true
}
curl -s https://api.valguard.ai/v1/chat/completions \
-H "Authorization: Bearer $VG_API_KEY" \
-H "X-VG-Agent: wrap-up-guard" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"Validate wrap-up JSON"}]}'
Rules on that object stay boring: enum on wrap_up_code, range on risk_score, escalate when risk_score >= 0.8 even if tone_ok is true.
Choose LLM-as-judge if
- The decision is primarily language quality
- You already freeze rubrics and models for audits
- Latency and token cost of a second call are acceptable
- No crisp numeric or schema contract exists yet
Choose ValGuard if
- A wrong enum or total causes real damage
- You need identical verdicts across services
- You must show rule IDs under review
- You want shadow mode before blocking customers
Using both
Put the judge inside the graph. Schema-check its JSON. Threshold the score with a rule. Act only on validation pass. Integrations: LangGraph, OpenAI Agents.
Objections
- Why not only Pydantic? Shape-check the judge output, then still enforce business limits. vs structured outputs.
- Latency. The judge is the expensive line item; the gate is not.
- ValGuard down. Call fails; no silent bypass. Trust.
- Data residency. SaaS default; Enterprise self-host for VPC.
- False positives. Shadow, tune, enforce.
- Misses. No rule means no catch; judges miss too, differently.
- Exit. Playbook and per-agent export; no org-wide policy file yet.
FAQ
Is ValGuard an eval product? No. Keep your eval harness. Add runtime rules.
Can the judge be the only gate? Not for irreversible side effects.
Where do I learn the sandwich? Why we built ValGuard and AI failure modes.
Related
- vs Guardrails AI
- vs structured outputs
- Validation
- Trust
- Pricing
- Evals vs validation
Next step
Quickstart and shadow mode.