Benchmarks

Benchmark methodology

Block and re-ask rules wait for the full reply before releasing tokens. Passthrough can stream tokens, but cannot withhold output that fails post-stream checks. Details.

A latency number without a method is marketing. This page explains how we measure ValGuard, what each number includes, and what it leaves out. Results live on Benchmarks, and per-scenario cards are on Benchmark scenarios. Read this page first if you plan to compare our figures with another vendor's chart. The long version of the method is in How We Measure ValGuard Latency.

Three layers, three numbers

LayerWhat it measuresLatest snapshot
EngineDeterministic rule evaluation inside the Go validation engineMicroseconds
HTTP layerThe full request through ValGuard with a mocked upstream0.36 ms p50, 3 ms p99
Playbook, 3 stepsRouting plus a validation on every step, mocked agentssupport 1.8, RAG 1.4 ms p50 (archetype-dependent)
ModelThe provider's completionHundreds of milliseconds

The figures in this table are read from the saved benchmark snapshot. They change when we re-run the benches.

These are ValGuard's own measurements, not an independent benchmark. They measure the listed paths and configurations, not production error reduction or answer quality.

Quote one layer and you misstate the others. A rule pack at a few microseconds says nothing about auth, Postgres, and the persist queue. An end-to-end figure with a live model tells you more about the provider than about us. We publish each layer separately.

What each bench does

Rule engine. go test -benchmem runs on the production validator templates. We record p50 latency, allocations, and bytes per operation. There is no network, no database, and no HTTP envelope. This bench answers one question: does the check itself stall a request?

HTTP layer. This is the live OpenAI-compatible stack: auth, input guard, a production output pack, and the persist queue. Concurrency is 1. The default path pairs the input guard with the customer_support_reply pack (7 rules). The upstream is a zero-latency mock on the local host. The result is ValGuard infrastructure time. It is not model time.

Playbook graphs. orchestration-bench runs published graphs of 3 to 5 steps through X-VG-Flow: linear, branch, merge, and fan-out. Agents are mocked. Concurrency sweeps from 1 to 1000 on the 3-step linear path are on the scenarios page.

A separate overhead bench measures four archetypes: support, invoice, RAG, and lead. Each is a three-step graph with a validation on every step, run at concurrency 1. Playbook pages use these results to estimate overhead for similar graphs. Treat them as estimates for a graph of that shape, not as a measurement of that exact template.

Playbook pages also show a multiplier against the model. It divides a fixed 820 ms reference model call by the measured overhead. The 820 ms is an assumption we chose, not something we measure. Replace it with your own model's latency to get your own ratio.

Streaming

Streaming scenarios consume a full Server-Sent Events upstream and validate the assembled reply. They show whether validation adds cost to stream handling.

They do not turn blocking into token streaming. Block and re-ask rules still buffer the reply first. See the box above and Trust.

The saved mocked-upstream results are not a production time-to-first-token comparison between buffering and passthrough. A fast rule engine does not remove the wait for model completion before a buffered reply can be released.

For your rollout, record both time to first visible token and time to a complete accepted reply. Compare direct-provider, warn/log/shadow passthrough, and buffered block/re-ask paths using the same model, prompts, output length, and concurrency.

Report blocked replies and re-ask attempts separately from first-attempt passes. Include failed requests in the outcome counts; measuring successful replies alone hides the cost of enforcement.

What the numbers do not include

  • A live provider. Every upstream in these benches is mocked.
  • Your environment. Your auth setup, rule packs, persistence settings, and hardware will move the result.
  • Cold starts and noisy neighbors. Each run measures steady state after warmup.
  • Display precision. Saved snapshots retain measurement values; visible labels may round them. Each snapshot on Benchmarks shows when it was generated.

We also do not claim that faster rules make a bad policy correct. A rule that is too strict will still false-positive. It will just do it quickly. Shadow mode exists so you can measure that before you block. See the post on validating before you enforce.

How to read the numbers

If you own latency SLOs: budget the model and the selected enforcement mode together. A buffered path delays the first visible token until completion and validation. Confirm both latency measures with your real rule pack. See latency budgets per step.

If you own risk: look at the pack time for the workflow you actually run. Checks that cost microseconds are cheap enough to run on every request, which is the only way an audit trail is complete.

If you own cost: a deterministic check adds no token bill. An LLM judge does.

Reproduce

On hosted ValGuard, run your own canary. Attach your rule pack to a flow you already run, mock the model, and compare p50 with the layer on and off.

On Enterprise self-host, operators with a repo checkout can re-run the benches on their own hardware:

make bench-validation-save
make bench-proxy-save
make bench-orchestration-save
make bench-playbook-overhead-save

Related