ValGuard + Langfuse: rule IDs next to your traces

Enforce packs on the ValGuard path. Attach X-Request-Id and validation status to Langfuse generation metadata so blocks show up in the same timeline.

Last verified:

Langfuse traces LLM generations, scores, and sessions. ValGuard enforces deterministic rules on the OpenAI-compatible path. Use both: ValGuard for pass/block with rule IDs, Langfuse for product analytics and human review timelines. Neither replaces the other.

When you need this

Your Langfuse dashboard shows green generations while refund amounts are wrong. Availability and traces do not equal policy. Attach ValGuard validation headers to the Langfuse generation metadata so a block is visible next to the span that caused it.

Architecture

flowchart LR
  App[App / agent] --> VG[ValGuard proxy]
  VG --> M[Model]
  VG -->|headers + body| App
  App -->|trace + metadata| LF[Langfuse]

ValGuard is on the request path. Langfuse is on the observability path. A direct provider call that skips ValGuard also skips rule IDs in your traces. Trust.

Setup

  1. Create a ValGuard agent and packs. Start in shadow mode.
  2. Point your OpenAI-compatible client at https://api.valguard.ai/v1 with X-VG-Agent.
  3. Create or update a Langfuse generation/span for the same call.
  4. Copy X-Request-Id / X-Trace-Id (and validation status headers) into Langfuse metadata.
  5. Optional: forward W3C traceparent on the ValGuard request if your tracer already has one.

Code

curl -s "$VG_PROXY/v1/chat/completions" \
  -H "Authorization: Bearer $VG_API_KEY" \
  -H "X-VG-Agent: $VG_AGENT" \
  -H "traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01" \
  -H "Content-Type: application/json" \
  -d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"Hi"}]}'
# Store X-Request-Id / X-Trace-Id and X-VG-Validation-Status on the Langfuse generation metadata.

Python sketch (OpenAI SDK + Langfuse metadata; adjust to your Langfuse SDK version):

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.valguard.ai/v1",
    api_key=os.environ["VG_API_KEY"],
    default_headers={"X-VG-Agent": "support-triage"},
)

raw = client.chat.completions.with_raw_response.create(
    model="openai/gpt-4o-mini",
    messages=[{"role": "user", "content": "Hi"}],
)
headers = {k.lower(): v for k, v in raw.headers.items()}
meta = {
    "vg_request_id": headers.get("x-request-id"),
    "vg_trace_id": headers.get("x-trace-id"),
    "vg_validation_status": headers.get("x-vg-validation-status"),
}
completion = raw.parse()
# langfuse_generation.update(metadata=meta)  # use your Langfuse client API
print(completion.choices[0].message.content, meta)

On validation failure

OutcomeWhat to log in Langfuse
PassStatus pass; optional rule evaluation summary
BlockStatus block; rule IDs; do not mark the business outcome success
Re-askFinal status after retry; count attempts
ShadowWould-block flags while traffic still succeeds

Prefer structured metadata over parsing free-text errors.

Latency

Rule packs: microseconds. Proxy path: about 0.36 ms p50 with a mocked upstream. Langfuse export is async from your app; do not block the user on the Langfuse HTTP call. Block/re-ask buffering still applies on the ValGuard path. Methodology.

Honest limits

  • Langfuse does not enforce ValGuard packs. Tracing a bad refund does not stop it.
  • ValGuard is not a full APM. Use Langfuse (or similar) for scores, datasets, and session replay.
  • Header names can vary by proxy version; confirm in a staging call.
  • Hosted ValGuard still receives request content. Enterprise self-host for VPC. Trust.

Related

Next step

Quickstart. Add vg_request_id to one Langfuse generation template, then alert on validation block rate.