Langfuse traces LLM generations, scores, and sessions. ValGuard enforces deterministic rules on the OpenAI-compatible path. Use both: ValGuard for pass/block with rule IDs, Langfuse for product analytics and human review timelines. Neither replaces the other.
When you need this
Your Langfuse dashboard shows green generations while refund amounts are wrong. Availability and traces do not equal policy. Attach ValGuard validation headers to the Langfuse generation metadata so a block is visible next to the span that caused it.
Architecture
flowchart LR
App[App / agent] --> VG[ValGuard proxy]
VG --> M[Model]
VG -->|headers + body| App
App -->|trace + metadata| LF[Langfuse]
ValGuard is on the request path. Langfuse is on the observability path. A direct provider call that skips ValGuard also skips rule IDs in your traces. Trust.
Setup
- Create a ValGuard agent and packs. Start in shadow mode.
- Point your OpenAI-compatible client at
https://api.valguard.ai/v1withX-VG-Agent. - Create or update a Langfuse generation/span for the same call.
- Copy
X-Request-Id/X-Trace-Id(and validation status headers) into Langfuse metadata. - Optional: forward W3C
traceparenton the ValGuard request if your tracer already has one.
Code
curl -s "$VG_PROXY/v1/chat/completions" \
-H "Authorization: Bearer $VG_API_KEY" \
-H "X-VG-Agent: $VG_AGENT" \
-H "traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01" \
-H "Content-Type: application/json" \
-d '{"model":"openai/gpt-4o-mini","messages":[{"role":"user","content":"Hi"}]}'
# Store X-Request-Id / X-Trace-Id and X-VG-Validation-Status on the Langfuse generation metadata.
Python sketch (OpenAI SDK + Langfuse metadata; adjust to your Langfuse SDK version):
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.valguard.ai/v1",
api_key=os.environ["VG_API_KEY"],
default_headers={"X-VG-Agent": "support-triage"},
)
raw = client.chat.completions.with_raw_response.create(
model="openai/gpt-4o-mini",
messages=[{"role": "user", "content": "Hi"}],
)
headers = {k.lower(): v for k, v in raw.headers.items()}
meta = {
"vg_request_id": headers.get("x-request-id"),
"vg_trace_id": headers.get("x-trace-id"),
"vg_validation_status": headers.get("x-vg-validation-status"),
}
completion = raw.parse()
# langfuse_generation.update(metadata=meta) # use your Langfuse client API
print(completion.choices[0].message.content, meta)
On validation failure
| Outcome | What to log in Langfuse |
|---|---|
| Pass | Status pass; optional rule evaluation summary |
| Block | Status block; rule IDs; do not mark the business outcome success |
| Re-ask | Final status after retry; count attempts |
| Shadow | Would-block flags while traffic still succeeds |
Prefer structured metadata over parsing free-text errors.
Latency
Rule packs: microseconds. Proxy path: about 0.36 ms p50 with a mocked upstream. Langfuse export is async from your app; do not block the user on the Langfuse HTTP call. Block/re-ask buffering still applies on the ValGuard path. Methodology.
Honest limits
- Langfuse does not enforce ValGuard packs. Tracing a bad refund does not stop it.
- ValGuard is not a full APM. Use Langfuse (or similar) for scores, datasets, and session replay.
- Header names can vary by proxy version; confirm in a staging call.
- Hosted ValGuard still receives request content. Enterprise self-host for VPC. Trust.
Related
- Observability
- Portkey (gateway + traces elsewhere)
- vs LLM-as-judge
- Trust · Pricing
Next step
Quickstart. Add vg_request_id to one Langfuse generation template, then alert on validation block rate.