Prompt injection: deterministic LLM output check

Block prompt override attempts in output.

Block prompt override attempts in output.

This page is the SEO reference for the prompt_injection validator. Configure it on an agent in the dashboard or via the agents API. Full catalog: validators.

What it checks

Block prompt override attempts in output.

FieldValue
Rule IDprompt_injection
Packcore
CategorySecurity
Default severityerror
Default on_failblock

Config notes

Attach the rule to an agent, set severity and on_fail, then start in shadow mode before enforce. Pack core must be allowed on your plan.

Pass / fail example

Pass: output satisfies the encoded contract for this rule (shape, pattern, or checksum as configured).

Fail: output violates the contract. ValGuard returns a validation failure with this rule ID when on_fail is block or reask.

Exact fixtures depend on your field paths and thresholds. Use the dashboard tester or validate LLM output.

on_fail behaviours

ModeEffect
blockRequest fails closed; next step should not run
reaskProxy may retry the model once within plan limits; buffers stream
warn / logTraffic continues; audit still records the evaluation
shadow agentWould-block logged; response still delivered

Prefer fail-closed for money and PHI paths.

Performance

Engine timing: 9.6 µs engine p50 (snapshot). HTTP path and model time dominate end-to-end latency. See methodology. Numbers are ValGuard's own measurements on published hardware snapshots.

Related

Next step

Quickstart. Attach prompt_injection in shadow on one agent, then enforce.