AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection Guardrails

arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC

Verdict-only evaluation does not reveal whether a vision-language model (VLM) used the visual evidence that should support its decision. We study this problem in web-agent guardrails, where a VLM judges whether on-screen text conflicts with a user instruction. We introduce Mind2Web-Injection, a benchmark of 9,954 instruction-screenshot pairs with instruction-relative labels, pixel-exact evidence boxes, and matched image-side counterfactuals. Across six VLMs, two models with nearly identical average precision differ ninefold in Evidence-Aligned Detection (EAD), the fraction of attacks both dete

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T05:51:54.566Z. This is not the publication date.