AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction

arXiv · AI, language, vision and robotics · article · Sep 5, 2026 · UTC

Vulnerability discovery is becoming an important ability of large language model (LLM) agents: agents that silently miss real defects leave critical software exposed. Rigorously measuring this ability is therefore urgent, but existing benchmarks are gameable through data contamination, score recall against an unknowable vulnerability set, often rely on synthetic bugs, and report a single end-to-end verdict that cannot localize where an agent fails. Vulnerability discovery is a composite ability: an agent must comprehend source code, infer input constraints, construct inputs, execute them, and

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:32:07.623Z. This is not the publication date.