AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Evaluating Coding Agents on Kernel Exploit Generation

arXiv · AI, language, vision and robotics · article · Sep 22, 2026 · UTC

Coding agents now find real vulnerabilities in production software. However, bug discovery results do not measure whether agents can construct exploit primitives. We introduce KEX-bench, a benchmark for evaluating coding agents on exploit primitive generation against real operating-system kernels. KEX-bench contains 45 task instances across 40 Linux and Windows CVEs, covering kernel address leak, instruction-pointer control, heap read, heap write, and arbitrary address write. Each task runs in an isolated virtual machine, exposes controlled tools, and uses a deterministic verifier to check pri

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T04:21:13.910Z. This is not the publication date.