AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security

arXiv · AI, language, vision and robotics · article · Sep 21, 2026 · UTC

LLM-based agents increasingly operate in environments where they interact with users, tools, and external systems. Yet most security evaluations assume passive users and static control, ignoring the interactive dynamics that shape real agent behavior. We introduce \textbf{DUMA-Bench}, a benchmark and evaluation protocol for measuring agent security under \emph{dual-control} interaction, where both the agent and the user can influence the shared environment state. DUMA-Bench extends $τ^2$-bench ~\cite{barres2025tau} with adversarial environments covering eight vulnerability classes, including R

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T06:11:12.848Z. This is not the publication date.