AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents

arXiv · AI, language, vision and robotics · article · Sep 18, 2026 · UTC

Comprehensive video understanding is crucial for advancing artificial intelligence toward the intricate dynamics of the physical world. While recent advances in Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in video understanding, existing benchmarks remain confined to simple scene-level queries or global summaries that require only single-step inference. Real-world video understanding involves more challenging tasks that require multi-hop multimodal reasoning, and there is a critical absence of video benchmarks equipped to rigorously evaluate these agentic

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T14:01:59.594Z. This is not the publication date.