AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

FailBench: How Reliable are VLMs at Judging Robot Task Success?

arXiv · AI, language, vision and robotics · article · Sep 3, 2026 · UTC

Vision-Language Models (VLMs) are increasingly used to evaluate robot manipulation outcomes, but existing benchmarks offer limited evidence of cross-domain generalization. We introduce FailBench, a benchmark for robot failure detection comprising 2,197 manipulation attempts across 14 public sources (12 real-world, 2 simulated). In FailBench, 75% of failures occur naturally, and six real-world sources come from non-failure-detection datasets. Evaluating 13 VLM-based detectors, we find the best model achieves only 0.77 mean balanced accuracy. Notably, models fine-tuned for failure detection cons

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T04:51:57.792Z. This is not the publication date.