AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Think, Look, and Revise: Inconsistency-Aware Visual Self-Correction in MLLMs

arXiv · AI, language, vision and robotics · article · Aug 29, 2026 · UTC

Tool-augmented multimodal reasoning integrates external tools (e.g., object detection, depth estimation) into multimodal large language models (MLLMs) to address perceptual bottlenecks in complex visual tasks. However, existing approaches rarely verify tool outputs, limiting their ability to detect and recover from tool failures. We propose ReVISE, a framework that equips MLLMs with verification and dynamic error recovery for tool-augmented reasoning. ReVISE introduces (1) a curated training dataset that supervises reflective behaviors, enabling models to validate tool-derived evidence, reform

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:51:58.603Z. This is not the publication date.