AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

arXiv · AI, language, vision and robotics · article · Sep 24, 2026 · UTC

Multimodal large language models (MLLMs) achieve strong performance on visual reasoning tasks, yet remain prone to hallucinations and over-reliance on language priors, often generating answers without adequately using task-relevant visual evidence. Existing approaches primarily improve reasoning through reasoning-oriented supervision or inference-time strategies. In this work, we study a complementary question: can multimodal reasoning be improved by strengthening implicit visual grounding without directly supervising the reasoning process? Motivated by the functional specialization of attenti

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-25T06:12:46.948Z. This is not the publication date.