AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

From Gaze to Meaning: A Training-Free AI Agent for Unified Grounding and Explanation

arXiv · AI, language, vision and robotics · article · Sep 5, 2026 · UTC

Understanding human attention is fundamental for scene interpretation, yet existing approaches often rely on heavily trained models that lack interpretability. Prior methods struggle to jointly reason about gaze targets, attended objects, and visual grounding without extensive supervision. To the best of our knowledge, this work introduces the first training-free Gaze Target Agent (GTA) for gaze-guided reasoning across tasks such as gaze target prediction, attention localization, and object identification. This is achieved by leveraging pretrained vision-language models, augmenting them with v

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:32:07.623Z. This is not the publication date.