AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Temporal Tree of Thought: Reasoning-Guided Visual Cue Search for Long-Video Understanding

arXiv · AI, language, vision and robotics · article · Aug 28, 2026 · UTC

Long-video understanding remains challenging for Multimodal Large Language Models (MLLMs) due to limited context length. Uniform sampling may miss crucial moments, while agent-based frame video understanding methods often evaluate frames independently, overlooking the temporal organization of videos. Ideally, evidence selection should mimic how humans answer questions about long videos: first locating the relevant segment from the global context, then zooming into local objects and details. We propose Temporal Tree of Thought T^3, a training-free framework for adaptive coarse-to-fine long-vide

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T08:21:55.975Z. This is not the publication date.