AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

DeCo: Efficient Decouple-to-Couple Learning for Multi-Task Visual Grounding

arXiv · AI, language, vision and robotics · article · Sep 21, 2026 · UTC

Multi-task visual grounding requires models to jointly understand linguistic semantics and perform accurate visual localization and segmentation. Despite the success of multimodal large language models, effectively adapting them to multiple grounding objectives remains challenging. Existing methods commonly enforce task cooperation through shared representations, while overlooking the intrinsic conflict between task-oriented feature interests. In this paper, we introduce $\textbf{DeCo}$, an efficient $\textbf{De}$couple-to-$\textbf{Co}$uple learning framework that resolves this dilemma through

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T08:01:43.213Z. This is not the publication date.