AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

From Alignment to Fusion in 3D Vision-Language

arXiv · AI, language, vision and robotics · article · Sep 23, 2026 · UTC

Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to language-guided reasoning. Existing methods often process point clouds, voxel grids, and multi-view images independently; directly combining these heterogeneous representations may leave substantial feature discrepancy unresolved, while subsequent unconstrained adaptation may distort their internal geometry. We propose an align-then-fuse framework that first applies triple pairwise cosine alignment to establish segment-level correspondence acro

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-24T08:22:30.429Z. This is not the publication date.