AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

GRIP: Gaussian Rendering as a Cross-Modal Bridge for Image-to-Point Cloud Registration

arXiv · AI, language, vision and robotics · article · Sep 22, 2026 · UTC

This paper introduces GRIP, a pose-conditioned refinement framework for pixel-to-point matching and 2D to 3D registration. Given an initial coarse pose estimate, GRIP addresses the structural mismatch between grid based image descriptors and unordered point cloud descriptors by softly rendering learned 3D point features onto the image grid through Gaussian feature splatting. The rendered point derived feature map is then fused with image features by a pixel aligned transformer, enabling visual semantic and geometric cues to interact in a shared 2D representation. The refined features are decod

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T04:11:12.117Z. This is not the publication date.