AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment

arXiv · AI, language, vision and robotics · article · Sep 5, 2026 · UTC

In this paper, we address the challenging problem of 4D reconstruction from sparse-view videos. This setup usually relies on monocular depth estimation to provide priors for the reconstruction model. A key challenge arises from limited cross-view overlap and temporal variation, making monocular depth predictions inconsistent across views and time. Existing methods align spatial and temporal dimensions in separate stages, requiring foreground segmentation masks while failing to leverage temporal cues for cross-view alignment. Contrary to these methods, we propose a unified spatial-temporal dept

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:32:07.623Z. This is not the publication date.