AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

STA-TFM: Spatio-Temporal Aggregation Across Views TransForMer for Pose Estimation

arXiv · AI, language, vision and robotics · article · Sep 21, 2026 · UTC

Monocular 3D human pose estimation (HPE) remains challenging due to depth ambiguity, occlu- sions, and the need for temporal consistency. While multi-view methods provide superior accuracy over monocular approaches, they often require complex setups. We introduce STA-TFM, a transformer-based architecture that combines spatial and temporal information for multi-view pose estimation. The approach leverages DSTformer, a monocular feature extractor, to capture long-range pose dependencies within each view. A fusion transformer then aggregates information across views to produce coherent 3D estimat

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T08:01:43.213Z. This is not the publication date.