SOURCE-LINKED INTELLIGENCE
Riemannian--Lorentz Fusion of Vision Transformers and State-Space Models
Scaling deep learning faces critical bottlenecks: data exhaustion, exponential training costs, and resource concentration. Model merging combines pre-trained checkpoints without gradient descent, offering orders-of-magnitude savings versus retraining. Combining independently trained vision models is difficult when their architectures and parameter shapes differ. Existing weight-space merging methods generally assume aligned, shape-compatible checkpoints, whereas a Vision Transformer (ViT) and a state-space model (SSM) implement token mixing with different operators. We study a hybrid Heterogen
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · Artificial Intelligence · 2026-09-16T20:05:11.000Z
- arXiv · AI, language, vision and robotics · 2026-09-16T20:05:11.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.