AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Train Overcomplete, Deploy Compact: Scaling Recovery Capacity for Structured LLM Pruning

arXiv · AI, language, vision and robotics · article · Sep 7, 2026 · UTC

Large language models achieve strong performance across diverse tasks, but deployment remains costly because of memory, latency, and energy demands. Structured pruning reduces these costs by removing architectural components, yet its recovery stage is often limited by a mismatch between the recovery module's representational capacity and the complexity of the removed knowledge. We call this bottleneck the capacity-knowledge asymmetry and propose OverRep, an Overcomplete Reparameterization framework for structured LLM pruning. Following the principle of "train overcomplete, deploy compact", Ove

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:52:10.320Z. This is not the publication date.