AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models

arXiv · AI, language, vision and robotics · article · Oct 6, 2026 · UTC

Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue,

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-10-07T11:02:23.601Z. This is not the publication date.