AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

arXiv · AI, language, vision and robotics · article · Sep 17, 2026 · UTC

Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevant token among many irrelevant ones. We call this failure mode mean dilution and propose RBS-Attention, a training-free sparse-prefill method with two complementary selection branches. A centroid base branch captures average relevance, while a rescue branch uses the maximum key-block radius and its prompt-, layer-, and head-dependent distribution

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T14:12:08.350Z. This is not the publication date.