AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

Massive KV caches can cause severe memory-bandwidth bottlenecks during long-context decoding. Sparse attention methods mitigate this via selective loading, but that comes at a cost: rigid heuristics drop necessary context, leading to quality degradation. We introduce \textbf{Elastic Threshold Attention (ETA)}, an end-to-end trainable architecture that achieves hardware-accelerated decoding speed without sacrificing dense model quality. ETA predicts dynamic, contextual thresholds directly from query representations, allowing the model to allocate dense-like context to difficult retrieval or rea

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T17:51:24.264Z. This is not the publication date.