AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware

arXiv · AI, language, vision and robotics · article · Aug 31, 2026 · UTC

Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache and quadratic attention cost. State-space models (SSMs) mitigate this through linear attention and fixed-size recurrent states, but their large dense linear projections remain computationally expensive even after quantization. We introduce a method that induces sparse neural activity in heavily quantized linear-attention models with minimal performance loss. Activations below a per-projection trainable threshold ($\pm Δ$) are nullified while preserving crucial outliers, achieving compara

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:01:58.596Z. This is not the publication date.