AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

The Structure of Quantization Damage in LLMs: Why the Next Bit Should Be Spent Globally

arXiv · AI, language, vision and robotics · article · Sep 1, 2026 · UTC

Post-training quantization (PTQ) is widely used to reduce the cost of serving large language models (LLMs), but its accuracy cost is uneven and is often tuned per model. We study where quantization damage occurs and how to allocate a small additional precision budget. Using causal mixed-precision intervention as ground truth (raise each layer to 8-bit in turn and measure the accuracy it recovers) across 9 open-weight models in 4 architecture families, we test 3 intuitive hypotheses: that quantization damage lives in task circuits, where the model computes, or in weight statistics. None of them

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T06:01:56.170Z. This is not the publication date.