AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

MicroQonv: Reshaping Convolution Tensors for Efficient Microscaling in Training and Inference

arXiv · AI, language, vision and robotics · article · Sep 23, 2026 · UTC

Microscaling quantization techniques are increasingly used to represent neural network parameters with 8 bits or fewer while preserving near-full precision accuracy. However, applying these methods efficiently in convolutional layers is not straightforward. A naive approach transfers full-precision weights and activations to processing units and quantizes each tensor twice, resulting in much more memory movement than expected. Additional overhead comes from the activation tensors, whose sizes grow substantially because of the im2col transformation applied before quantization. We propose MicroQ

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-24T08:22:30.429Z. This is not the publication date.