AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Task-Aware QUBO Allocation for Mixed-Precision Quantization

arXiv · AI, language, vision and robotics · article · Sep 4, 2026 · UTC

Mixed-precision quantization requires discrete allocation of weight and activation bit-widths, followed by recovery of the selected network. We develop a task-aware quadratic unconstrained binary optimization (QUBO) surrogate with separate weight and activation profiles, a bit-operation (BOP) cost, and selected structural priors. QUBO provides a network-wide allocation that can be refined through direct validation-based PROTES search. On a compact NAFBlock-based denoiser, the refined route achieves 37.192 dB after LSQ+ at 4.035\% routed-layer BOPs, versus 37.092 dB at 4.101\% for a HAWQ-style

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-25T21:32:25.884Z. This is not the publication date.