AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model Families

arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC

Quantization enables deployment of large language models on resource-constrained clinical edge devices, but its effect on clinical accuracy and safety remains understudied. We evaluate five 7-8B parameter models at FP16, GPTQ-INT8, and GPTQ-INT4 precision across five benchmarks: MedQA, MedMCQA, Med-HALT, a risk-stratified sample of HealthBench, and MedSafetyBench. The study jointly varies quantization bit width, model family, and clinical task type, with explicit risk stratification and safety measures. INT8 GPTQ is universally safe (max. degradation -1.9%-1.9%), while INT4 degradation is subs

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-26T08:21:45.852Z. This is not the publication date.