AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

From Reward Signal to Visual Utility: A Controlled Audit of Medical VLM Post-Training

arXiv · AI, language, vision and robotics · article · Sep 25, 2026 · UTC

Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy. We examine how changes in accuracy and training objectives relate to image-conditioned decisions in a controlled Qwen2.5-VL-3B study on PMC-VQA. We compare supervised fine-tuning (SFT) with low-rank adaptation (LoRA) restricted to the language model, expanded multimodal adaptation scopes, standard answer-only Group Relative Policy Optimization (GRPO), and a counterfactual evidence objective. On 2,000 clean-test questions, language model LoRA SFT changes correct-image accuracy by +1.10 percentage po

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-28T07:21:24.486Z. This is not the publication date.