AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione

arXiv · AI, language, vision and robotics · article · Sep 6, 2026 · UTC

Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms. In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention. We discover that a sparse subset of Hypersensitive Safety Heads misfires on Hard-Safe prompts, exhibiting abnormal attention entanglement that forcefully binds harmless target entities t

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-25T16:52:32.424Z. This is not the publication date.