AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

A Unified Mechanistic Analysis of Knowledge- and Safety-Based Refusals

arXiv · AI, language, vision and robotics · article · Sep 1, 2026 · UTC

Large language models (LLMs) are increasingly trained to decline queries that fall outside their knowledge (knowledge-based refusal, KR) or violate safety policies (safety-based refusal, SR). Although KR and SR result in superficially similar responses, they have largely been studied in isolation, leaving open whether they share an underlying mechanism. We address this gap with a systematic study on a new dataset of 213 contrastive quadruples that jointly probe both refusal types. We find that KR and SR are governed by overlapping yet distinguishable mechanisms. Both share a refusal direction,

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T06:11:57.537Z. This is not the publication date.