AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

arXiv · AI, language, vision and robotics · article · Sep 1, 2026 · UTC

Benign fine-tuning severely weakens the safety alignment of large language models (LLMs), so we study why refusal behavior is so fragile. While prior work often attributes this failure to gradient conflict, we propose a fundamentally different Fisher-geometric explanation: safety Fisher is low-rank, and alignment makes the safety geometry flatter while preserving an output-routing pathway. After 100 benign fine-tuning examples, this pathway is selectively re-sharpened in output-side MLP modules, explaining the asymmetric fragility: safety can collapse to high attack success rates, while genera

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T06:01:56.170Z. This is not the publication date.