AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

A Single Suffix to Break Them All: Basin-Aware Jailbreaks for Merged Model Families

arXiv · AI, language, vision and robotics · article · Aug 27, 2026 · UTC

Model merging enables combining multiple fine-tuned models without additional training, but its safety implications remain poorly understood. Prior work primarily attributes merging risks to unsafe constituent models, implicitly assuming that merging individually aligned models preserves safety. In contrast, we show that model merging reveals a previously overlooked jailbreak risk rooted in the pretrained foundation model, even when all constituent models are individually safety-aligned. Motivated by this observation, we study a new threat setting where an attacker constructs jailbreak prompts

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T09:11:58.312Z. This is not the publication date.