AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive Alternatives

arXiv · AI, language, vision and robotics · article · Aug 31, 2026 · UTC

Safety alignment is essential for deploying large language models, requiring systems to prevent harmful compliance while preserving helpfulness on benign requests. Activation steering offers a training-free inference-time approach to safety control, but effective safety steering requires addressing two coupled questions: when to intervene and how generation should be shaped after intervention. However, existing safety steering methods remain limited along both dimensions, as their triggering mechanisms can be unstable across domains and refusal-oriented steering often yields rigid refusals rat

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:22:03.933Z. This is not the publication date.