AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA

arXiv · AI, language, vision and robotics · article · Sep 22, 2026 · UTC

Reinforcement learning (RL)-based post-training has become an effective approach for eliciting reasoning capabilities in large language models (LLMs). However, adapting post-RL models to new knowledge domains or behaviors through subsequent supervised fine-tuning (SFT) can severely overwrite these capabilities. Existing approaches mitigate such forgetting through experience replay, specialized initialization, or constrained optimization using gradient projection, but either provide limited preservation or incur substantial training overhead. Our analysis shows that reasoning activations concen

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T04:21:13.910Z. This is not the publication date.