AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Towards Full Pipeline FP8 Reinforcement Learning for LLMs

arXiv · AI, language, vision and robotics · article · Sep 19, 2026 · UTC

Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8 RL pipeline remains challenging. While previous works have focused on resolving train-inference mismatches using correction techniques like TIS, we reveal that full-pipeline FP8 RL still suffers from severe training instability, manifesting as anomalous mid-training entropy surges and garbled outputs. We trace this instability to a previously overlooked cause: compo

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T12:01:45.602Z. This is not the publication date.