AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Information-Time Proximal Policy Optimization

arXiv · AI, language, vision and robotics · article · Sep 21, 2026 · UTC

RLVR has substantially improved the reasoning capabilities of LLMs. However, existing methods typically parameterize temporal progression in the Markov Decision Process by token-by-token generation, despite the highly non-uniform information flow along autoregressive trajectories. In this paper, we propose InfoPPO, which reparameterizes temporal progression using information density rather than raw token count. This reparameterization induces a common state-dependent structure for both temporal credit propagation and policy updates. InfoPPO restores the effectiveness of non-trivial discounting

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T08:01:43.213Z. This is not the publication date.