AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

A Convergence Framework for Deep $V$-Learning: Error Propagation and Sharp Action-Gap Bounds

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

We establish convergence bounds for deep $V$-learning with horizon $H$. The algorithm fits a scalar value function to targets from executed transitions and selects actions using a predictive model and the value function. For current observed-successor targets with fresh true-kernel outcomes, the conditional mean is $\mathcal{T}^βV$, which averages over behavior-policy actions. The Bellman optimality update is $\mathcal{T} V$. We decompose the update error into six residuals: fitting, transition reuse, target construction, replay, action selection, and exploration. Under $L^s$ concentrability,

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-19T20:28:26.698Z. This is not the publication date.