AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence

arXiv · AI, language, vision and robotics · article · Aug 31, 2026 · UTC

Recent advances in large reasoning models (LRMs) have shown that reinforcement learning with verifiable rewards (RLVR) can substantially improve reasoning in mathematics and code, where outcomes can be checked automatically. Extending this progress to open-ended and agentic tasks remains difficult because reliable rewards are harder to obtain and direct human supervision cannot keep pace with the scale and complexity of model-generated experience. This paper studies how LRMs can continue to improve as human supervision gradually recedes from the learning loop. We examine two connected dimensio

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T06:41:57.136Z. This is not the publication date.