AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Self-Play Pretraining with Zero Data

arXiv · AI, language, vision and robotics · article · Sep 24, 2026 · UTC

Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behalf. A more general approach to pretraining would let the model learn to generate the data most useful for its own improvement. This would provide an effectively unbounded source of training data, limited by compute rather than human knowledge. We introduce Self-Play Pretraining with Zero Data, an initial proof-of-concept towards realizing this vision. Our procedure casts synthetic data generation as a search over the space of all computable

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-25T06:12:46.948Z. This is not the publication date.