Recursive self-improvement in large language models
How do large language models improve themselves recursively, through which mechanisms, with what empirical results, and under what limits?
This review asks how large language models improve themselves recursively. Across 156 sources spanning 2006 to 2026, single-loop self-improvement — self-training, self-correction, self-play, self-rewarding — is repeatedly shown to deliver real gains, but every mechanism also has a documented failure mode, and the recursion reliably stalls or degrades without an external anchor such as verifiable feedback, fresh data, or human judgment. Confidence is moderate: the field is young, dominated by author-run benchmarks and preprints, and several headline results are actively contested.
Updated 10 Aug 2026156 sources2006–2026Deep28 min read
recursive self-improvement · self-training · self-correction · self-play · RLAIF · model collapse · weak-to-strong generalization · self-evolving agents · intelligence explosion