
LEGO-Anything shows promise, but benchmark results expose weak geometric self-assessment.
OpenAI says GPT-5.6 Sol adapted and ran post-training for the smaller Luna model, while scoring 16.2 points above GPT-5.5 on an internal recursive self-improvement benchmark.


OpenAI says GPT-5.6 Sol independently post-trained Luna, a smaller model, after Luna’s initial pre-training. The task was initiated through Codex with a “fairly under-specified prompt” that asked Sol to find training configurations, select suitable GPUs, launch the training script, and verify the run.
The key takeaway: OpenAI is positioning Sol not just as a model that answers questions, but as a system that can operate inside parts of the AI development workflow.

OpenAI built an internal evaluation suite for real-world AI research tasks, including debugging research systems, optimizing kernels and training recipes, running machine learning experiments, and improving another model. On the aggregated Recursive Self-Improvement index, GPT-5.6 Sol scores 16.2 points higher than GPT-5.5, according to OpenAI.
The reported model hierarchy places Sol above Terra and Luna variants, followed by GPT-5.5 and GPT-5.4.
OpenAI employee Jason Liu later clarified that Sol did not invent a complete training recipe from scratch. Most of the configuration already existed from Sol’s own post-training, and the task was to adapt that setup for Luna and run the training job.
Liu said the work would otherwise have “taken two staff researchers maybe an extra two weeks,” framing the result as still significant while narrowing what “autonomous” means in this case.
The article frames Sol as part of a broader push by AI labs to use AI systems to accelerate their own development. OpenAI says researchers use GPT-5.6 Sol across debugging, training-system optimization, experiments, and result analysis.
OpenAI also says average daily token output per active researcher more than doubled the previous peak set by GPT-5.5, while pull requests and experiments per researcher increased. Those metrics do not directly prove faster scientific progress, but they show how rapidly AI-assisted research workflows are scaling inside the company.

LEGO-Anything shows promise, but benchmark results expose weak geometric self-assessment.
Altman says AI benefits outweigh some harms, while rejecting catastrophic risks.
A former OpenAI safety leader says the company is not being careful enough.

RRSI aims to help self-improving AI agents generalize beyond their test tasks while using fewer tokens.