Open AI3 mins read

OpenAI says GPT-5.6 Sol autonomously post-trained Luna from an underspecified prompt

OpenAI says GPT-5.6 Sol adapted and ran post-training for the smaller Luna model, while scoring 16.2 points above GPT-5.5 on an internal recursive self-improvement benchmark.

What OpenAI says Sol did

Partially redacted prompt used to instruct GPT-5.6 Sol to post-train Luna
Image credits:OpenAI

OpenAI says GPT-5.6 Sol independently post-trained Luna, a smaller model, after Luna’s initial pre-training. The task was initiated through Codex with a “fairly under-specified prompt” that asked Sol to find training configurations, select suitable GPUs, launch the training script, and verify the run.

The key takeaway: OpenAI is positioning Sol not just as a model that answers questions, but as a system that can operate inside parts of the AI development workflow.

The benchmark claim: 16.2 points above GPT-5.5

Recursive Self-Improvement benchmark showing GPT-5.6 Sol ahead of GPT-5.5
Image credits:OpenAI

OpenAI built an internal evaluation suite for real-world AI research tasks, including debugging research systems, optimizing kernels and training recipes, running machine learning experiments, and improving another model. On the aggregated Recursive Self-Improvement index, GPT-5.6 Sol scores 16.2 points higher than GPT-5.5, according to OpenAI.

The reported model hierarchy places Sol above Terra and Luna variants, followed by GPT-5.5 and GPT-5.4.

Important context: not a full recipe from scratch

OpenAI employee Jason Liu later clarified that Sol did not invent a complete training recipe from scratch. Most of the configuration already existed from Sol’s own post-training, and the task was to adapt that setup for Luna and run the training job.

Liu said the work would otherwise have “taken two staff researchers maybe an extra two weeks,” framing the result as still significant while narrowing what “autonomous” means in this case.

Why it matters for AI research teams

The article frames Sol as part of a broader push by AI labs to use AI systems to accelerate their own development. OpenAI says researchers use GPT-5.6 Sol across debugging, training-system optimization, experiments, and result analysis.

OpenAI also says average daily token output per active researcher more than doubled the previous peak set by GPT-5.5, while pull requests and experiments per researcher increased. Those metrics do not directly prove faster scientific progress, but they show how rapidly AI-assisted research workflows are scaling inside the company.

Discover More