‘Already, Sol has been transforming our research program. As one example, GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna.’ https://youtu.be/Wq45rvPGNHs?t=1240
I asked for clarification here and Ted Sanders’ response indicates the thing that actually happened was something similar to my understanding:
My understanding is that you gave Sol a small task involved in the post-training process (taking a config, making small modifications to a run scheduler file, and starting a run using that config and modified run scheduler file), and it successfully completed that task in a controlled environment (and this wasn’t part of the actual Luna post-training process). Could you confirm whether this is correct?
(this is very different from the conclusion I jumped to when I saw the text of your post, which was “Sol, in the real world, with minimal instruction, conducted all of the work involved in pre-training the real Luna”)
This would be totally insane if they mean literally autonomously. 5.6 Luna is something like a $10 billion product, so even at $1k per human hour it would be worth spending 50 engineer-years to improve its quality by 1%. I would predict that well over 50 human engineer-years went into 5.6 Luna in total, and lots of this into post-training. Maybe 5.6 Sol was monitoring posttraining using a pipeline that needed tens of engineer years to work out all the issues in.
This assumes that economic outcomes are linearly correlated with quality, but historically that hasn’t really been the case. Especially not for something that the market is currently unable to meaningfully evaluate on the grounds of quality.
Getting 1% better on a cheap model doesn’t necessarily correlate to any tangible revenue for OpenAI; the economics of consumer behavior operates in step changes and vibes.
‘Already, Sol has been transforming our research program. As one example, GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna.’ https://youtu.be/Wq45rvPGNHs?t=1240
More details. (This looks like it involved less judgment/research taste/novelty than ‘autonomously post-trained’ made me think of.)
I asked for clarification here and Ted Sanders’ response indicates the thing that actually happened was something similar to my understanding:
Ted implies the model:
I also say:
This would be totally insane if they mean literally autonomously. 5.6 Luna is something like a $10 billion product, so even at $1k per human hour it would be worth spending 50 engineer-years to improve its quality by 1%. I would predict that well over 50 human engineer-years went into 5.6 Luna in total, and lots of this into post-training. Maybe 5.6 Sol was monitoring posttraining using a pipeline that needed tens of engineer years to work out all the issues in.
This assumes that economic outcomes are linearly correlated with quality, but historically that hasn’t really been the case. Especially not for something that the market is currently unable to meaningfully evaluate on the grounds of quality.
Getting 1% better on a cheap model doesn’t necessarily correlate to any tangible revenue for OpenAI; the economics of consumer behavior operates in step changes and vibes.