I had some updates on Twitter here: https://x.com/a_karvonen/status/2049891045736055260
Basically, there has been no progress since this post. Even Claude Fable 5 makes no improvement. Deepseek V4-Preview actually did a bit better than Claude Opus 4.8.
Over the last two years there was some progress on visual ability to correctly identify more features of the part (although no model can consistently identify all features correctly) and little to no apparent progress on spatial or mechanical reasoning.
Copy pasting the most relevant part:
“Wow, Deepseek’s vision is pretty good and better than Claude’s.
One of the trickiest vision parts is identifying there are two opposing flats, where you have to look very closely at the left part. Deepseek also identifies the cross hole on the upper end of the left part (which GPT-5.5-Pro misidentified as a slot instead of a hole).
But then Deepseek totally botches it on the radial cross hole by saying “Stand the part perfectly upright, resting the bottom flat face on the parallels.”
Where it wants to drill the cross-hole in the top of the left part axially, even though the hole is radial.
This is the type of mistake that all LLMs make that makes me believe they have very poor spatial reasoning ability.”
My understanding is that “GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna” meant something like “make some modifications to a run scheduler file and launch the post-training script”, not “selected, configured, adapted, and ran RL gyms”.
There’s more info / discussion here: https://x.com/nikolaj2030/status/2075297831376793764