The nanogpt speedrun feels more like developing better methods to culture e coli at a hobbyist level, and quite unlikely to lead to any substantial advancement applicable to the operational efficiency of well-funded companies at the frontier.
If you’ll permit a bit of snark, I think that your comment was wrong even when it was written in October 2025.
The Muon optimizer is the clearest example of a hobbyiest-to-frontier transfer of all the techniques I know. Keller Jordan introduced Muon on specifically the nanoGPT speedrun challenge in a tweet thread from October 2024. (He was unsurprisingly hired to work at OpenAI on pretraining shortly after.) Muon seems to enable stable training at large scales, at least moreso than Adam. As evidence of this, by the time you wrote your comment, Muon was used as an optimizer by MoonShot AI for Kimi K2 as well as Zhipu for GLM-4.5, and has seen continued use (e.g. for GLM-5).
If you’ll permit a bit of snark, I think that your comment was wrong even when it was written in October 2025.
The Muon optimizer is the clearest example of a hobbyiest-to-frontier transfer of all the techniques I know. Keller Jordan introduced Muon on specifically the nanoGPT speedrun challenge in a tweet thread from October 2024. (He was unsurprisingly hired to work at OpenAI on pretraining shortly after.) Muon seems to enable stable training at large scales, at least moreso than Adam. As evidence of this, by the time you wrote your comment, Muon was used as an optimizer by MoonShot AI for Kimi K2 as well as Zhipu for GLM-4.5, and has seen continued use (e.g. for GLM-5).