This doesn’t change the directional conclusions of the post, but the scaling factor seems off comparing distilled models to the model they were distilled on, like GPT 5.6 Luna / GLM 5.3 Flash vs. Sol/GLM 5.3: these are smaller (320/18 vs 753⁄40 billion params for GLM, Luna/Sol probably have a larger difference based on pricing) but the smaller models do not seem like they are 30-100 IQ points smaller (at least to me). [Also note: 2x smaller models are >>2x cheaper to run, I think]
I’d estimate scaling down inference compute using distillation would also cost ~50 IQ points per OOM of parameter count, similar to training. And it doesn’t seem possible to scale up inference compute in the same way, other than increasing the reasoning time/effort per task, and that seems like it gives even fewer returns (especially due to context window limits), probably sub-logarithmic or reaching an asymptotic IQ.
I think the first “following a skills assessment and an interview” should not be there?