I see—I wonder how sensitive this measurement is to labs doing post-training steps which include lots of low/no reasoning token math problems. With K3 currently supporting only ‘max’ reasoning I’m also led to wonder if if Kimi deliberately put lower focus on low/no-reasoning tasks in their post-training and if you could in fact be seeing the results of that rather than of a substantially weak pre-training phase.
I see—I wonder how sensitive this measurement is to labs doing post-training steps which include lots of low/no reasoning token math problems. With K3 currently supporting only ‘max’ reasoning I’m also led to wonder if if Kimi deliberately put lower focus on low/no-reasoning tasks in their post-training and if you could in fact be seeing the results of that rather than of a substantially weak pre-training phase.