I’m curious, what tests are you running to isolate the quality of the K3 pretrain? If I was trying to do this I’d have pretty low confidence in my ability to correctly credit the model’s pre-training vs post-training for its performance on any given test or suite of tests.
I see—I wonder how sensitive this measurement is to labs doing post-training steps which include lots of low/no reasoning token math problems. With K3 currently supporting only ‘max’ reasoning I’m also led to wonder if if Kimi deliberately put lower focus on low/no-reasoning tasks in their post-training and if you could in fact be seeing the results of that rather than of a substantially weak pre-training phase.
I’m curious, what tests are you running to isolate the quality of the K3 pretrain? If I was trying to do this I’d have pretty low confidence in my ability to correctly credit the model’s pre-training vs post-training for its performance on any given test or suite of tests.
Measuring single forward pass performance using the math dataset here. (This is imperfect, but I think it gives a decent sense.)
I see—I wonder how sensitive this measurement is to labs doing post-training steps which include lots of low/no reasoning token math problems. With K3 currently supporting only ‘max’ reasoning I’m also led to wonder if if Kimi deliberately put lower focus on low/no-reasoning tasks in their post-training and if you could in fact be seeing the results of that rather than of a substantially weak pre-training phase.