As a pretrain, it’s probably somewhere between 4.8 and Mythos (around halfway between?).
I thought about this more and realized this was probably overestimating how good of a pretrain it is (given how good prior models like K2.6 were as pretrains). So I ran some tests.
My quick tests indicated that the Kimi K3 pretrain is around halfway between Opus 4 and Opus 4.5. So ~10 months behind Anthropic. These tests probably understate data improvements, so overall I think it’s a similarly good pretrain to Opus 4.5 (~8 months behind). These tests are better at measuring “general pretrain capability” than at incorporating (coding-specific) data quality.
So my claim that “As a pretrain, it’s probably somewhere between 4.8 and Mythos (around halfway between?)” seems significantly too bullish on the model!
I think Mythos is a pretty big step up in pretraining, so K3 might be more than 8 months behind on the historical pretraining trend relative to Mythos (as in, Mythos is >>3 months ahead of K3 and Mythos was fully done training ~5 months ago).
Overall, this makes me suspect more of the improvements are due to distillation-type effects and makes me think the full catch-up times would be somewhat longer (if Ant/OpenAI stopped but investment still followed current trends). Minimally, more of the improvement probably lives in post-training/mid-training.
For reference, the same test indicates K2.6 is around halfway between Sonnet 4 and Sonnet 4.5. (And this roughly corresponds to some other similar measures.)
I’m curious, what tests are you running to isolate the quality of the K3 pretrain? If I was trying to do this I’d have pretty low confidence in my ability to correctly credit the model’s pre-training vs post-training for its performance on any given test or suite of tests.
I see—I wonder how sensitive this measurement is to labs doing post-training steps which include lots of low/no reasoning token math problems. With K3 currently supporting only ‘max’ reasoning I’m also led to wonder if if Kimi deliberately put lower focus on low/no-reasoning tasks in their post-training and if you could in fact be seeing the results of that rather than of a substantially weak pre-training phase.
I thought about this more and realized this was probably overestimating how good of a pretrain it is (given how good prior models like K2.6 were as pretrains). So I ran some tests.
My quick tests indicated that the Kimi K3 pretrain is around halfway between Opus 4 and Opus 4.5. So ~10 months behind Anthropic. These tests probably understate data improvements, so overall I think it’s a similarly good pretrain to Opus 4.5 (~8 months behind). These tests are better at measuring “general pretrain capability” than at incorporating (coding-specific) data quality.
So my claim that “As a pretrain, it’s probably somewhere between 4.8 and Mythos (around halfway between?)” seems significantly too bullish on the model!
I think Mythos is a pretty big step up in pretraining, so K3 might be more than 8 months behind on the historical pretraining trend relative to Mythos (as in, Mythos is >>3 months ahead of K3 and Mythos was fully done training ~5 months ago).
Overall, this makes me suspect more of the improvements are due to distillation-type effects and makes me think the full catch-up times would be somewhat longer (if Ant/OpenAI stopped but investment still followed current trends). Minimally, more of the improvement probably lives in post-training/mid-training.
For reference, the same test indicates K2.6 is around halfway between Sonnet 4 and Sonnet 4.5. (And this roughly corresponds to some other similar measures.)
Sorry about the error.
I’m curious, what tests are you running to isolate the quality of the K3 pretrain? If I was trying to do this I’d have pretty low confidence in my ability to correctly credit the model’s pre-training vs post-training for its performance on any given test or suite of tests.
Measuring single forward pass performance using the math dataset here. (This is imperfect, but I think it gives a decent sense.)
I see—I wonder how sensitive this measurement is to labs doing post-training steps which include lots of low/no reasoning token math problems. With K3 currently supporting only ‘max’ reasoning I’m also led to wonder if if Kimi deliberately put lower focus on low/no-reasoning tasks in their post-training and if you could in fact be seeing the results of that rather than of a substantially weak pre-training phase.