This seems like a terrible take. Why didn’t GPT-4.5 crush at math then, if it was such a big pretrain? Obviously you need RL to get good performance at solving math problems (and in order to effectively leverage inference time compute and work together for a long time in an agent swarm.)
This seems like a terrible take. Why didn’t GPT-4.5 crush at math then, if it was such a big pretrain? Obviously you need RL to get good performance at solving math problems (and in order to effectively leverage inference time compute and work together for a long time in an agent swarm.)