I would be interested in how well base models do compared to the final reasoning models. E.g. DeepSeek-V4-Pro-Base with a prompt like “[blog post]<br>Posted by ”, versus just asking the post-trained DeepSeek-V4-Pro.
I would be interested in how well base models do compared to the final reasoning models. E.g. DeepSeek-V4-Pro-Base with a prompt like “[blog post]<br>Posted by ”, versus just asking the post-trained DeepSeek-V4-Pro.