We don’t really care about predicting the next token though
Ultimately, we care about finishing properly some tasks.
Without SFT and RL, using an LLM would be insufferable
I wonder what would be the capacity of a base model without RL nowadays in 2026. has anyone run pass@1000 or majority-vote sampling on a 2026 frontier base model (pre-RLVR) on something like current SWE-bench or a recent competition-math set?
We don’t really care about predicting the next token though
Ultimately, we care about finishing properly some tasks.
Without SFT and RL, using an LLM would be insufferable
I wonder what would be the capacity of a base model without RL nowadays in 2026. has anyone run pass@1000 or majority-vote sampling on a 2026 frontier base model (pre-RLVR) on something like current SWE-bench or a recent competition-math set?