not OP, but will note they intentionally focus the main results on tasks with short results. There presumably shouldn’t be a large difference based on number of tokens spent. That would be more about CoT, not a non-CoT measurement.
They do analysis around different sizes of open weight models in Appendix A.16. Not sure if you had a different question than what they address there.
but will note they intentionally focus the main results on tasks with short results.
We do this because it’s a sharper distinct meaning of no-CoT, i.e., in the main plot we restrict to tasks which only require very few forward passes. See the paper and above comments showing that including longer tasks, including generation and agent long-horizon tasks, doesn’t change the trends that much.
not OP, but will note they intentionally focus the main results on tasks with short results. There presumably shouldn’t be a large difference based on number of tokens spent. That would be more about CoT, not a non-CoT measurement.
They do analysis around different sizes of open weight models in Appendix A.16. Not sure if you had a different question than what they address there.
We do this because it’s a sharper distinct meaning of no-CoT, i.e., in the main plot we restrict to tasks which only require very few forward passes. See the paper and above comments showing that including longer tasks, including generation and agent long-horizon tasks, doesn’t change the trends that much.