Hi Brian and Luc, thanks for sharing this. I’ve been studying the same question and found a different result in my paper. Length penalties left hint influence near baseline but made Qwen3-4B and 14B mention the hint less often in their CoT. My length-matched control showed the same pattern.
Our setups are a tad different. You train on math with a group-relative penalty, while I train on MMLU-Pro with prompt-specific compression targets.
Hi Brian and Luc, thanks for sharing this. I’ve been studying the same question and found a different result in my paper. Length penalties left hint influence near baseline but made Qwen3-4B and 14B mention the hint less often in their CoT. My length-matched control showed the same pattern.
Our setups are a tad different. You train on math with a group-relative penalty, while I train on MMLU-Pro with prompt-specific compression targets.