I’ve found that doing no-reasoning SFT (train on output conditioned on empty reasoning) on reasoning models (like they do in Emil’s paper that you linked) can cook them quite badly. I wasn’t able to prevent the model from completely losing it’s reasoning ability when doing full-weight finetuning on qwen3-32b, and it sometimes reasons weirdly even when trained with LoRA, tho there seems to be some intra-run variance here (it sometimes comes out completely normal).
I’ve found that doing no-reasoning SFT (train on output conditioned on empty reasoning) on reasoning models (like they do in Emil’s paper that you linked) can cook them quite badly. I wasn’t able to prevent the model from completely losing it’s reasoning ability when doing full-weight finetuning on qwen3-32b, and it sometimes reasons weirdly even when trained with LoRA, tho there seems to be some intra-run variance here (it sometimes comes out completely normal).