I sometimes see discussions about whether a new weirdness observed in CoTs is a sign of coming steganography. I think people in these discussions miss that the most important consideration is preservation of our ability to say whether we still can monitor CoT. Checking if model still thinks in English is easier than checking if gibberish is still decipherable, same applies to checking if CoT is still faithful, etc.
I sometimes see discussions about whether a new weirdness observed in CoTs is a sign of coming steganography. I think people in these discussions miss that the most important consideration is preservation of our ability to say whether we still can monitor CoT. Checking if model still thinks in English is easier than checking if gibberish is still decipherable, same applies to checking if CoT is still faithful, etc.