I think the problem boils down to the claim that training on a specific probe makes that probe useless, but it doesn’t make other uncorrelated probes (especially stronger ones) useless.
For example, if the claim is true, you can train on the CoT and then, say, J-lens probes will keep working. So you lose the ability to do meaning alignment checks on CoT but you can still use J-lens to do them. Thus, from an alignment monitoring perspective, training on CoT is as bad as switching to neuralese which removes CoT completely, but it could be a good tradeoff if you still have other interpretability/​monitoring methods available. The real trouble comes if you are irresponsible and exhaust all of your reliable methods by training on all of them.
I think the problem boils down to the claim that training on a specific probe makes that probe useless, but it doesn’t make other uncorrelated probes (especially stronger ones) useless.
For example, if the claim is true, you can train on the CoT and then, say, J-lens probes will keep working. So you lose the ability to do meaning alignment checks on CoT but you can still use J-lens to do them. Thus, from an alignment monitoring perspective, training on CoT is as bad as switching to neuralese which removes CoT completely, but it could be a good tradeoff if you still have other interpretability/​monitoring methods available. The real trouble comes if you are irresponsible and exhaust all of your reliable methods by training on all of them.