yes, we are uniquely vulnerable to trusting our intuition with natural language CoT (in a natural language that we speak). suppose it used a natural language that you didn’t speak for CoT. you might object: “I’d just have to get a translator, you’re just raising the cost of the signal”—but in addition you’d be more cautious, you’d be aware of the ambiguities of translation. and that is the appropriate way to actually think about the signals provided by CoT, otherwise your “cheap” signal is just helping you fool yourself—because the natural language you or any other human speaks is already a “translation” for the LLM. now let the model use whatever CoT is native to it: “neuralese,” an “alien language,” probably more like “LLM circuit activation-ese.” now you have the most costly signal, but at least you aren’t actively undermining yourself by providing training on how to avoid detection in the language human’s trust most.
yes, we are uniquely vulnerable to trusting our intuition with natural language CoT (in a natural language that we speak). suppose it used a natural language that you didn’t speak for CoT. you might object: “I’d just have to get a translator, you’re just raising the cost of the signal”—but in addition you’d be more cautious, you’d be aware of the ambiguities of translation. and that is the appropriate way to actually think about the signals provided by CoT, otherwise your “cheap” signal is just helping you fool yourself—because the natural language you or any other human speaks is already a “translation” for the LLM. now let the model use whatever CoT is native to it: “neuralese,” an “alien language,” probably more like “LLM circuit activation-ese.” now you have the most costly signal, but at least you aren’t actively undermining yourself by providing training on how to avoid detection in the language human’s trust most.