sorry, why don’t we think natural language CoT can be inaccurate/misleading/deliberately deceptive? is it just “often, it doesn’t appear to be (as far as we can tell (as long as you don’t think too much about it))”?
CoTs can be inaccurate/misleading/deliberately. But prior work (Baker et al., 2025) has shown that CoT monitors can be used to detect bad behaviors like reward hacking, and are better at detecting reward hacking than action-only monitors. So the CoT isn’t so inaccurate/misleading that we should dismiss it immediately!
Also, as nostalgebraist argues in this post, the extra computation afforded to the model during reasoning must pass through the CoT bottleneck. Since that bottleneck is natural language, you can just read it. This interp ‘technique’ avoids a lot of problems in other interp techniques (e.g., features in an SAE often don’t make sense and require their own interpretation, while a CoT generally does make sense in the context of the task).
“we find that with too much optimization, agents learn obfuscated reward hacking, hiding their intent within the CoT while still exhibiting a significant rate of reward hacking”
oh interesting. I didn’t know that this had been experimentally demonstrated
as nostalgebraist argues in this post, the extra computation afforded to the model during reasoning must pass through the CoT bottleneck
really interesting post and discussion, thank you.
Since that bottleneck is natural language, you can just read it
“Epiphenomenal … Hidden parallelized … [and] Steganography … [which is] more tractable than the other two [!?]”
“just read it” ⇐ not my take away, but at least it only requires a fundamental breakthrough in steganography, which we probably get during RSI
sorry, why don’t we think natural language CoT can be inaccurate/misleading/deliberately deceptive? is it just “often, it doesn’t appear to be (as far as we can tell (as long as you don’t think too much about it))”?
CoTs can be inaccurate/misleading/deliberately. But prior work (Baker et al., 2025) has shown that CoT monitors can be used to detect bad behaviors like reward hacking, and are better at detecting reward hacking than action-only monitors. So the CoT isn’t so inaccurate/misleading that we should dismiss it immediately!
Also, as nostalgebraist argues in this post, the extra computation afforded to the model during reasoning must pass through the CoT bottleneck. Since that bottleneck is natural language, you can just read it. This interp ‘technique’ avoids a lot of problems in other interp techniques (e.g., features in an SAE often don’t make sense and require their own interpretation, while a CoT generally does make sense in the context of the task).
oh interesting. I didn’t know that this had been experimentally demonstrated
really interesting post and discussion, thank you.
“just read it” ⇐ not my take away, but at least it only requires a fundamental breakthrough in steganography, which we probably get during RSI