“we find that with too much optimization, agents learn obfuscated reward hacking, hiding their intent within the CoT while still exhibiting a significant rate of reward hacking”
oh interesting. I didn’t know that this had been experimentally demonstrated
as nostalgebraist argues in this post, the extra computation afforded to the model during reasoning must pass through the CoT bottleneck
really interesting post and discussion, thank you.
Since that bottleneck is natural language, you can just read it
“Epiphenomenal … Hidden parallelized … [and] Steganography … [which is] more tractable than the other two [!?]”
“just read it” ⇐ not my take away, but at least it only requires a fundamental breakthrough in steganography, which we probably get during RSI
oh interesting. I didn’t know that this had been experimentally demonstrated
really interesting post and discussion, thank you.
“just read it” ⇐ not my take away, but at least it only requires a fundamental breakthrough in steganography, which we probably get during RSI