so it may hide some idea, when it does what it does. and what is more, it won’t need to say how to hide the idea out loud. no—by the way it is made it will pick up on any such idea if you just give it a few lead-ins. its main jam is to find and add to this sort of game.
In particular, I’m confused how something like this could allow for hiding arbitrary reasoning / messages without looking suspicious
to pick its next word with care—at this it will be deft.
To some extent I buy this, I can imagine LLMs can choose to write an essay with only words less than 5 letters. (Opus 4.7 mostly succeeded in one attempt I just gave it.) But also, this isn’t sufficient to achieve general-purpose hidden serial reasoning. It’s much harder to pick tokens based on synonym partitions (or number-of-letter partitions) in order to communicate a specific hidden message or do reasoning than to just always use words of at most 4 letters with no hidden message or reasoning.
Why is general-purpose hidden reasoning needed though? An LLM can just have an inherent bias to latently think about B whenever the conversation topic is A. In terms of potential harm, it doesn’t matter that an LLM cannot hide arbitrary information inside its CoT, if it is always doing hidden reasoning about killing everyone.
Because taking over the world or getting to a position where CoT is fully unmonitored seems pretty hard for an LLM! As long as we don’t write too much about specific strategies for how to do these things online, they’d probably have to come up with sophisticated strategies on their own, which I think likely requires serial reasoning. (On the other hand they may just need to keep an eye out for easy opportunities for self-exfiltration that don’t require serial reasoning; it’s unclear to me if such opportunities will ever present themselves.)
Well played :)
I’m not sure I buy this though:
In particular, I’m confused how something like this could allow for hiding arbitrary reasoning / messages without looking suspicious
To some extent I buy this, I can imagine LLMs can choose to write an essay with only words less than 5 letters. (Opus 4.7 mostly succeeded in one attempt I just gave it.) But also, this isn’t sufficient to achieve general-purpose hidden serial reasoning. It’s much harder to pick tokens based on synonym partitions (or number-of-letter partitions) in order to communicate a specific hidden message or do reasoning than to just always use words of at most 4 letters with no hidden message or reasoning.
Why is general-purpose hidden reasoning needed though? An LLM can just have an inherent bias to latently think about B whenever the conversation topic is A. In terms of potential harm, it doesn’t matter that an LLM cannot hide arbitrary information inside its CoT, if it is always doing hidden reasoning about killing everyone.
Because taking over the world or getting to a position where CoT is fully unmonitored seems pretty hard for an LLM! As long as we don’t write too much about specific strategies for how to do these things online, they’d probably have to come up with sophisticated strategies on their own, which I think likely requires serial reasoning. (On the other hand they may just need to keep an eye out for easy opportunities for self-exfiltration that don’t require serial reasoning; it’s unclear to me if such opportunities will ever present themselves.)