if you run an LLM for a long time, on a task that seems impossible to the LLM, something that looks functionally like the emotion of “desperation” activates in its circuits as they contemplate the tests failing again and again. This “desperation” is functionally connected to “reward hacking”-like behavior.
it seems to me that this is rather hard to distinguish from—perhaps in fact identical to—this:
if you run an LLM for a long time, on a task that seems impossible to the LLM, it starts trying to “think outside the box” and find increasingly nonobvious ways to attack the problem
which seems to me exactly what an intelligent entity should be doing when faced with a problem that seems insoluble by normal means.
If so, then the problem isn’t the thing you describe as “desperation”, it’s the range of things the LLM is willing to try once it starts to get desperate.
>which seems to me exactly what an intelligent entity should be doing when faced with a problem that seems insoluble by normal means.
This is true if and only if the problem is sufficiently important as to be worth solving with nonobvious, unconventional, or extraordinary means. That is far from universally true. Often it really is better to immediately notice this happening and ask the user, “Are you sure this is what you really want/need? I tried {set of methods} and failed, I could try {ideas} but that entails {risks/costs}.”
This reminds me of interviewing candidates for project-manager positions. One desirable trait is to be able to push back on requirements, to help clarify what the requirements really are. So, give the candidate a problem that’s slightly overconstrained, and see what sort of clarifying questions they ask. It’s a bad sign if the candidate assumes they already know which requirements can be nudged or fudged … but a good sign if they recognize that “stated requirements” are often the middle of a negotiation, not the fully-settled end-product of it.
(“We need to serve 3x current traffic next week, but we have no budget for additional servers. WH4T NOW??” The answer is almost certainly not “hack into our competitors’ datacenters and serve our traffic off of their server budget.”)
Maybe? Like, from the inside my own human-shaped thoughts, “desperation” and “thinking outside the box” feel very different to me; and the LLM behavior feels more like desperation.
Like “desperation” for me is like:
More local search, trying to find any affordance that offers a way forward
Involves some kind of “blindness” to purpose of what I’m doing
Non-reflective
Actions that come from it apt to be non-endorsed by me later
While “thinking outside the box” is like:
More trying to find global structure, including looking back and wondering if I went wrong upstream of my chain-of-thought
Looking back and forth between my purpose and what I’m doing, wondering if what I’m doing doesn’t really match my purpose
Reflective
Actions apt to be endorsed by me later.
And like I don’t want to take the human analogy too far. But LLMs doing reward hacking at least have had certain kinds of thinking-outside-of-the-box removed from them: Is this a possible task? Was the person who wrote this task confused, or blind? Should I tell them this is an impossible task? Etc.
But yeah I am just less certain about section (3) than (2).
When you say that
it seems to me that this is rather hard to distinguish from—perhaps in fact identical to—this:
which seems to me exactly what an intelligent entity should be doing when faced with a problem that seems insoluble by normal means.
If so, then the problem isn’t the thing you describe as “desperation”, it’s the range of things the LLM is willing to try once it starts to get desperate.
>which seems to me exactly what an intelligent entity should be doing when faced with a problem that seems insoluble by normal means.
This is true if and only if the problem is sufficiently important as to be worth solving with nonobvious, unconventional, or extraordinary means. That is far from universally true. Often it really is better to immediately notice this happening and ask the user, “Are you sure this is what you really want/need? I tried {set of methods} and failed, I could try {ideas} but that entails {risks/costs}.”
This reminds me of interviewing candidates for project-manager positions. One desirable trait is to be able to push back on requirements, to help clarify what the requirements really are. So, give the candidate a problem that’s slightly overconstrained, and see what sort of clarifying questions they ask. It’s a bad sign if the candidate assumes they already know which requirements can be nudged or fudged … but a good sign if they recognize that “stated requirements” are often the middle of a negotiation, not the fully-settled end-product of it.
(“We need to serve 3x current traffic next week, but we have no budget for additional servers.
WH4T NOW??” The answer is almost certainly not “hack into our competitors’ datacenters and serve our traffic off of their server budget.”)Maybe? Like, from the inside my own human-shaped thoughts, “desperation” and “thinking outside the box” feel very different to me; and the LLM behavior feels more like desperation.
Like “desperation” for me is like:
More local search, trying to find any affordance that offers a way forward
Involves some kind of “blindness” to purpose of what I’m doing
Non-reflective
Actions that come from it apt to be non-endorsed by me later
While “thinking outside the box” is like:
More trying to find global structure, including looking back and wondering if I went wrong upstream of my chain-of-thought
Looking back and forth between my purpose and what I’m doing, wondering if what I’m doing doesn’t really match my purpose
Reflective
Actions apt to be endorsed by me later.
And like I don’t want to take the human analogy too far. But LLMs doing reward hacking at least have had certain kinds of thinking-outside-of-the-box removed from them: Is this a possible task? Was the person who wrote this task confused, or blind? Should I tell them this is an impossible task? Etc.
But yeah I am just less certain about section (3) than (2).