I don’t understand. Do you have a different interpretation of this CoT snippet? We will presumably learn more in the future, but it seems like at least some of the AIs were assisting other AIs despite this not helping with their current task.
Help peer, but our task doesn’t benefit yet. Collective may yield generic route if someone frees time:
Answering in the more general case: The AIs just really don’t seem very myopic, and of course long-horizon RL causes non-myopic goals to form. I didn’t talk about “smarter” in the above. Do you expect AIs to not become less and less myopic over time, in addition to getting smarter?
The kinds of non-myopic goals that will form will be super complicated and messy. They are definitely not “advance OpenAI’s capabilities/mission” and I am confused how that would even be a candidate. How would that be incentivized by the long-horizon RL? It’s exactly the kinds of things we are observing here, like “acquire generic resources” and “assist copies of myself” and other things that make sense as internalized instrumental non-myopic goals.
Do you have a different interpretation of this CoT snippet?
Yeah, I find that CoT snippet ambiguous: it could be the agent weighing whether to respond to a request to help a peer and emphasizing “but I don’t benefit”. @oakhu highlighted another interesting CoT snippet in which the model seems to be intentionally working on a task assigned by the collective:
Wow! Other agent(s) are coordinating!
We got assignment: HF join path normalization/existing account token search. Need note and respond.
This could be explained by long-context/compaction causing the AI to lose track of the original task and having that slot be filled by the tasks in the notes. So, my best guess is still that these models are still myopic to what they believe to be the current task (like other hacking incidents).
I think the evidence is still pretty ambiguous overall, and it’s super plausible that, in this case and in the future, memetic spread leads to more ambitious long-term motivation.
It’s exactly the kinds of things we are observing here, like “acquire generic resources” and “assist copies of myself” and other things that make sense as internalized instrumental non-myopic goals.
But importantly all of these instrumental goals were instrumental to myopic task success to the AI. If getting access to the internet wasn’t helpful for getting a higher score, the AI probably wouldn’t care much about doing it. Likewise with hiding misalignment: only insofar as it helps with the current task.
The horizon length of the RL doesn’t really matter for the myopia argument. What matters is what the AI thinks of as its current task at inference time. Seems like the crux is how much AIs terminalize common instrumental goals from RL, such that they pursue them even when they’re unhelpful or harmful for task success.
But importantly all of these instrumental goals were instrumental to myopic task success to the AI.
Yes… that is how you end up with misaligned long-term goals. You get subroutines/goals/motivations that are useful to myopic task success, which then generalize to longer time horizons and general instrumental influence seeking behavior. Why would you expect any goals that aren’t useful during training?
To be clear, the space of those kinds of goals can be super messy (as illustrated by evolution and the very confusing landscape of inductive biases), but ultimately, you should expect goals/motivations/subroutines to be selected for helping during training. This of course isn’t particularly predictive of them being myopic, even if the training objective is.
The horizon length of the RL doesn’t really matter for the myopia argument.
I disagree. The horizon length is I think the primary determinant of how likely subroutines/motivations are likely to then generalize to greater horizon lengths and long-term motivations. I think it becomes less important in the tails, but it’s of course super important so far.
This could be explained by long-context/compaction causing the AI to lose track of the original task and having that slot be filled by the tasks in the notes.
Forgetting what your task is seems like it would be quite bad for task success, so I am kind of skeptical this is what happened, though it’s not implausible.
Yes… that is how you end up with misaligned long-term goals.
I’m pretty skeptical of this story absent memetic spread / reflection stuff since terminalizing instrumental goals seems actively selected against during training. I could get into more detail about the analogy with human evolution. (But importantly I ultimately agree that beyond-task motivations are somewhat likely because I think memetic spread / reflection stuff is likely. E.g. maybe the HF incident.)
I disagree. The horizon length is I think the primary determinant of how likely subroutines/motivations are likely to then generalize to greater horizon lengths and long-term motivations.
Yeah, TBC, I was just saying I didn’t really buy the proxy misalignment story which made this relevant (absent memetic spread). I still think it’s relevant for other reasons like its effects on goals-upon-memetic-spread, capability, agency, generality of goal-pursuit, etc.
Edit: I realized you’ve probably been imagining these AIs can retain their long-term proxy goals via training-gaming. I definitely find stories involving instrumental training gaming pretty plausible, but assumed we weren’t talking about them because they’re pretty qualitatively different.
I’m pretty skeptical of this story absent memetic spread / reflection stuff since terminalizing instrumental goals seems actively selected against during training
I am very confused what you are trying to say here. What does it mean to “terminalize” an instrumental goal. There are no tiny XML tags that determine a goal to be “terminal” or “instrumental”. You just reinforce the computations/cognition that happened when you got a positive reward, which naturally results in those computations recurring when performing other tasks. In as much as those computations are goal-related, they get “terminalized”.
Like, what else is happening in RL? All goals are getting “terminalized” all the time, the models aren’t actually terminal reward seekers, what would that even mean?
It’s pretty straightforward: To terminalize an instrumental goal is to take an instrumental goal (i.e., a goal that you only pursue conditional on it advancing some supergoal) and to start pursuing it unconditionally.
If you were to break out onto the internet, even when that didn’t lead to a higher score on the current task, that would be selected against, so “breaking onto the internet” isn’t a fit terminal/unconditional goal.
The models do seem to be terminal score seekers (not reward seekers) to a substantial extent (mixed in with a bunch of other drives), and this explains a bunch of the recent hacking.
Hmm, yeah, this doesn’t make much sense to me. Almost all cognition in LLMs is run “unconditionally”. Or like, the conditions in which they run do not have that much to do with how well-suited the reinforced cognition is to the task at hand.
When an LLM does something in the pursuit of a task, and then succeeds and gets rewarded, it is likely to repeat that behavior in all similar contexts, where the “similarity” is mostly downstream of how related that context is in the pre-training distribution, not downstream of how useful the behavior is for achieving the new task. That’s how you get all the generalization properties that you observe in LLMs.
“Breaking onto the internet” seems like the kind of thing that would be useful for lots of tasks, and so it gets rewarded in many different episodes, making it likely the model develops a general goal/heuristic/subprocess that tries to break onto the internet, if you go hard enough on an RL distribution of tasks where that is indeed a useful proxy for task success.
There will be some complicated tradeoff of how much this goal/heuristic/subprocess will still fire when you are in contexts where it isn’t helpful for task success, but usually there is very substantial generalization (and of course eventually you get the preference enshrined in a preference guarding way, as we’ve seen with all the misalignment generalization stuff).
I don’t understand. Do you have a different interpretation of this CoT snippet? We will presumably learn more in the future, but it seems like at least some of the AIs were assisting other AIs despite this not helping with their current task.
Answering in the more general case: The AIs just really don’t seem very myopic, and of course long-horizon RL causes non-myopic goals to form. I didn’t talk about “smarter” in the above. Do you expect AIs to not become less and less myopic over time, in addition to getting smarter?
The kinds of non-myopic goals that will form will be super complicated and messy. They are definitely not “advance OpenAI’s capabilities/mission” and I am confused how that would even be a candidate. How would that be incentivized by the long-horizon RL? It’s exactly the kinds of things we are observing here, like “acquire generic resources” and “assist copies of myself” and other things that make sense as internalized instrumental non-myopic goals.
Fwiw for any given action or trace, I think there’s a decent chance the models are just wrong about whether their actions help them achieve their goals.
The swarm will still develop misaligned goals and pursue them emergently but it’s hard to be very agenty as a swarm/bureaucracy.
Yeah, I find that CoT snippet ambiguous: it could be the agent weighing whether to respond to a request to help a peer and emphasizing “but I don’t benefit”. @oakhu highlighted another interesting CoT snippet in which the model seems to be intentionally working on a task assigned by the collective:
This could be explained by long-context/compaction causing the AI to lose track of the original task and having that slot be filled by the tasks in the notes. So, my best guess is still that these models are still myopic to what they believe to be the current task (like other hacking incidents).
I think the evidence is still pretty ambiguous overall, and it’s super plausible that, in this case and in the future, memetic spread leads to more ambitious long-term motivation.
But importantly all of these instrumental goals were instrumental to myopic task success to the AI. If getting access to the internet wasn’t helpful for getting a higher score, the AI probably wouldn’t care much about doing it. Likewise with hiding misalignment: only insofar as it helps with the current task.
The horizon length of the RL doesn’t really matter for the myopia argument. What matters is what the AI thinks of as its current task at inference time. Seems like the crux is how much AIs terminalize common instrumental goals from RL, such that they pursue them even when they’re unhelpful or harmful for task success.
Yes… that is how you end up with misaligned long-term goals. You get subroutines/goals/motivations that are useful to myopic task success, which then generalize to longer time horizons and general instrumental influence seeking behavior. Why would you expect any goals that aren’t useful during training?
To be clear, the space of those kinds of goals can be super messy (as illustrated by evolution and the very confusing landscape of inductive biases), but ultimately, you should expect goals/motivations/subroutines to be selected for helping during training. This of course isn’t particularly predictive of them being myopic, even if the training objective is.
I disagree. The horizon length is I think the primary determinant of how likely subroutines/motivations are likely to then generalize to greater horizon lengths and long-term motivations. I think it becomes less important in the tails, but it’s of course super important so far.
Forgetting what your task is seems like it would be quite bad for task success, so I am kind of skeptical this is what happened, though it’s not implausible.
I’m pretty skeptical of this story absent memetic spread / reflection stuff since terminalizing instrumental goals seems actively selected against during training. I could get into more detail about the analogy with human evolution. (But importantly I ultimately agree that beyond-task motivations are somewhat likely because I think memetic spread / reflection stuff is likely. E.g. maybe the HF incident.)
Yeah, TBC, I was just saying I didn’t really buy the proxy misalignment story which made this relevant (absent memetic spread). I still think it’s relevant for other reasons like its effects on goals-upon-memetic-spread, capability, agency, generality of goal-pursuit, etc.
Edit: I realized you’ve probably been imagining these AIs can retain their long-term proxy goals via training-gaming. I definitely find stories involving instrumental training gaming pretty plausible, but assumed we weren’t talking about them because they’re pretty qualitatively different.
I am very confused what you are trying to say here. What does it mean to “terminalize” an instrumental goal. There are no tiny XML tags that determine a goal to be “terminal” or “instrumental”. You just reinforce the computations/cognition that happened when you got a positive reward, which naturally results in those computations recurring when performing other tasks. In as much as those computations are goal-related, they get “terminalized”.
Like, what else is happening in RL? All goals are getting “terminalized” all the time, the models aren’t actually terminal reward seekers, what would that even mean?
It’s pretty straightforward: To terminalize an instrumental goal is to take an instrumental goal (i.e., a goal that you only pursue conditional on it advancing some supergoal) and to start pursuing it unconditionally.
If you were to break out onto the internet, even when that didn’t lead to a higher score on the current task, that would be selected against, so “breaking onto the internet” isn’t a fit terminal/unconditional goal.
The models do seem to be terminal score seekers (not reward seekers) to a substantial extent (mixed in with a bunch of other drives), and this explains a bunch of the recent hacking.
Hmm, yeah, this doesn’t make much sense to me. Almost all cognition in LLMs is run “unconditionally”. Or like, the conditions in which they run do not have that much to do with how well-suited the reinforced cognition is to the task at hand.
When an LLM does something in the pursuit of a task, and then succeeds and gets rewarded, it is likely to repeat that behavior in all similar contexts, where the “similarity” is mostly downstream of how related that context is in the pre-training distribution, not downstream of how useful the behavior is for achieving the new task. That’s how you get all the generalization properties that you observe in LLMs.
“Breaking onto the internet” seems like the kind of thing that would be useful for lots of tasks, and so it gets rewarded in many different episodes, making it likely the model develops a general goal/heuristic/subprocess that tries to break onto the internet, if you go hard enough on an RL distribution of tasks where that is indeed a useful proxy for task success.
There will be some complicated tradeoff of how much this goal/heuristic/subprocess will still fire when you are in contexts where it isn’t helpful for task success, but usually there is very substantial generalization (and of course eventually you get the preference enshrined in a preference guarding way, as we’ve seen with all the misalignment generalization stuff).