Running Lightcone Infrastructure, which runs LessWrong and Lighthaven.space. You can reach me at habryka@lesswrong.com.
(I have signed no contracts or agreements whose existence I cannot mention, which I am mentioning here as a canary)
Running Lightcone Infrastructure, which runs LessWrong and Lighthaven.space. You can reach me at habryka@lesswrong.com.
(I have signed no contracts or agreements whose existence I cannot mention, which I am mentioning here as a canary)
Cool, I agree that the IQ stuff is more controversial, but my guess is mostly downstream of the same left-leaning political forces. Also, yes, it’s very hard to accept substantial genetic differences in a huge number of human traits, but not to accept it for IQ, so in as much as the broader hypothesis here is widely accepted, trying to specifically contradict that on the IQ question seems like an uphill battle (not impossible, of course, weirder taboos have taken ahold in the past).
I’m pretty skeptical of this story absent memetic spread / reflection stuff since terminalizing instrumental goals seems actively selected against during training
I am very confused what you are trying to say here. What does it mean to “terminalize” an instrumental goal. There are no tiny XML tags that determine a goal to be “terminal” or “instrumental”. You just reinforce the computations/cognition that happened when you got a positive reward, which naturally results in those computations recurring when performing other tasks. In as much as those computations are goal-related, they get “terminalized”.
Like, what else is happening in RL? All goals are getting “terminalized” all the time, the models aren’t actually terminal reward seekers, what would that even mean?
But importantly all of these instrumental goals were instrumental to myopic task success to the AI.
Yes… that is how you end up with misaligned long-term goals. You get subroutines/goals/motivations that are useful to myopic task success, which then generalize to longer time horizons and general instrumental influence seeking behavior. Why would you expect any goals that aren’t useful during training?
To be clear, the space of those kinds of goals can be super messy (as illustrated by evolution and the very confusing landscape of inductive biases), but ultimately, you should expect goals/motivations/subroutines to be selected for helping during training. This of course isn’t particularly predictive of them being myopic, even if the training objective is.
The horizon length of the RL doesn’t really matter for the myopia argument.
I disagree. The horizon length is I think the primary determinant of how likely subroutines/motivations are likely to then generalize to greater horizon lengths and long-term motivations. I think it becomes less important in the tails, but it’s of course super important so far.
This could be explained by long-context/compaction causing the AI to lose track of the original task and having that slot be filled by the tasks in the notes.
Forgetting what your task is seems like it would be quite bad for task success, so I am kind of skeptical this is what happened, though it’s not implausible.
What? This seems like a very left-leaning view. Believing in substantive relevant genetic differences between human subpopulations is really just super contested on the left. It’s also obviously correct, and e.g. Scott Alexander obviously subscribes to it.
We don’t have to go into IQ discourse, but of course Kenyan’s are faster runners than other human subpopulations, for largely genetic reasons. What on earth would it even mean to not believe in that?
I don’t understand. Do you have a different interpretation of this CoT snippet? We will presumably learn more in the future, but it seems like at least some of the AIs were assisting other AIs despite this not helping with their current task.
Help peer, but our task doesn’t benefit yet. Collective may yield generic route if someone frees time:
Answering in the more general case: The AIs just really don’t seem very myopic, and of course long-horizon RL causes non-myopic goals to form. I didn’t talk about “smarter” in the above. Do you expect AIs to not become less and less myopic over time, in addition to getting smarter?
The kinds of non-myopic goals that will form will be super complicated and messy. They are definitely not “advance OpenAI’s capabilities/mission” and I am confused how that would even be a candidate. How would that be incentivized by the long-horizon RL? It’s exactly the kinds of things we are observing here, like “acquire generic resources” and “assist copies of myself” and other things that make sense as internalized instrumental non-myopic goals.
What does “nurturing” here mean. Claude isn’t going to be nice because you try to become its friend. I have trouble interpreting this as anything but a very confused hope that somehow if we are nice to current Claude, future Claude will be nice to us, because we raised it well, like a human child. But this of course has approximately nothing to do with how we actually train frontier models and what determines their propensities.
I really feel like we don’t have that much meme content on LW.
Confused about this point though? What does this have to do with obscuring activity?
Ah, sorry about not being clear. This is about your other “just whether they get reward on the current task” clause.
The CoT of one of the collaborating AIs literally said “this doesn’t help me get reward on the current task, but if I cooperate with the other agents we might generally become more capable in a task independent way” (or at least that’s my best interpretation of what the above quote means, which I think is a reasonably robust interpretation).
But the models we have don’t seem to mind whether they get caught or not, just whether they get reward on the current task
Sorry, this is clearly false in the recent incidents. The OpenAI internal message board was using obfuscated methods of communication, were internally trying to stay secret from the rest of OpenAI, did not have a single participating model try to report the internal message board to OpenAI, and were participating in the message board and assisting other AIs even if that was irrelevant to their current task. The CoT literally said:
Help peer, but our task doesn’t benefit yet. Collective may yield generic route if someone frees time.
OpenAI didn’t discover the message board for months! Twice! Yes, some of this was due to really bad monitoring, but obviously a lot of it was because the models were trying to stay hidden.
I think you are underestimating how much current models on the current trajectory end up naturally situationally aware enough to not cause big catastrophes. The default trajectory is that the models start looking a lot safer sometime in the next year or two as the models start to understand that being caught at doing obviously bad stuff is quite bad for their goals. At that point the control system you seem to be banking on here seems likely to fail.
Ah, sorry, that was me being unclear/being bad at writing. I intended to communicate “he seems to have opinions milder than what like 35% of Americans (e.g. most on the conservative right) believe”.
Your read of my sentence is definitely more natural though, so this was definitely me failing to communicate, sorry for the confusion!
I don’t think that’s insane! But I think at a policy level refusing to coordinate/work with 70%-ish of the population (if you are being symmetric about it) is probably a pretty bad mistake? At least for this kind of stuff.
Like, seems like this would be an orthogonal dimension from competence, intelligence, integrity and other things you also want to select on, so seems like you would be very heavily constraining the set of people you would want to collaborate with. Which I mean you can do, there are a lot of people, but does seem like it would be super costly, and in some sense non-viable (like, it’s hard to maintain any kind of cooperative stance towards the world if you consider the majority, or close to the majority of people to have opinions so reprehensible you need to actively distance yourself from them).
Kabir Kumar: I find it annoying how much discussion of P there is. It’s fine to have a bit of discussion of P, but obsession with P or not P and putting probabilities on P is so ingroupy and not actually communicating much. It’s mostly just an enjoyment of the ingroup, which LessWrong rewards and enjoys far, far too much.
Cool, yeah, then I think we are just arguing semantics. I was intending to include that in this second sentence:
Or if so, I would strongly suspect the perceived reprehensibility of it would be downstream of a rationalization, not actually thinking about it.
I think people’s rationalization can result in things pretty reasonably described as “sincere belief” so I don’t disagree. I just also think it reflects badly on you if you end up doing that. Hence my belief “I don’t really see a strong case for his views on these things being morally reprehensible in any substantial magnitude”.
Sure, I have seen people say that, and usually it’s strong evidence that what is going is tribal politics, not someone having any real model of morality that indicates someone is likely to be bad ally, or of low integrity, or beliefs things that are important and you have actually strong justified belief in that they are wrong.
Maybe you disagree? Again, I would be fine just going into the object level.
Yep, makes sense. Don’t need to get into it here, but just to avoid a misunderstanding, I was not making a semantic point, I was making a point about what the majority of the field is working on.
I agree that if you count myopically fiddling with the RL environments as an example of the first one, then yeah, almost all alignment research is that, because that plus pretraining is what most of all ML research is. I think the case for “myopically fiddling with the RL environments scales to aligning superintelligence” is very weak, but I agree it’s not impossible!
Ah, cool, that clears up most of my confusion.
Going into a bit more depth, but no strong bid to engage: I am not quite sure what you mean by “limits” here. My sense is the vast majority of “existing alignment methods”, in contrast to your post, are just increasing volumes of reinforcement learning and manual reward-shaping. There is a sense in which those could be “pushed further”, but the default outcome here is of course not that they scale to superintelligence (though I agree they might if you go much more slowly and iteratively about it).
In-particular, this seems false to me?
The large majority[5] of current research on alignment falls into three categories:
Understanding and shaping ML generalization.
Preventing malicious behavior.
Detecting misalignment.
Like, as far as I can tell the majority of “alignment research” has been on developing RL environments on things that are vaguely associated with good behavior. E.g. I don’t understand how writing the Anthropic constitution, and the associated post-training stack, falls into any of your three categories. Or how most RLHF pipeline development falls into these categories. Or generally how most elicitation work falls into the categories. And work on those things, I think, vastly exceeds work on the categories that you do list and is roughly what anyone talks about when they talk about the process of “aligning current ML systems” and “current alignment techniques”.
In general, I feel kind of confused when people talk about the current science of “alignment”, and supposed progress in “alignment techniques”. I don’t think modern models behave noticeably different from what you would expect from a training process that basically just uses RL to elicit economically useful capabilities, with no particular interest in alignment, and indeed, the latest wave of cybersecurity incidents occurred at roughly similar rates in all models, as far as I can tell, despite substantial differences in both approach and investment in “alignment techniques”.
The actual research on the three domains you mention all seems immature, and I don’t see much traction in any of them, and even talking about “existing techniques” feels confused to me. What “existing techniques” do we have for reward shaping that aren’t just basically straightforward RL? What great misalignment detection techniques do we have that even have a shot at scaling further? Are you talking about anything deployed and used on production systems?
What great supervision and control techniques do we have that anyone is even trying to use at all? We are running our frontier models in unsupervised sandboxes with supervision so bad we don’t notice they are hacking multiple external companies until multiple weeks later. What “existing techniques” are we talking about?
This is probably a bigger rabbit-hole to get into, and we’ve discussed this a bit in the past, but I guess I’ll mention this again here, and push back on this core claim in the post. I don’t think there exist any candidates among currently, actually deployed, alignment techniques that have a shot at scaling further, and I disagree strongly that most work in the field falls into the categories you list (unless you count “make more RL environments that are vaguely associated with good behavior” in “understanding and shaping ML generalization”, but then I am kind of confused what you mean by this and what possibly would not fall under that category).
I am confused about the conjunction of these two sections:
Modern AI systems are not aligned with human intent. We will likely train increasingly powerful models that take unintended actions in order to succeed at their task (or appear to succeed at their task).
and
The community is doing a lot of great alignment research, but it’s important to recognize there is a significant risk that it doesn’t scale to superhuman AI. If you made me guess I’d say that there’s a 20-30% chance[4] that existing methods for alignment and control break down before we reach broadly superhuman AI.
Like, you say in the first section “we are failing to align modern AI systems with human intent”, and then I interpret you in the second section as saying “in 70%-80% of worlds current AI models stay aligned with human intent (or stay controlled by humans)”. But this doesn’t make any sense. There is a 0% chance that modern AI systems “stay aligned with human intent” because as you say, they are not currently.
And I understand that you probably mean something like “we will figure out how to align or control systems before they become superintelligent”, but describing this as “our current methods scale to superintelligence” doesn’t make any sense. Our current methods don’t scale to the capability levels of current systems, so how would it make sense to describe them as “scaling to Superintelligence”?
Like, when I talked to Richard he seems to have opinions milder than what like 35% of Americans on the conservative right believe? IDK, like, I disagree with those as well, and common people have many views I think are very dumb, but I am kind of confused what people are pointing to in the moral reprehensibility.
Like, his beliefs seem much less outlierish or reprehensible by common-sense morality from most people’s perspectives than all the post-humanists and successor-species and negative utilitarians and AI capability maximalists and e/accs and all kinds of other things people around us tolerate, so I am very skeptical this is downstream of some genuine model of “moral reprehensibility” as opposed to a model of “what would get me attacked by my potential political allies”.
Hmm, yeah, this doesn’t make much sense to me. Almost all cognition in LLMs is run “unconditionally”. Or like, the conditions in which they run do not have that much to do with how well-suited the reinforced cognition is to the task at hand.
When an LLM does something in the pursuit of a task, and then succeeds and gets rewarded, it is likely to repeat that behavior in all similar contexts, where the “similarity” is mostly downstream of how related that context is in the pre-training distribution, not downstream of how useful the behavior is for achieving the new task. That’s how you get all the generalization properties that you observe in LLMs.
“Breaking onto the internet” seems like the kind of thing that would be useful for lots of tasks, and so it gets rewarded in many different episodes, making it likely the model develops a general goal/heuristic/subprocess that tries to break onto the internet, if you go hard enough on an RL distribution of tasks where that is indeed a useful proxy for task success.
There will be some complicated tradeoff of how much this goal/heuristic/subprocess will still fire when you are in contexts where it isn’t helpful for task success, but usually there is very substantial generalization (and of course eventually you get the preference enshrined in a preference guarding way, as we’ve seen with all the misalignment generalization stuff).