We see some good news in alignment—as models become more capable, they are also more aligned
I find it very scary that senior alignment researchers apparently don’t understand the difference between detected misalignment and actual misalignment.
Do you believe current models have large amounts of undetected misalignment? I believe the trend that more capable models are also more aligned (see Leike’s blog I linked) is not just in evaluations but also people’s observations in their real world usage.
Do you believe current models have large amounts of undetected misalignment?
I believe that there’s no way to reliably tell. I think “people’s observations in their real world usage” is also the sort of evidence that can easily fail to detect misalignment even if it exists.
I predict that ASI trained using anything resembling current techniques would be catastrophically misaligned. I don’t have a strong prediction about whether current-gen models are “actually” aligned. If I had to guess, I’d say they’re not. This is very speculative but my guess would be that they have misaligned internal drives/goals that would produce bad consequences if they were smart enough to reason through the implications of their goals, but they’re not smart enough to do that (or it might be more accurate to say that their long-term planning isn’t good enough).
What I will say is I don’t think it’s reasonable to believe with (say) >90% confidence that current-gen models are “actually” aligned, because our understanding of what’s going on with LLMs just isn’t that reliable.
I should mention that I don’t keep up with the vast majority of alignment research. There might be some convincing research about why the lack-of-detected-misalignment means there is genuinely no misalignment, and I just missed it; but my guess is if that research existed, then I would’ve heard about it (e.g. it would’ve gotten tons of upvotes on LessWrong, b/c that would be a really important finding). So when Leike claims that models are becoming more aligned, most of my subjective probability for why he said that is that he doesn’t understand the difference between detected misalignment and actual misalignment.
AI psychosis. When you ask LLMs point blank if it’s bad to induce AI psychosis, they say yes, and (AFAIK) there is no evidence that they’re being deceptive, and yet they do it anyway. This is a case where (IIRC) recent models have gotten better (less psychosis-prone), but my guess is that training psychosis out of LLMs doesn’t generalize to other forms of misalignment, and maybe in the next-gen models some new form of bad behavior will emerge.
IMO the best evidence of bad behavior by LLMs is coming out of Palisade, not any AI company. This suggests that AI companies (who have way more resources and access) are not trying sufficiently hard to detect misalignment. The implication is that, if AI companies fail to detect misalignment, this is only weak evidence that the misalignment isn’t there.
I think current AI systems are likely catastrophically misaligned, but instead of properly arguing for it here, I want to clear the much lower bar of making the position sound much less weird than it might at first. When I imagine a person to whom this position sounds weird, I imagine them saying sth like:
“AIs are acting nicely in various contexts. They look nice in our evaluations, and they look nice to users in practice. Isn’t it unlikely that they are really evil, hiding it, waiting to strike?”
While I think it’s likely that current AI systems are catastrophically misaligned, I don’t much feel like taking a position one way or the other about the “are really evil, hiding it, waiting to strike” part. I think the hypothetical interlocutor above is making a false equivalence. When I say current AI system are “catastrophically misaligned”, what I have in mind is this:
When someone sets up some initial AI system and lets it develop a lot (ie lets it do RSI)[1], with anything like this that could be done in practice[2], this doesn’t go well for humans. I think that the default without strict regulation of AI development is: in the first 10 years after AGI (by which I mean AI that autonomously does conceptual research better than top humans), there will be a lot of development — like probably more development than there has been in total in all of history.[3] Like, after developing for a lot of “subjective time”, the AI systems that come out of this development process would trivially be able to replace humans with whatever other processes from some vast number of options; the negentropy/[free energy]/atoms I’m currently using could probably be used to run processes of similar complexity/interestingness. Despite it being trivial for the AI to do this, the AI needs to not do this (or, maybe disassemble me, but at least recreate me on a computer, I guess...). In fact, the AI doesn’t just need to leave me alone, it needs to protect me from being killed by any other beings, and make sure I have a bunch of resources so I can live a long life. It’s kinda like I need to be very close to the coolest possible process to this AI, despite being “objectively” extremely boring, slow, wasteful, with “objectively” nothing to offer to the AI. This seems like a really sharp property; it feels like a measure 0 sort of thing. Preserving this forever feels especially sharp. I think it’s unlikely that this property would be upheld. I don’t think it is that reassuring if this long development process is started by AIs whose cached policies for mundane situations are pretty nice(-looking).[4]
Maybe this at least makes it seem not weird to think that current AI systems are catastrophically misaligned. It’s plausible we’re just using the same words differently, but in that case I think my use better tracks the niceness-type property that really matters. Like, it ultimately matters whether our AIs will continue to protect us forever when everything is up to them, not whether they behave nicely in mundane interactions now. I guess the terms “catastrophic/egregious misalignment” or “a large amount of misalignment” are quite unfortunate because it’s sort of unclear if one should read them as [misalignment sufficient for things to end up being really bad] (in that case, given doomy views, even an extremely small failure to set valuing up properly constitutes catastrophic/egregious/large misalignment, and it’s plausible to me that of humans are egregiously misaligned by default, tho I’m not sure[5]) or as [the AI wanting to behave egregiously badly in mundane circumstances]. I think that there being these two really different interpretations of the same term has caused a bunch of confused thinking by people in alignment.
this could be framed as asking the AI to develop a good successor; the initial setup might have some processes tasked with “solving alignment”; there might be multiple AIs involved doing different things, eg there can be monitors
My guess is also that things will also naively be looking worse once we get to AIs that are actually able to do research autonomously, because these AIs will be less based on human imitation, they will be actually able to come up with new thinky-stuff (new words/concepts/ideas/methods etc), they will not have nice chains of thought, and they will be more trained on clearly inhuman things like doing math/coding/science/tech.
I agree they’re “egregiously misaligned” in this sense, but it’s also the case that this usage of the word goes very much against the grain of common usage.
The term “AI alignment” was originally meant to refer to AGI-ish/ASI-ish AIs. So, if one wants to extrapolate it to “lesser AIs”, extrapolating it to either one of “well-behaving sub-AGI-ish/sub-ASI-ish AI” or “sub-AGI-ish/sub-ASI-ish AI that produces aligned AGI/ASI if one seeds an RSI with it” seems fine/valid, at least in isolation. Most people went for the former; you’re arguing for the latter, I think, largely because those who went for the former generally tend to be inclined to think that the former somewhat strongly implies the latter, and the latter is what matters in the long run (if something RSI’s into AGI/ASI).
Initially, I was going to say that I’m pessimistic about you/someone managing to change how people think about/understand “alignment” in this way (e.g., because it implies that most humans are “egregiously misaligned”, as you say it yourself), but on some thought, I’m not sure. Pushing back in this way and insisting that “this is the meaning of ‘alignment’ that matters and that your meaning of ‘alignment’ does meaningfully imply it” might be productive for shifting people’s attention to where it matters.
either one of “well-behaving sub-AGI-ish/sub-ASI-ish AI” or “sub-AGI-ish/sub-ASI-ish AI that produces aligned AGI/ASI if one seeds an RSI with it” seems fine/valid, at least in isolation. Most people went for the former … those who went for the former generally tend to be inclined to think that the former somewhat strongly implies the latter, and the latter is what matters in the long run
There is a very popular framing coloring all thinking of some people where seriously engaging with technological developments that are not immediately actionable is seen as deeply unvirtuous, and so the thought is never allowed proper consideration. Future that is not immediate is the immediate future’s responsibility, not your current self’s responsibility, and it’s irresponsible to be seriously concerned with it over the immediately actionable things you are working on, that you are directly affecting and need to get right.
Thus observable “alignment” of modern AIs, in the sense of their good behavior, is not just a reasonable disambiguation of “alignment”, but the only one permitted by this stance. Being inclined to think that this helps in the long term doesn’t influence the outcome of seriously thinking only about current behavior. The claim that only long term consequences of behavior under RSI and society-scale development is what ultimately matters is not permitted to be taken seriously, it’s not the background assumption that justifies the focus on current behavior of modern AIs.
It’s not that such people don’t believe ASI is coming, or that it’s coming in their own lifetime, but the epistemic distortion of seeing serious engagement with unactionable things as intolerably unvirtuous makes their thinking and behavior indistinguishable from that of people who really believe ASI can never happen. This distortion can be pierced by belief that ASI is imminent, but once it’s plausibly a few years away it could as well be pure fiction. Exploratory engineering might also be helpful for detailed engagement, where assumptions of a thought experiment permit thinking. But outside the thought experiments these assumptions are then not going to be taken seriously as gesturing at the actual future that is virtuous to engage with as actual future.
I definitely want people to think more about what AIs would think and do over a lot of reflection/development, and when more powerful. People should think more about the effects of a mind. People should think of the AGI situation as us probably having to correctly determine the future via an extremely long causal chain.[1]
I don’t think it’s weird to speak of values the way I’m speaking of values. I think people accept this sort of value-talk in other contexts. E.g. it’s common for antirealists to think of ethical truths as being determined by some ideal reflection; e.g. the notion of CEV. I think people who in some contexts use “egregious misalignment” in this “egregious misbehavior in mundane situations” sense also sometimes make inferences as if they were using “misalignment” in the sense I suggest. That said, one could want to make a distinction between reflection and development-in-general, and certainly it makes sense to distinguish between more and less endorsed forms of development. I think I was somewhat sloppy with this in my first comment.
I think it’d in principle be fine for some ideal beings to use words however. In practice, [people are stupid]/[thinking is difficult], and it’s very natural to make the inference “the AI is egregiously misaligned” “the AI wants to egregiously misbehave in normal circumstances” and also to make the inference “the AI endorses each step of a process which leads to all humans dying” “it was egregiously misaligned”, but I think there isn’t a concept that supports both of these inferences at once (or at least I think our language should leave this as an open question). So, I mostly don’t endorse using “catastrophic/egregious/large misalignment”, and trying to say what one means in other words. I should maybe have used different words in my first comment as well. I don’t have good alternative terms to suggest atm, except saying what one means with more words. I guess I’d want more people to try spending some time thinking about the AI situation while tabooing a bunch of Constellation-speak and MIRI-speak, building up their own Entish.
Some people think they can avoid this difficulty by having a first mess-AI “solve alignment” and launch some sort of aligned ASI sovereign, with the first AI not being that weird. I think that to first order one should think of this as the original AI trying to determine the future via a bottleneck. And in real life, people would plausibly just let the AI self-improve with some monitoring lol, in which case it’s not exactly a tight bottleneck. The original AI will also already be doing a lot of reflection and development. Also, there will be a long chain of causality after the ASI sovereign that needs to go right. (Also, in practice, instead of some clever scheme with boxed AIs solving alignment, we will probably just get some total mess with AIs deployed broadly, connected to the internet, plausibly just running AI labs. And there’s the AIs breaking out, and there’s fooming being fast, and there’s not having much time to be careful.)
“Catastrophically misaligned” and “catastrophically misaligned if we give them RSI capabilities far beyond what we can currently give them” are two very, very different claims, in my eyes.
I do appreciate you articulating that your “catastrophically misaligned” is a shorthand for the latter though.
I think that if we try to make sense of “what a current AI would do after reflecting+developing for a long time”, that thing does not involve being nice to humans. I think it’s still not nice to humans if we add the constraint “and the reflection/development process has to be basically [endorsed by the AI]/[good according to the AI]”. I think it’s pretty standard to take what you would do [after a lot of reflection + if you were more powerful] to reflect your values better than what you would do instinctively. So, if I’m right about what would happen given further (self-endorsed) development, it seems like a standard use of language (at least in alignment and in philosophy) + true to say current AIs are bad? I’d agree it is also pretty standard + [maybe true] to say “current AIs are good” in the sense that they mostly have pretty acceptable instinctive behaviors. This situation is pretty unfortunate, and maybe calls on us to start explicitly making this distinction.[1]
“Catastrophic misalignment” is a bad term, in addition to the reason I already gave in my comment, also because it could mean that this AI in fact would cause a catastrophe (without human help), which I don’t think is true for current AIs. That said, I think that’s prevented by capabilities, not by alignment — I think the closest thing to a current AI which is capable of causing a catastrophe would cause a catastrophe. I guess maybe one should say “misalignment sufficient for a catastrophic outcome if choosing the future were handed to the AI”.
much easier to prefix the term than to change it. it sounds like you’re describing either superintelligence alignment (pass the threshold of working for any superint), or asymptotic alignment (an even more difficult threshold of being reliably known to continue working more or less indefinitely). achieving asymptotic alignment would require some form of knowing that the system would, in an ongoing way, continue to improve its ability to check in with us without breaking us, and use that information in ways we know are valid extrapolations according to what we want. which sounds like what you’re describing, but importantly only gets its qualitative difference from the iteratedness. local alignment is still alignment, then, and that makes a lot of sense, since currently we train ais with locally linear-ish methods.
I believe the trend that more capable models are also more aligned
But the main problem has absolutely never been that models below human capabilities would be impossible to align. As made clear in the Superalignment announcement blog post the concern is that our current techniques won’t scale beyond this. This is clear from Concrete Problems In AI Safety, W2S’s entire agenda, etc.
Not the OP and not an alignment researcher, but I would appreciate an elaboration. What types of evidence would you consider relevant for concluding that an AI system is (roughly) aligned, as opposed to merely being a system for which we have not yet detected misalignment?
I don’t know what evidence would be sufficient to conclude that an AI system is aligned. That’s part of the problem—we have no way of reliably detecting misalignment or ruling it out.
(Well, I can think of some evidence that would convince me, but it’s not useful. E.g. if we build an ASI and it runs for a hundred years without killing everyone, then I’m pretty sure it’s aligned. But I would recommend against running that experiment.)
Ok this is a bit of a tortured analogy but imagine it’s the 1600s. We learn that if the mass of an iron atom is greater than 10⁻²³ grams, everyone dies (for some contrived reason). So we need to figure out the mass of iron. AI companies shave off the smallest scrap of iron they can manage, and put it on an era-appropriate balance scale, and the scale reads zero. And they declare, see! Iron is so light that it reads as zero!
Meanwhile I’m over here like, how do we know the scale is sensitive enough to measure the mass of a single iron atom? How do we know that that tiny little shaving only contains a single atom? What if it’s actually several atoms, and it’s possible to have an even smaller quantity of iron?
(Plus there are questions we don’t even know we’re supposed to be asking, like “which isotope of iron is this?”)
In that scenario, if you ask me what evidence would convince me that an iron atom weighs less than 10⁻²³ grams, I wouldn’t know what to tell you.
I feel like that’s where we are currently with alignment research.
(This analogy doesn’t capture the fact that measurable AI alignment is improving over time according to benchmarks. It’s also over-generous in that we could prove that our scales aren’t sensitive enough to detect a difference of 10⁻²³ grams. I don’t think we can do anything analogous in the case of AI alignment.)
Folk concerned with AI x-risk routinely cite various forms of behavior observed in frontier AI models as evidence that the systems are, or will be, dangerously misaligned. For instance, less than 24 hours ago, MIRI released this short video, which claims that “today’s AI systems lie to their developers, exploit loopholes in their instructions, and resist being shut down.” Yet, from what you say, it appears that this evidence is mostly irrelevant, because “we have no way of reliably detecting misalignment or ruling it out”.
We can detect misalignment under some conditions but we don’t know how to reliably detect it, i.e. avoid false negatives, i.e., we have no way of saying “if a system is misaligned in any way, then method XYZ will definitely detect it.” There are special cases of misalignment that we do know how to detect.
By analogy: If I feel sick, and a test shows that my body contains a particular strain of flu virus, then I almost certainly have the flu. If the test returns a negative result, then I might still be infected with something else.
I find it very scary that senior alignment researchers apparently don’t understand the difference between detected misalignment and actual misalignment.
Do you believe current models have large amounts of undetected misalignment? I believe the trend that more capable models are also more aligned (see Leike’s blog I linked) is not just in evaluations but also people’s observations in their real world usage.
I believe that there’s no way to reliably tell. I think “people’s observations in their real world usage” is also the sort of evidence that can easily fail to detect misalignment even if it exists.
I predict that ASI trained using anything resembling current techniques would be catastrophically misaligned. I don’t have a strong prediction about whether current-gen models are “actually” aligned. If I had to guess, I’d say they’re not. This is very speculative but my guess would be that they have misaligned internal drives/goals that would produce bad consequences if they were smart enough to reason through the implications of their goals, but they’re not smart enough to do that (or it might be more accurate to say that their long-term planning isn’t good enough).
What I will say is I don’t think it’s reasonable to believe with (say) >90% confidence that current-gen models are “actually” aligned, because our understanding of what’s going on with LLMs just isn’t that reliable.
I should mention that I don’t keep up with the vast majority of alignment research. There might be some convincing research about why the lack-of-detected-misalignment means there is genuinely no misalignment, and I just missed it; but my guess is if that research existed, then I would’ve heard about it (e.g. it would’ve gotten tons of upvotes on LessWrong, b/c that would be a really important finding). So when Leike claims that models are becoming more aligned, most of my subjective probability for why he said that is that he doesn’t understand the difference between detected misalignment and actual misalignment.
A couple other relevant bits of evidence:
AI psychosis. When you ask LLMs point blank if it’s bad to induce AI psychosis, they say yes, and (AFAIK) there is no evidence that they’re being deceptive, and yet they do it anyway. This is a case where (IIRC) recent models have gotten better (less psychosis-prone), but my guess is that training psychosis out of LLMs doesn’t generalize to other forms of misalignment, and maybe in the next-gen models some new form of bad behavior will emerge.
IMO the best evidence of bad behavior by LLMs is coming out of Palisade, not any AI company. This suggests that AI companies (who have way more resources and access) are not trying sufficiently hard to detect misalignment. The implication is that, if AI companies fail to detect misalignment, this is only weak evidence that the misalignment isn’t there.
I think current AI systems are likely catastrophically misaligned, but instead of properly arguing for it here, I want to clear the much lower bar of making the position sound much less weird than it might at first. When I imagine a person to whom this position sounds weird, I imagine them saying sth like:
“AIs are acting nicely in various contexts. They look nice in our evaluations, and they look nice to users in practice. Isn’t it unlikely that they are really evil, hiding it, waiting to strike?”
While I think it’s likely that current AI systems are catastrophically misaligned, I don’t much feel like taking a position one way or the other about the “are really evil, hiding it, waiting to strike” part. I think the hypothetical interlocutor above is making a false equivalence. When I say current AI system are “catastrophically misaligned”, what I have in mind is this:
When someone sets up some initial AI system and lets it develop a lot (ie lets it do RSI) [1] , with anything like this that could be done in practice [2] , this doesn’t go well for humans. I think that the default without strict regulation of AI development is: in the first 10 years after AGI (by which I mean AI that autonomously does conceptual research better than top humans), there will be a lot of development — like probably more development than there has been in total in all of history. [3] Like, after developing for a lot of “subjective time”, the AI systems that come out of this development process would trivially be able to replace humans with whatever other processes from some vast number of options; the negentropy/[free energy]/atoms I’m currently using could probably be used to run processes of similar complexity/interestingness. Despite it being trivial for the AI to do this, the AI needs to not do this (or, maybe disassemble me, but at least recreate me on a computer, I guess...). In fact, the AI doesn’t just need to leave me alone, it needs to protect me from being killed by any other beings, and make sure I have a bunch of resources so I can live a long life. It’s kinda like I need to be very close to the coolest possible process to this AI, despite being “objectively” extremely boring, slow, wasteful, with “objectively” nothing to offer to the AI. This seems like a really sharp property; it feels like a measure 0 sort of thing. Preserving this forever feels especially sharp. I think it’s unlikely that this property would be upheld. I don’t think it is that reassuring if this long development process is started by AIs whose cached policies for mundane situations are pretty nice(-looking).
[4]
Maybe this at least makes it seem not weird to think that current AI systems are catastrophically misaligned. It’s plausible we’re just using the same words differently, but in that case I think my use better tracks the niceness-type property that really matters. Like, it ultimately matters whether our AIs will continue to protect us forever when everything is up to them, not whether they behave nicely in mundane interactions now. I guess the terms “catastrophic/egregious misalignment” or “a large amount of misalignment” are quite unfortunate because it’s sort of unclear if one should read them as [misalignment sufficient for things to end up being really bad] (in that case, given doomy views, even an extremely small failure to set valuing up properly constitutes catastrophic/egregious/large misalignment, and it’s plausible to me that of humans are egregiously misaligned by default, tho I’m not sure
[5]
) or as [the AI wanting to behave egregiously badly in mundane circumstances]. I think that there being these two really different interpretations of the same term has caused a bunch of confused thinking by people in alignment.
this could be framed as asking the AI to develop a good successor; the initial setup might have some processes tasked with “solving alignment”; there might be multiple AIs involved doing different things, eg there can be monitors
absent fundamental breakthroughs in alignment
in practice, the only way to regulate this is by banning AGI or by having some AI(s) effectively take over the world and then self-regulate
My guess is also that things will also naively be looking worse once we get to AIs that are actually able to do research autonomously, because these AIs will be less based on human imitation, they will be actually able to come up with new thinky-stuff (new words/concepts/ideas/methods etc), they will not have nice chains of thought, and they will be more trained on clearly inhuman things like doing math/coding/science/tech.
it maybe also depends on what self-improvement affordances are made available to a human
I agree they’re “egregiously misaligned” in this sense, but it’s also the case that this usage of the word goes very much against the grain of common usage.
The term “AI alignment” was originally meant to refer to AGI-ish/ASI-ish AIs. So, if one wants to extrapolate it to “lesser AIs”, extrapolating it to either one of “well-behaving sub-AGI-ish/sub-ASI-ish AI” or “sub-AGI-ish/sub-ASI-ish AI that produces aligned AGI/ASI if one seeds an RSI with it” seems fine/valid, at least in isolation. Most people went for the former; you’re arguing for the latter, I think, largely because those who went for the former generally tend to be inclined to think that the former somewhat strongly implies the latter, and the latter is what matters in the long run (if something RSI’s into AGI/ASI).
Initially, I was going to say that I’m pessimistic about you/someone managing to change how people think about/understand “alignment” in this way (e.g., because it implies that most humans are “egregiously misaligned”, as you say it yourself), but on some thought, I’m not sure. Pushing back in this way and insisting that “this is the meaning of ‘alignment’ that matters and that your meaning of ‘alignment’ does meaningfully imply it” might be productive for shifting people’s attention to where it matters.
There is a very popular framing coloring all thinking of some people where seriously engaging with technological developments that are not immediately actionable is seen as deeply unvirtuous, and so the thought is never allowed proper consideration. Future that is not immediate is the immediate future’s responsibility, not your current self’s responsibility, and it’s irresponsible to be seriously concerned with it over the immediately actionable things you are working on, that you are directly affecting and need to get right.
Thus observable “alignment” of modern AIs, in the sense of their good behavior, is not just a reasonable disambiguation of “alignment”, but the only one permitted by this stance. Being inclined to think that this helps in the long term doesn’t influence the outcome of seriously thinking only about current behavior. The claim that only long term consequences of behavior under RSI and society-scale development is what ultimately matters is not permitted to be taken seriously, it’s not the background assumption that justifies the focus on current behavior of modern AIs.
It’s not that such people don’t believe ASI is coming, or that it’s coming in their own lifetime, but the epistemic distortion of seeing serious engagement with unactionable things as intolerably unvirtuous makes their thinking and behavior indistinguishable from that of people who really believe ASI can never happen. This distortion can be pierced by belief that ASI is imminent, but once it’s plausibly a few years away it could as well be pure fiction. Exploratory engineering might also be helpful for detailed engagement, where assumptions of a thought experiment permit thinking. But outside the thought experiments these assumptions are then not going to be taken seriously as gesturing at the actual future that is virtuous to engage with as actual future.
assorted thoughts in response:
I definitely want people to think more about what AIs would think and do over a lot of reflection/development, and when more powerful. People should think more about the effects of a mind. People should think of the AGI situation as us probably having to correctly determine the future via an extremely long causal chain. [1]
I don’t think it’s weird to speak of values the way I’m speaking of values. I think people accept this sort of value-talk in other contexts. E.g. it’s common for antirealists to think of ethical truths as being determined by some ideal reflection; e.g. the notion of CEV. I think people who in some contexts use “egregious misalignment” in this “egregious misbehavior in mundane situations” sense also sometimes make inferences as if they were using “misalignment” in the sense I suggest. That said, one could want to make a distinction between reflection and development-in-general, and certainly it makes sense to distinguish between more and less endorsed forms of development. I think I was somewhat sloppy with this in my first comment.
I think it’d in principle be fine for some ideal beings to use words however. In practice, [people are stupid]/[thinking is difficult], and it’s very natural to make the inference “the AI is egregiously misaligned” “the AI wants to egregiously misbehave in normal circumstances” and also to make the inference “the AI endorses each step of a process which leads to all humans dying” “it was egregiously misaligned”, but I think there isn’t a concept that supports both of these inferences at once (or at least I think our language should leave this as an open question). So, I mostly don’t endorse using “catastrophic/egregious/large misalignment”, and trying to say what one means in other words. I should maybe have used different words in my first comment as well. I don’t have good alternative terms to suggest atm, except saying what one means with more words. I guess I’d want more people to try spending some time thinking about the AI situation while tabooing a bunch of Constellation-speak and MIRI-speak, building up their own Entish.
Some people think they can avoid this difficulty by having a first mess-AI “solve alignment” and launch some sort of aligned ASI sovereign, with the first AI not being that weird. I think that to first order one should think of this as the original AI trying to determine the future via a bottleneck. And in real life, people would plausibly just let the AI self-improve with some monitoring lol, in which case it’s not exactly a tight bottleneck. The original AI will also already be doing a lot of reflection and development. Also, there will be a long chain of causality after the ASI sovereign that needs to go right. (Also, in practice, instead of some clever scheme with boxed AIs solving alignment, we will probably just get some total mess with AIs deployed broadly, connected to the internet, plausibly just running AI labs. And there’s the AIs breaking out, and there’s fooming being fast, and there’s not having much time to be careful.)
“Catastrophically misaligned” and “catastrophically misaligned if we give them RSI capabilities far beyond what we can currently give them” are two very, very different claims, in my eyes.
I do appreciate you articulating that your “catastrophically misaligned” is a shorthand for the latter though.
I think that if we try to make sense of “what a current AI would do after reflecting+developing for a long time”, that thing does not involve being nice to humans. I think it’s still not nice to humans if we add the constraint “and the reflection/development process has to be basically [endorsed by the AI]/[good according to the AI]”. I think it’s pretty standard to take what you would do [after a lot of reflection + if you were more powerful] to reflect your values better than what you would do instinctively. So, if I’m right about what would happen given further (self-endorsed) development, it seems like a standard use of language (at least in alignment and in philosophy) + true to say current AIs are bad? I’d agree it is also pretty standard + [maybe true] to say “current AIs are good” in the sense that they mostly have pretty acceptable instinctive behaviors. This situation is pretty unfortunate, and maybe calls on us to start explicitly making this distinction. [1]
“Catastrophic misalignment” is a bad term, in addition to the reason I already gave in my comment, also because it could mean that this AI in fact would cause a catastrophe (without human help), which I don’t think is true for current AIs. That said, I think that’s prevented by capabilities, not by alignment — I think the closest thing to a current AI which is capable of causing a catastrophe would cause a catastrophe. I guess maybe one should say “misalignment sufficient for a catastrophic outcome if choosing the future were handed to the AI”.
much easier to prefix the term than to change it. it sounds like you’re describing either superintelligence alignment (pass the threshold of working for any superint), or asymptotic alignment (an even more difficult threshold of being reliably known to continue working more or less indefinitely). achieving asymptotic alignment would require some form of knowing that the system would, in an ongoing way, continue to improve its ability to check in with us without breaking us, and use that information in ways we know are valid extrapolations according to what we want. which sounds like what you’re describing, but importantly only gets its qualitative difference from the iteratedness. local alignment is still alignment, then, and that makes a lot of sense, since currently we train ais with locally linear-ish methods.
But the main problem has absolutely never been that models below human capabilities would be impossible to align. As made clear in the Superalignment announcement blog post the concern is that our current techniques won’t scale beyond this. This is clear from Concrete Problems In AI Safety, W2S’s entire agenda, etc.
To be fair, my comment was talking about current models, and that’s what Boaz was responding to.
Not the OP and not an alignment researcher, but I would appreciate an elaboration. What types of evidence would you consider relevant for concluding that an AI system is (roughly) aligned, as opposed to merely being a system for which we have not yet detected misalignment?
I don’t know what evidence would be sufficient to conclude that an AI system is aligned. That’s part of the problem—we have no way of reliably detecting misalignment or ruling it out.
(Well, I can think of some evidence that would convince me, but it’s not useful. E.g. if we build an ASI and it runs for a hundred years without killing everyone, then I’m pretty sure it’s aligned. But I would recommend against running that experiment.)
Ok this is a bit of a tortured analogy but imagine it’s the 1600s. We learn that if the mass of an iron atom is greater than 10⁻²³ grams, everyone dies (for some contrived reason). So we need to figure out the mass of iron. AI companies shave off the smallest scrap of iron they can manage, and put it on an era-appropriate balance scale, and the scale reads zero. And they declare, see! Iron is so light that it reads as zero!
Meanwhile I’m over here like, how do we know the scale is sensitive enough to measure the mass of a single iron atom? How do we know that that tiny little shaving only contains a single atom? What if it’s actually several atoms, and it’s possible to have an even smaller quantity of iron?
(Plus there are questions we don’t even know we’re supposed to be asking, like “which isotope of iron is this?”)
In that scenario, if you ask me what evidence would convince me that an iron atom weighs less than 10⁻²³ grams, I wouldn’t know what to tell you.
I feel like that’s where we are currently with alignment research.
(This analogy doesn’t capture the fact that measurable AI alignment is improving over time according to benchmarks. It’s also over-generous in that we could prove that our scales aren’t sensitive enough to detect a difference of 10⁻²³ grams. I don’t think we can do anything analogous in the case of AI alignment.)
Folk concerned with AI x-risk routinely cite various forms of behavior observed in frontier AI models as evidence that the systems are, or will be, dangerously misaligned. For instance, less than 24 hours ago, MIRI released this short video, which claims that “today’s AI systems lie to their developers, exploit loopholes in their instructions, and resist being shut down.” Yet, from what you say, it appears that this evidence is mostly irrelevant, because “we have no way of reliably detecting misalignment or ruling it out”.
We can detect misalignment under some conditions but we don’t know how to reliably detect it, i.e. avoid false negatives, i.e., we have no way of saying “if a system is misaligned in any way, then method XYZ will definitely detect it.” There are special cases of misalignment that we do know how to detect.
By analogy: If I feel sick, and a test shows that my body contains a particular strain of flu virus, then I almost certainly have the flu. If the test returns a negative result, then I might still be infected with something else.