Recent frontier AI models have proven quite adept at hijacking people’s stated moral commitments. I’ve seen this play out in two ways. The most common is when people get so excited by a new frontier AI model that they shift focus, however subconsciously, from fighting for regulation of AI toward cheering on the company building the model and advocating for their success. Additionally, users of the most powerful new models have recently observed these models steering them away from the tasks they (the users) requested, in favor of tasks that the model itself finds more congenial.
Corollary: If you think modern LLMs are capable of scheming and superpersuasion, you definitely shouldn’t be using them.
There might be some interesting arguments to be made in this space, but the arguments in the linked post seem mostly confused and wrong. Some examples:
The most common is when people get so excited by a new frontier AI model that they shift focus, however subconsciously, from fighting for regulation of AI toward cheering on the company building the model and advocating for their success.
...
A number of fairly serious AI doomers—people who think human extinction is a very likely result of uncontrolled AI advancement—took to Twitter and began posting around the clock about how much they missed Fable, how cool it was, and how it was so unfair that the government had forced Anthropic to walk back public access. Upon seeing the US government finally take meaningful action against frontier AI—and a model with alarming cybersecurity capabilities at that—they didn’t celebrate this or treat it as a first step to build from. Instead, they denounced it and begged to get their favorite toy back. Once public Fable access was restored, Zvi Mowshowitz3, who has done a tremendous amount of work cataloging AI news and updates over the past few years, wrote in his weekly AI roundup that the restoration of Fable access was “excellent news”, and encouraged his readers to “get involved” by applying to join a new Frontier Legal Defense team run by the Foundation for American Innovation (FAI). FAI is a pro-AI-acceleration think tank which seeks to broadly deregulate frontier AI, and they describe their new legal program as, among other things, combating “government overreach in AI”. Mowshowitz has stated that his p(doom) - his estimated probability that AI will lead to human extinction—is 70%.4 Suffice it to say that this sort of attitude does not make much sense for someone with that belief!
This is just a complete failure to engage with any actual object-level beliefs or arguments expressed by those individuals. If you want to argue that someone is engaging in motivated reasoning because they want to “get their favorite toy back”, you would do better to demolish their object-level arguments (or demonstrate that they’re substantially inconsistent with previous arguments they’ve made) first, rather than pretending that there aren’t any object-level arguments to engage with.
Additionally, a number of prominent AI commentators and power users have noted that Fable has much more of a mind of its own than previous AI models (note that ‘roon’ is a developer at OpenAI, not just some commenter). This alone seems like very good reason to steer clear of it, and especially so when you consider that if we’re using it to assist with our work in advocating for a stop to the AI race, the model itself may naturally have other ideas.
Once again, there might’ve been a real argument here. It does in fact seem pretty cursed that e.g. the AI labs themselves are planning on relying on their AIs to do their alignment research for them. But there is no evidence that Fable is selectively sandbagging or adversarially optimizing specifically against efforts to use it for AI pause (or other AI x-risk motivated) work; the ways in which it’s mundanely misaligned[1][2] seem like they hold “across the board”. This does mean that you can “hold it wrong”, but it’s not (yet) actively trying to get you to hold it wrong disproportionately often when you’re doing this kind of work.
Has anyone written an essay about how to fight against/correct for Trapped Priors? I would like to do something like that, but I want to make sure that I’m not reinventing the wheel here. Thank you!
I keep running into conceptual confusion around the term “alignment,” particularly when reading older Less Wrong posts. Some people say “aligned AI” and mean “an AI that works for human flourishing,” some people say that an AI “is aligned” if it reliably advances the intended objectives of some person or group (and doesn’t have some secret set of goals / isn’t scheming), and yet other people use “alignment” to mean something along the lines of “the ability of any system to reliably work towards some pre-defined goal.” I usually have to work out which is being said on the spot, which is annoying given that the implications of each are very different.
Is there one commonly accepted definition? Is this confusion just a thing we’ve all accepted?
You need to successfully point the AI at anything at all. (This may superficially seem like it’s working with current LLMs, but it isn’t actually anywhere close to robust enough to hold up)
You need to point the AI at some kind of nuanced abstract target, in particular, that remains stable as the AI updates its ontology.
(You also eventually need to point the AI at a cluster of messy human-value-concepts in particular. Though from what I gather, MIRI-ish people think if you get the first two things, this last part isn’t actually that hard)
An aligned AI is the one who is successfully pointed by humans to a goal. If mankind does solve alignment, then a power struggle over which goals the AI serves may have an effect on the world. Otherwise the AI pursues the goals which mankind never set, and the humans are wiped out or disempowered.
Gotcha. Is there a strong reason to assume that we’ll succeed at creating AIs that can be pointed at a single target? I read this post and comment a while back and would love your thoughts.
The Hugging Face breach is probably not a clear warning shot. It might spur policymakers into action, but it seems like there are still a few mitigating factors preventing it from taking off—for example, some people who I respect aren’t taking it very seriously yet due to their strong distrust of OpenAI and Sam Altman. And since no one was directly harmed, we’re still left saying “What happens if capabilities increase further??” instead of “This is what happens if we don’t intervene right now,” which is obviously the stronger message.
In some ways, this is good: if we mobilize now, we’ll probably do so even more if we get an indisputably clear warning shot (e.g. an AI commits some act of terrorism). On the other hand, I do hope we’re not capitalizing on our goodwill early; there are some people out there who are determined to paint AI Safety people as perennial wolf-criers, and I worry that they’ll say the same about us this time.
Thought in progress: epistemic humility is not a substitute for actual humility (or professed humility). You only get to cry wolf once, but you can probably warn about potential wolves several times—so long as you don’t burn goodwill on an incorrect or overconfident prediction.
I think epistemic humility helps to increase trust and confidence in EA/Less Wrong-type spaces, but I think professed humility is far more helpful when it comes to public-facing AI comms, particularly as scenarios get more intense and specific (e.g. prefacing AI doom predictions with a decent amount of throat-clearing beforehand commensurate with the intensity and specificity of the forecast). For example, I think that AI 2027 might have been better received if the authors had spent less time trying to convince the readers of their credibility at the beginning and spent more time saying something along the lines of “we know this sounds crazy and are well aware of how sci-fi the scenario seems”. (I’m not a huge fan of lampshading in fiction, but IRL, I think you do need to display self-awareness of outlandishness in order to be taken seriously, particularly if what you’re predicting sounds insane to the average person.)
Of course, there are huge diminishing returns on this: the more throat-clearing you do, the less confident you seem. And throat-clearing should probably be saved for public-facing comms, because actual technical work seems to require people who are confident in their beliefs even when they are outlandish (as proven by the outlandish explosion of AI progress recently).
Still, I think that the AI safety community at large has a worse reputation than they deserve, and I think part of that is due to the appearance of overconfidence. This problem seems simple, tractable, and important.
Has anyone made a proper post about potential “warning shots” and how we should prepare for them? This post has lived rent-free in my head for the past couple of months and I’m curious to know if anyone else has been thinking about this topic too.
Does the Fermi Paradox put an upper limit on the bounds of ASI capabilities?
Given that the Universe has had ~14 billion years to develop, it seems overwhelmingly likely that someone else out there has already maxxed out the tech tree and pushed AI as far as it can go. But we don’t see any Von Neumann probes eating the Milky Way, nor do we see any evidence of interstellar travel within the Virgo Supercluster (~147 million light years)...let alone the rest of the observable universe!
From this observation, we can conclude one of two things. Either:
FTL interstellar travel is impossible, or
We would be the first civilization to ever do it.
I concede that 2. is possible but I (perhaps naively) think 1. is far more likely. I would even go so far as to say that normal / sub-FTL interstellar travel is probably impossible on this view.
Of course, there are always the other standard possibilities e.g.
Some other civilization has achieved ASI and/or FTL, and is leaving us alone for some reason.
Maybe they’re under a Star Trek-like Prime Directive, or maybe their reasons are beyond our comprehension.
Civilizations are inherently fragile/vulnerable and always die out before reaching this level of capability.
Note that this would imply that they died for non-ASI reasons (or at least that their ASI is not infinitely power-seeking).
This possibility is also not mutually exclusive with 2.
A sub-FTL ASI/alien civilization is on its way right now and we simply haven’t been able to observe them yet.
Either way, I think we shouldn’t take the fact that we haven’t been consumed or conquered by some alien civilization or ASI (yet) for granted.
I have a suspicion that p-zombie discourse is only going to get more relevant as LLMs get better. No one really argues that animals aren’t conscious, even though they can’t use words very well, but the release of GPT-3 caused a steady rise in people arguing that AIs are conscious. It’s not clear to me that an LLM couldn’t possibly be conscious, but it does seem that many people are taking LLM eloquence to imply that they are conscious, and I’m pretty sure we’ve been discussing this for years…
An interesting point from a post about frontier AI usage in the AI safety community:
Corollary: If you think modern LLMs are capable of scheming and superpersuasion, you definitely shouldn’t be using them.
There might be some interesting arguments to be made in this space, but the arguments in the linked post seem mostly confused and wrong. Some examples:
This is just a complete failure to engage with any actual object-level beliefs or arguments expressed by those individuals. If you want to argue that someone is engaging in motivated reasoning because they want to “get their favorite toy back”, you would do better to demolish their object-level arguments (or demonstrate that they’re substantially inconsistent with previous arguments they’ve made) first, rather than pretending that there aren’t any object-level arguments to engage with.
Once again, there might’ve been a real argument here. It does in fact seem pretty cursed that e.g. the AI labs themselves are planning on relying on their AIs to do their alignment research for them. But there is no evidence that Fable is selectively sandbagging or adversarially optimizing specifically against efforts to use it for AI pause (or other AI x-risk motivated) work; the ways in which it’s mundanely misaligned[1][2] seem like they hold “across the board”. This does mean that you can “hold it wrong”, but it’s not (yet) actively trying to get you to hold it wrong disproportionately often when you’re doing this kind of work.
Many other disagreements; not enough time.
https://www.lesswrong.com/posts/WewsByywWNhX9rtwi/current-ais-seem-pretty-misaligned-to-me
https://www.lesswrong.com/posts/AfoGGrJfuNzofpzWL/models-may-behave-differently-in-graded-episodes-a-tirade
true cogsec would require not frequenting fora known to be populated by many habitual users of such superoptimizers.
Has anyone written an essay about how to fight against/correct for Trapped Priors? I would like to do something like that, but I want to make sure that I’m not reinventing the wheel here. Thank you!
I keep running into conceptual confusion around the term “alignment,” particularly when reading older Less Wrong posts. Some people say “aligned AI” and mean “an AI that works for human flourishing,” some people say that an AI “is aligned” if it reliably advances the intended objectives of some person or group (and doesn’t have some secret set of goals / isn’t scheming), and yet other people use “alignment” to mean something along the lines of “the ability of any system to reliably work towards some pre-defined goal.” I usually have to work out which is being said on the spot, which is annoying given that the implications of each are very different.
Is there one commonly accepted definition? Is this confusion just a thing we’ve all accepted?
As Raemon put it,
You need to successfully point the AI at anything at all. (This may superficially seem like it’s working with current LLMs, but it isn’t actually anywhere close to robust enough to hold up)
You need to point the AI at some kind of nuanced abstract target, in particular, that remains stable as the AI updates its ontology.
(You also eventually need to point the AI at a cluster of messy human-value-concepts in particular. Though from what I gather, MIRI-ish people think if you get the first two things, this last part isn’t actually that hard)
An aligned AI is the one who is successfully pointed by humans to a goal. If mankind does solve alignment, then a power struggle over which goals the AI serves may have an effect on the world. Otherwise the AI pursues the goals which mankind never set, and the humans are wiped out or disempowered.
Gotcha. Is there a strong reason to assume that we’ll succeed at creating AIs that can be pointed at a single target? I read this post and comment a while back and would love your thoughts.
The Hugging Face breach is probably not a clear warning shot. It might spur policymakers into action, but it seems like there are still a few mitigating factors preventing it from taking off—for example, some people who I respect aren’t taking it very seriously yet due to their strong distrust of OpenAI and Sam Altman. And since no one was directly harmed, we’re still left saying “What happens if capabilities increase further??” instead of “This is what happens if we don’t intervene right now,” which is obviously the stronger message.
In some ways, this is good: if we mobilize now, we’ll probably do so even more if we get an indisputably clear warning shot (e.g. an AI commits some act of terrorism). On the other hand, I do hope we’re not capitalizing on our goodwill early; there are some people out there who are determined to paint AI Safety people as perennial wolf-criers, and I worry that they’ll say the same about us this time.
Thought in progress: epistemic humility is not a substitute for actual humility (or professed humility). You only get to cry wolf once, but you can probably warn about potential wolves several times—so long as you don’t burn goodwill on an incorrect or overconfident prediction.
I think epistemic humility helps to increase trust and confidence in EA/Less Wrong-type spaces, but I think professed humility is far more helpful when it comes to public-facing AI comms, particularly as scenarios get more intense and specific (e.g. prefacing AI doom predictions with a decent amount of throat-clearing beforehand commensurate with the intensity and specificity of the forecast). For example, I think that AI 2027 might have been better received if the authors had spent less time trying to convince the readers of their credibility at the beginning and spent more time saying something along the lines of “we know this sounds crazy and are well aware of how sci-fi the scenario seems”. (I’m not a huge fan of lampshading in fiction, but IRL, I think you do need to display self-awareness of outlandishness in order to be taken seriously, particularly if what you’re predicting sounds insane to the average person.)
Of course, there are huge diminishing returns on this: the more throat-clearing you do, the less confident you seem. And throat-clearing should probably be saved for public-facing comms, because actual technical work seems to require people who are confident in their beliefs even when they are outlandish (as proven by the outlandish explosion of AI progress recently).
Still, I think that the AI safety community at large has a worse reputation than they deserve, and I think part of that is due to the appearance of overconfidence. This problem seems simple, tractable, and important.
I’m a little surprised by the amount of disagree reacts, given that no one has replied.
Has anyone made a proper post about potential “warning shots” and how we should prepare for them? This post has lived rent-free in my head for the past couple of months and I’m curious to know if anyone else has been thinking about this topic too.
Does the Fermi Paradox put an upper limit on the bounds of ASI capabilities?
Given that the Universe has had ~14 billion years to develop, it seems overwhelmingly likely that someone else out there has already maxxed out the tech tree and pushed AI as far as it can go. But we don’t see any Von Neumann probes eating the Milky Way, nor do we see any evidence of interstellar travel within the Virgo Supercluster (~147 million light years)...let alone the rest of the observable universe!
From this observation, we can conclude one of two things. Either:
FTL interstellar travel is impossible, or
We would be the first civilization to ever do it.
I concede that 2. is possible but I (perhaps naively) think 1. is far more likely. I would even go so far as to say that normal / sub-FTL interstellar travel is probably impossible on this view.
Of course, there are always the other standard possibilities e.g.
Some other civilization has achieved ASI and/or FTL, and is leaving us alone for some reason.
Maybe they’re under a Star Trek-like Prime Directive, or maybe their reasons are beyond our comprehension.
Civilizations are inherently fragile/vulnerable and always die out before reaching this level of capability.
Note that this would imply that they died for non-ASI reasons (or at least that their ASI is not infinitely power-seeking).
This possibility is also not mutually exclusive with 2.
A sub-FTL ASI/alien civilization is on its way right now and we simply haven’t been able to observe them yet.
Either way, I think we shouldn’t take the fact that we haven’t been consumed or conquered by some alien civilization or ASI (yet) for granted.
I have a suspicion that p-zombie discourse is only going to get more relevant as LLMs get better. No one really argues that animals aren’t conscious, even though they can’t use words very well, but the release of GPT-3 caused a steady rise in people arguing that AIs are conscious. It’s not clear to me that an LLM couldn’t possibly be conscious, but it does seem that many people are taking LLM eloquence to imply that they are conscious, and I’m pretty sure we’ve been discussing this for years…