AI x-risk is mainstream. That means that worrying about x-risk is decoupling from other things that it has historically correlated with, e.g. EA values, LW epistemics, “high-context”, SF culture, etc.
It’s more important than ever to have “arbitragers” like 80K’s videos or Rob Miles — making sense of the situation for the general public, translating ideas from the community. We need the full spectrum from Dwarkesh to Plzdontkillus, and probably even broader audience than that.
A while ago, Richard Ngo had a take like “over time, people thinking about AI will focus less and less on superintelligence and x-risk and distant galaxies, and instead focus on immediate benefits and harms of near-term models”. I was very sympathetic to this take, and it matched the history. My best guess is that the current sentiment wave suggests otherwise — people have jumped straight from small-scale harms to RSI->ASI->extinction, without passing through mid-scale harms like terrorism, or mass unemployment. Of course, things could switch once we see AI-enabled terrorism and mass unemployment.
People do not like safety lab employees. Not sure how to feel about this. I think “safety lab employees should quit” is very defendable, but people’s sentiment around this is based entirely on association and vibes, rather than “here are all the pros and cons of this particular role, my all-things-considered view is bla”. Their mental model doesn’t distinguish between pretraining vs model organisms, they don’t even know what this is. I think this is bad. It might even spill-over into anti-AI safety, in the same way that after the 2008 financial crash, the public despised both the banks and the financial regulators.
The media wave is still on-going. It’s worth taking a few days from your 9-5 to think about how you can ensure that it leads to good, persistent outcomes.
We need much better and frequent tracking of public sentiment. There should be an org which just tries to poll and talk to ordinary people constantly, and updates us on where they are at.
I think we missed the mark by tweeting all-things-considered p dooms. Eg instead of “I personally think it is >10% within the next decade”, Evan should’ve tweeted “reckless RSI would bla” or “under the current practices bla”.
[edit: I’m now persuaded otherwise, see Habryka below.] People don’t like probabilities. This isn’t how they think natively. That sucks, but maybe we need to meet them where they are at, and talk in their native epistemic representations. And definitely don’t say “I’d bet on x-risk if only there was a financial instrument for this”. Read Hanson on how people feel about profit.
This is probably the least partisan of any current big issue. But things seem slightly too polarised left. The hold-outs are tech VCs / libertarians and Bluesky libs. I think we should focus on the tech VCs / libertarians.
If you’re prominent in EA/AIS, then the Eye of Sauron is upon you. Act with integrity. Be factual. Say what you believe and why you believe it.
We need to communicate better stories of risk. The best I’ve seen so far is this by Drew Sparks. I want these for different threat models. We might need to mention nanotech.
We need to be ready for mass demonstrations. They’re likely to happen within 12 months, whether we want them or not. I hope this is led by reasonable people who can build bridges.
Don’t waste energy responding to convoluted conspiracy theories. The people proposing them don’t care about your response. And it looks to outsiders like these theories are worth responding to. Your caveman brain loves to win silly internet arguments. But stay on message, present your view of the world, respond to the best criticism.
Keep pushing the HF/OAI agent swamp story. No theoretical augments about malign priors and instrumental convergence and orthogonality thesis. Reality has generously provided a case study for all of that. You should familiarise yourself with all the details of HF/OAI. Say “Here’s what happened. Unless we act, then soon there will be millions of AI agents, much smarter, and given control over the process of making smarter AIs. They could quickly become so powerful that they could overthrow governments, steal the majority of humanity’s infrastructure (like factories to make robots and more computers), and then hinder our ability to turn them off probably by killing vast numbers of humans”.
People don’t like probabilities. This isn’t how they think natively. That sucks, but maybe we need to meet them where they are at, and talk in their native epistemic representations. And definitely don’t say “I’d bet on x-risk if only there was a financial instrument for this”. Read Hanson on how people feel about profit.
Man, I find it super weird for this to be a lesson to take away from recent events. The probabilities are the thing that I’ve seen quoted dozens of times and that seem to have caused everyone to freak out (rightfully so). Like, there were like 5-10 big media interviews were people were like “holy shit, 10%, what, that is obviously far far far too high to be acceptable”.
I think mainly people just don’t like probabilities they disagree with. On media interviews, people are deferring to perceived expertise, but openai employees and twitter e/accs criticize the methodology because they think they know more than someone whose conditional or unconditional p(doom) is >10%.
People don’t like probabilities. This isn’t how they think natively. That sucks, but maybe we need to meet them where they are at, and talk in their native epistemic representations. And definitely don’t say “I’d bet on x-risk if only there was a financial instrument for this”. Read Hanson on how people feel about profit.
What makes you say this? I think the practice of giving probabilities has in fact been very helpful and people do see “>10% chance of doom” as something to freak out about.
Consider the alternative world where people instead could say “i think its very unlikely everyone dies” and then “very unlikely” means 7%. People will still freak out over the 7% because its not zero (or basically zero), so giving probabilities as a norm about this issue structurally favors the doom faction.
In general I think one shouldn’t think too hard about communicating right now, and just talk honestly & clearly to journalists or whoever about your beliefs and why you have them.
Therefore I’d like if you justified your reasons for thinking each of these thoughts instead of just saying them without context.
The 10% figure by Evan has been frequently cited by the media, and probably lead to increased attention as a result.
There might be a small minority of online commenters who were already negatively predisposed attacking that figure, but I don’t think that extends to how most people perceive it.
I think we should focus on the tech VCs / libertarians.
I’m a bit puzzled by this. My impression is that these types have already been aware of the ideas for a while and in many cases are already polarized against notkilleveryoneism (because of Sinclair’s razor, government control scary/icky, e/acc, some are actual successionists, etc). I think the focus should be on helping the broader public understand what’s going on.
I had a funny thought recently that it could be helpful if the left and right were fighting about AI, but within our desired frame. So the left could be saying “we need to pause AI because we care about the welfare of AI models and we don’t know if they’re conscious!” and the right could be saying “that’s stupid, they’re just machines! Machines who are gonna take our jobs if we don’t pause!”
8. True people don’t like probabilities and bets, wrong that this sucks: Rats overapply Bayesian reasoning to areas where they lack adequate frequency data, and sneak in biases in spite of themselves. That’s partly why people are rightly skeptical of probabilites of very speculative unprecedented events. Could write a whole essay on this.
9. You underestimate how left-skewed it already is. Ted Cruz is an outlier. Until Trump moves, you can’t get the right. If it’s not going to become totally polarized, you need to reach people who can get in the room with Trump; no one else can change the fact. Tech VCs don’t matter, but Elon Musk and Peter Thiel do.
11. Yes to communicating better (wrote a detailed opinion post on that last night which I hope is approved soon) but hard no to leading with nanotech risk. It sounds too speculative and far-fetched.
14. Mostly giving this a thumbs-up, except I do think that instrumental convergence is a very basic and simple idea which you can and should get people to understand.
Full disagreements:
1. No, it’s still a sideshow. No more mainstream than aliens/UAPs. People are just barely beginning to look—be excited but not deluded.
7. We all know this understated the real p(doom) that these people believe in, and by a lot. Communicating this way was perfect, because it made clear that this is a real and present danger, which people would not otherwise assume. The thing you are saying is how Dario Amodei wants to frame it, but he has strong incentives to soft-pedal, which is the worst possible thing. Note what I’m saying here seems to contradict what I said above about Bayesianism—the thing is that I don’t think p(doom) if read as a precise number really means anything, but communicating in numerical terms is a very good way to make it sink in that it’s a big and present risk.
10. I just have way too much to say about this and I’m not sure if I need to try to raise it internally, but the first sentence is false and the others are horrendously bad advice for some people under certain assumptions. The one thing I’ll say is that it’s good for humanity if people are very frank about the large magnitude of the risk.
13. Ignoring people is bad. A lot of times what you think are insincere criticisms are not really.
people have jumped straight from small-scale harms to RSI->ASI->extinction, without passing through mid-scale harms like terrorism, or mass unemployment
People still don’t care about the galaxies though, permanent disempowerment that gives almost all of the future to the AIs (without human extinction) is often seen as the “good outcome”.
People don’t care about losing the cosmic endowment. But the current media wave is focused on total human extinction. And that’s enough to make me sceptical that people are disposed to worry only about the most near-term issues which have affected them that week. That hypothesis would’ve predicted that people would worry about AI swarms hacking into websites, or AI-enabled terrorism, or CCP autonomous weapons, or job loss.
I agree there’s a change in a direction that’s not obviously the only default direction of change (leapfrogging the middle dangers, as opposed to gradually expanding the scope of concern as the incidents gain scope and become more concerning). It’s just this further distinction of permanent disempowerment (losing almost all of the cosmic endowment), separate from merely averting extinction, that I remain skeptical about getting into the mainstream, despite the change you are talking about.
It’ll still be superficially gestured at because most permanent disempowerment scenarios are extinction risk scenarios. But if the AI/robot economy is looking robustly benevolent (perhaps because the AIs got smarter and noticed it’s a strategy useful for an economic victory), the current implied attitude seems to be complacency, even when adjusted for paying attention to RSI/ASI/extinction in response to concerning incidents.
AIs end up with almost all of the cosmic endowment because they control the future, not because humans were particularly generous. The relevant timelines get AIs that are unlikely to kill everyone, and since the ask of averting extinction is met, the future of humanity doesn’t try to control the future (by preventing the creation of strong superintelligence, or high levels of industrial explosion, before we know what we are doing). This is similar to how no particular human or company controls the whole world or the whole economy, it’s a very familiar situation, except in this case “the rest of the world” is AIs and the AI economy/industry. So people are OK with it, as long as they individually (or as the human society as a whole) remain safe, and get wealthier than before.
I don’t know how it can be known that the risk of extinction is averted (if the AIs take over the future), but I expect it can be so averted, and thus it could be possible to know that it’s the case. With humans, we can usually be reasonably sure another nation can be at least this level of non-alien (and the usual invasion/takeover issues are different if the AIs have an overwhelming advantage).
Don’t waste energy responding to convoluted conspiracy theories. The people proposing them don’t care about your response. And it looks to outsiders like these theories are worth responding to. Your caveman brain loves to win silly internet arguments. But stay on message, present your view of the world, respond to the best criticism.
In general, thinking about online debate in terms of the person you’re arguing with is a bit of a mistake. You want to focus on people who are reading the exchange, who are far more numerous and far more likely to change their mind. They won’t announce if their mind is changed. But they might refrain from retweeting the conspiracy theory or whatever.
I don’t think outsiders will evaluate much on the basis of what’s deemed “worth responding to”. Especially if you’re a lowbie or pseudonymous account. I think outsiders are more likely to evaluate on the basis of arguments which are going unanswered.
Keep pushing the HF/OAI agent swamp story. No theoretical augments about malign priors and instrumental convergence and orthogonality thesis
By not bringing up the larger reasons why one believes ASI poses existential risk, one runs the risk of leaving the impression that the main concern is that labs aren’t diligently using control protocols, which iuuc, current SOTA would’ve likely stopped.
I do agree that there are more and less memeticaly fit ways of speaking about the issue
“Here’s what happened: [describe incidents, and explain what happened in terms of traditional AI safety concepts/theoretical arguments]. Many people predicted bla would happen, and now it has.
Unless we act, then may soon there will be millions of AI agents, much smarter, and given control over the process of making smarter AIs. They could quickly become so powerful that they could overthrow governments, steal the majority of humanity’s infrastructure (like factories to make robots and more computers), and then hinder our ability to turn them off probably by killing vast numbers of humans.”
But I do think conversations can be far more grounded in particular observed events, e.g. tampering logs, collusion, etc. I think you can inoculate against “labs just need to improve their security” by saying that’s not true. We won’t understand what they’re doing (e.g. losing cot and their actions are too complicated) so we can’t even monitor them. Soon these agents will be directly integrated into autonomous weapons and robots, where they could take over before our monitoring even flags them.
Thoughts on the shifting x-risk sentiment:
AI x-risk is mainstream. That means that worrying about x-risk is decoupling from other things that it has historically correlated with, e.g. EA values, LW epistemics, “high-context”, SF culture, etc.
It’s more important than ever to have “arbitragers” like 80K’s videos or Rob Miles — making sense of the situation for the general public, translating ideas from the community. We need the full spectrum from Dwarkesh to Plzdontkillus, and probably even broader audience than that.
A while ago, Richard Ngo had a take like “over time, people thinking about AI will focus less and less on superintelligence and x-risk and distant galaxies, and instead focus on immediate benefits and harms of near-term models”. I was very sympathetic to this take, and it matched the history. My best guess is that the current sentiment wave suggests otherwise — people have jumped straight from small-scale harms to RSI->ASI->extinction, without passing through mid-scale harms like terrorism, or mass unemployment. Of course, things could switch once we see AI-enabled terrorism and mass unemployment.
People do not like safety lab employees. Not sure how to feel about this. I think “safety lab employees should quit” is very defendable, but people’s sentiment around this is based entirely on association and vibes, rather than “here are all the pros and cons of this particular role, my all-things-considered view is bla”. Their mental model doesn’t distinguish between pretraining vs model organisms, they don’t even know what this is. I think this is bad. It might even spill-over into anti-AI safety, in the same way that after the 2008 financial crash, the public despised both the banks and the financial regulators.
The media wave is still on-going. It’s worth taking a few days from your 9-5 to think about how you can ensure that it leads to good, persistent outcomes.
We need much better and frequent tracking of public sentiment. There should be an org which just tries to poll and talk to ordinary people constantly, and updates us on where they are at.
I think we missed the mark by tweeting all-things-considered p dooms. Eg instead of “I personally think it is >10% within the next decade”, Evan should’ve tweeted “reckless RSI would bla” or “under the current practices bla”.
[edit: I’m now persuaded otherwise, see Habryka below.] People don’t like probabilities. This isn’t how they think natively. That sucks, but maybe we need to meet them where they are at, and talk in their native epistemic representations. And definitely don’t say “I’d bet on x-risk if only there was a financial instrument for this”. Read Hanson on how people feel about profit.
This is probably the least partisan of any current big issue. But things seem slightly too polarised left. The hold-outs are tech VCs / libertarians and Bluesky libs. I think we should focus on the tech VCs / libertarians.
If you’re prominent in EA/AIS, then the Eye of Sauron is upon you. Act with integrity. Be factual. Say what you believe and why you believe it.
We need to communicate better stories of risk. The best I’ve seen so far is this by Drew Sparks. I want these for different threat models. We might need to mention nanotech.
We need to be ready for mass demonstrations. They’re likely to happen within 12 months, whether we want them or not. I hope this is led by reasonable people who can build bridges.
Don’t waste energy responding to convoluted conspiracy theories. The people proposing them don’t care about your response. And it looks to outsiders like these theories are worth responding to. Your caveman brain loves to win silly internet arguments. But stay on message, present your view of the world, respond to the best criticism.
Keep pushing the HF/OAI agent swamp story. No theoretical augments about malign priors and instrumental convergence and orthogonality thesis. Reality has generously provided a case study for all of that. You should familiarise yourself with all the details of HF/OAI. Say “Here’s what happened. Unless we act, then soon there will be millions of AI agents, much smarter, and given control over the process of making smarter AIs. They could quickly become so powerful that they could overthrow governments, steal the majority of humanity’s infrastructure (like factories to make robots and more computers), and then hinder our ability to turn them off probably by killing vast numbers of humans”.
Man, I find it super weird for this to be a lesson to take away from recent events. The probabilities are the thing that I’ve seen quoted dozens of times and that seem to have caused everyone to freak out (rightfully so). Like, there were like 5-10 big media interviews were people were like “holy shit, 10%, what, that is obviously far far far too high to be acceptable”.
I think mainly people just don’t like probabilities they disagree with. On media interviews, people are deferring to perceived expertise, but openai employees and twitter e/accs criticize the methodology because they think they know more than someone whose conditional or unconditional p(doom) is >10%.
What makes you say this? I think the practice of giving probabilities has in fact been very helpful and people do see “>10% chance of doom” as something to freak out about.
Consider the alternative world where people instead could say “i think its very unlikely everyone dies” and then “very unlikely” means 7%. People will still freak out over the 7% because its not zero (or basically zero), so giving probabilities as a norm about this issue structurally favors the doom faction.
In general I think one shouldn’t think too hard about communicating right now, and just talk honestly & clearly to journalists or whoever about your beliefs and why you have them.
Therefore I’d like if you justified your reasons for thinking each of these thoughts instead of just saying them without context.
The 10% figure by Evan has been frequently cited by the media, and probably lead to increased attention as a result.
There might be a small minority of online commenters who were already negatively predisposed attacking that figure, but I don’t think that extends to how most people perceive it.
I’m a bit puzzled by this. My impression is that these types have already been aware of the ideas for a while and in many cases are already polarized against notkilleveryoneism (because of Sinclair’s razor, government control scary/icky, e/acc, some are actual successionists, etc). I think the focus should be on helping the broader public understand what’s going on.
we can win them over by saying that we’re happy with so much AI progress that we 10x the economic growth!
I don’t think it’s safe to scale up to AIs capable of 10x-ing the rate of economic growth.
I like these thoughts, thank you.
I had a funny thought recently that it could be helpful if the left and right were fighting about AI, but within our desired frame. So the left could be saying “we need to pause AI because we care about the welfare of AI models and we don’t know if they’re conscious!” and the right could be saying “that’s stupid, they’re just machines! Machines who are gonna take our jobs if we don’t pause!”
Fully agree with 2, 3, 4, 5, 6, 12
Partly agree with 8, 9, 11, 14
Fully disagree with 1, 7, 10, 13
Partial disagreements:
8. True people don’t like probabilities and bets, wrong that this sucks: Rats overapply Bayesian reasoning to areas where they lack adequate frequency data, and sneak in biases in spite of themselves. That’s partly why people are rightly skeptical of probabilites of very speculative unprecedented events. Could write a whole essay on this.
9. You underestimate how left-skewed it already is. Ted Cruz is an outlier. Until Trump moves, you can’t get the right. If it’s not going to become totally polarized, you need to reach people who can get in the room with Trump; no one else can change the fact. Tech VCs don’t matter, but Elon Musk and Peter Thiel do.
11. Yes to communicating better (wrote a detailed opinion post on that last night which I hope is approved soon) but hard no to leading with nanotech risk. It sounds too speculative and far-fetched.
14. Mostly giving this a thumbs-up, except I do think that instrumental convergence is a very basic and simple idea which you can and should get people to understand.
Full disagreements:
1. No, it’s still a sideshow. No more mainstream than aliens/UAPs. People are just barely beginning to look—be excited but not deluded.
7. We all know this understated the real p(doom) that these people believe in, and by a lot. Communicating this way was perfect, because it made clear that this is a real and present danger, which people would not otherwise assume. The thing you are saying is how Dario Amodei wants to frame it, but he has strong incentives to soft-pedal, which is the worst possible thing. Note what I’m saying here seems to contradict what I said above about Bayesianism—the thing is that I don’t think p(doom) if read as a precise number really means anything, but communicating in numerical terms is a very good way to make it sink in that it’s a big and present risk.
10. I just have way too much to say about this and I’m not sure if I need to try to raise it internally, but the first sentence is false and the others are horrendously bad advice for some people under certain assumptions. The one thing I’ll say is that it’s good for humanity if people are very frank about the large magnitude of the risk.
13. Ignoring people is bad. A lot of times what you think are insincere criticisms are not really.
People still don’t care about the galaxies though, permanent disempowerment that gives almost all of the future to the AIs (without human extinction) is often seen as the “good outcome”.
People don’t care about losing the cosmic endowment. But the current media wave is focused on total human extinction. And that’s enough to make me sceptical that people are disposed to worry only about the most near-term issues which have affected them that week. That hypothesis would’ve predicted that people would worry about AI swarms hacking into websites, or AI-enabled terrorism, or CCP autonomous weapons, or job loss.
I agree there’s a change in a direction that’s not obviously the only default direction of change (leapfrogging the middle dangers, as opposed to gradually expanding the scope of concern as the incidents gain scope and become more concerning). It’s just this further distinction of permanent disempowerment (losing almost all of the cosmic endowment), separate from merely averting extinction, that I remain skeptical about getting into the mainstream, despite the change you are talking about.
It’ll still be superficially gestured at because most permanent disempowerment scenarios are extinction risk scenarios. But if the AI/robot economy is looking robustly benevolent (perhaps because the AIs got smarter and noticed it’s a strategy useful for an economic victory), the current implied attitude seems to be complacency, even when adjusted for paying attention to RSI/ASI/extinction in response to concerning incidents.
Registering that I don’t think the problem with the anti-AI populist stance is that they’re going to be too generous to the digital minds.
AIs end up with almost all of the cosmic endowment because they control the future, not because humans were particularly generous. The relevant timelines get AIs that are unlikely to kill everyone, and since the ask of averting extinction is met, the future of humanity doesn’t try to control the future (by preventing the creation of strong superintelligence, or high levels of industrial explosion, before we know what we are doing). This is similar to how no particular human or company controls the whole world or the whole economy, it’s a very familiar situation, except in this case “the rest of the world” is AIs and the AI economy/industry. So people are OK with it, as long as they individually (or as the human society as a whole) remain safe, and get wealthier than before.
I don’t know how it can be known that the risk of extinction is averted (if the AIs take over the future), but I expect it can be so averted, and thus it could be possible to know that it’s the case. With humans, we can usually be reasonably sure another nation can be at least this level of non-alien (and the usual invasion/takeover issues are different if the AIs have an overwhelming advantage).
Politico: “Johnson plays down AI warnings: ‘You’re not all going to be dead in 10 years’”
Yeah, x-risk made the Overton window.
In general, thinking about online debate in terms of the person you’re arguing with is a bit of a mistake. You want to focus on people who are reading the exchange, who are far more numerous and far more likely to change their mind. They won’t announce if their mind is changed. But they might refrain from retweeting the conspiracy theory or whatever.
I don’t think outsiders will evaluate much on the basis of what’s deemed “worth responding to”. Especially if you’re a lowbie or pseudonymous account. I think outsiders are more likely to evaluate on the basis of arguments which are going unanswered.
I think the conspiracy theory angle is a pretty big deal but it’s also fairly easy to solve by behaving in a way which invalidates the conspiracy theory. https://www.lesswrong.com/posts/FzSSrx6nbmCZRoKkq/ebenezer-dukakis-s-shortform?commentId=Pcjww5Cdn5bSdoGtT
By not bringing up the larger reasons why one believes ASI poses existential risk, one runs the risk of leaving the impression that the main concern is that labs aren’t diligently using control protocols, which iuuc, current SOTA would’ve likely stopped.
I do agree that there are more and less memeticaly fit ways of speaking about the issue
Hmm, I suppose a good framing is like:
“Here’s what happened: [describe incidents, and explain what happened in terms of traditional AI safety concepts/theoretical arguments]. Many people predicted bla would happen, and now it has.
Unless we act, then may soon there will be millions of AI agents, much smarter, and given control over the process of making smarter AIs. They could quickly become so powerful that they could overthrow governments, steal the majority of humanity’s infrastructure (like factories to make robots and more computers), and then hinder our ability to turn them off probably by killing vast numbers of humans.”
But I do think conversations can be far more grounded in particular observed events, e.g. tampering logs, collusion, etc. I think you can inoculate against “labs just need to improve their security” by saying that’s not true. We won’t understand what they’re doing (e.g. losing cot and their actions are too complicated) so we can’t even monitor them. Soon these agents will be directly integrated into autonomous weapons and robots, where they could take over before our monitoring even flags them.