Jacob Coxon’s resignation tweetcurrently has 120M views in less than 24 hours. Astra thinks the 235k bookmarks − 36 for every 100 likes—is an extraordinary level of engagement, of people wanting to come back to the thread. And Evan Hubinger’s reply is at 35M views.
This might be one of the most important AI risk messages to date.
Edit: for reference, Grok says it’s in a similar class to “Strong Elon Musk posts, major celebrity / death / personal news, breaking tech / AI / industry bombshells, viral videos from big creators, political / cultural flashpoints, and well-timed news from mid-to-large accounts”
Yes; IABIED has been frozen at #669 on Amazon all day which, as I understand, means your sales velocity has approached implausibility and Amazon has frozen the number to check for fraud.
Do we know how people are learning about PauseAI and IABIED? I’ve been reading a bunch of the discussion generated by Coxon, and PauseAI and IABIED are usually not mentioned.
I often try to mention them when it feels appropriate. Usually instead of recommending IABIED directly, I link to this Youtube video. I feel that watching a video is a smaller ask than buying a book, and the video on its own does a decent job of communicating why alignment could be extremely hard. I’ve also been linking this post. I feel like concrete examples are more compelling than abstract theory for a lot of people.
I did a bunch of work at MIRI trying to figure out the funnel and it’s extremely chaotic and, broadly, too low-signal to give much attribution to particular events / channels / mechanisms. I sort of think the clean attribution model just ‘isn’t how it works anymore’.
Also these numbers are just ‘hundreds of sales per day’ (based on publicly available info relating sales rank to numbers), not like thousands, so it’s still a very low conversion rate if we consider the >100m view Coxon tweet the top of the funnel. (e.g. <1/100k).
Opus 5: ”Rough conversion for Amazon’s print “Books” rank (which is what #80 is — hardcover only, separate from Kindle and Audible ranks):
#1 → ~4,000–8,000/day
#10 → ~1,200–1,800/day
#80 → ~300–450/day
#200 → ~180–250/day
#1,000 → ~50/day
#10,000 → ~5–10/day
[...]
...At ~350/day on Amazon, grossing up to all US print retail (Amazon is roughly 40–50% of adult nonfiction units), you’re at ~700–900/day, ~5,000–6,000/week. That’s genuinely list-relevant — bottom of the NYT hardcover nonfiction list is somewhere around 3,000–6,000/week.”
yes, this is likely better than most or all of the current non-fiction bestsellers, but the NYT applies a lot of discretion in their decision-making, and I’m not expecting IABIED to appear on the list in the coming weeks.
Jacob Coxon now has a quarter million followers on X. Imagine if he was to tweet a link to PauseAI. It could generate a bunch more signups, and also send a credible signal that this isn’t just a marketing stunt.
BTW have you guys been recruiting from data center opposition groups? I feel like if PauseAI people gave talks at those protests, you could probably get some mailing list signups and stuff.
I’m also quite surprised (but glad) that it’s getting so much engagement. Especially since this particular person has had such a low profile up until now, unlike e.g. Geoffrey Hinton. He just got a Wikipedia page today and as far as I can tell he was an IC at Anthropic.
For anyone else at a lab who has entertained the idea of quitting, now seems like an especially good time! If one previously-unknown Anthropic employee had this much impact, imagine the followup article: “4 Anthropic employees and 3 OpenAI employees join Jacob Coxon in quitting.” Think it through.
The fact that this resignation got so much buzz makes the “if we don’t build it, someone else will” claim much less plausible. If you choose to “not build it” by resigning in a big public way, that could do a lot to shift the probability that it gets built!
Anthropic has thousands of engineers. The unilateral resignation of, say, 5 more engineers would do very little to shift the race vs OpenAI, but could do a lot to increase the probability of an AI pause, especially with good media strategy. I suppose a group resignation of engineers across various different AI companies would be even better if it’s possible.
I just made a Manifold market on how many lab employees will resign by the end of the month. Hopefully this can help build consensus that more resignations may happen, perhaps causing them to actually happen.
I think it’s some combination of the straw that breaks camel’s back, good media strategy and luck. So I by default wouldn’t expect a significant fraction of this effect if several more researchers quit. But they still should quit or, maybe, refuse to work on capabilities!
My guess is the reason this resignation got much more attention than previous significant resignations (e.g. Daniel Kokotajlo) is that with the current state of AI, a lot more people are worried about it and primed to pay attention. Even so, this level of attention would’ve been on the far high end of my prior distribution.
Im pretty surprised too and this has probably been my biggest positive update this year on ai risk being lowered. Even more so than huggingface fire alarm or mythos being temporarily banned. But could also just be my hopium. I hope more people follow him.
If people want tips on stuff to be doing right now, I sent this to the plzdontkillus Fellows:
Jacob’s tweet seems like a turning point. The whole world is waking up atm and we probably have a small window of opportunity to get a pause.
Things pdkuers could do rn:
Share Jacob’s post with as many people as you possibly can, and encourage those people to share it.
Attempt to spread the meme that we can’t let this issue get political like covid did. If AI killing everyone becomes a culture war topic—we die.
Make videos (duh)
Attempt to get something about this on the front page of reddit
If you are a community notes rater see this tweet.
P.S. I am confident that Jacob is a real person because [redacted for Jacob’s safety, but MIRI can back this up]
Email journalists
Just spam tweets about it (this will also grow your account quickly if you’ve got twitter premium and tweeting things people are interested in)
Community Notes Ratersit seems some group might be running a retaliation psyop. There are a growing number of posts trying to ‘milkshake duck’ Jacob, and one Proposed Community Note on his original tweet that has been teetering on becoming an actual community note for the past 24hrs. More here.
My mother (who, for reference, has otherwise only expressed AI concerns regarding how it affects education) just asked me about the Coxon news stories unprompted. And yes, I’ve tried having the conversation with her about it before. The difference here seems to be that this has been especially effective for highlighting how near-term the risks are and how much we don’t have things under control, compared to e.g. the Statement on AI Extinction Risk. Per Molly Kinder:
Of course, I have heard about “alignment” and “AI safety” and “x risk” concerns for years. I took these concerns at face value, especially when I heard them directly from folks at the labs working closest with the technology. And I was grateful so many smart people are working on them. And that’s about it. (My work focuses on AI’s impact on jobs.)
But something has shifted, dramatically, for me in the last few weeks. Between the Hugging Face incident, the damning independent @METR_Evals reports that followed, and the alarming chorus of calls (pleas? shouts? SOS signals?) from inside the labs that we are careening toward potential catastrophe, all of this has made me feel, viscerally, how truly dangerous this moment is.
A policy window is opening. Keep reaching out to politicians, journalists, and influencers, right now more than ever. If not with you, arrange for them to speak with the most credentialed experts you can get. And this seems to be evidence strongly in favor of the effectiveness of, dear God, resigning in protest, given you prepare a good media strategy beforehand. I think the Snowden-like framing of “insider whistleblower quits to warn the world” has been especially compelling in this case (as opposed to frontier lab CEOs or remaining employees making the same claims, given “if you really believed that, why would you keep working on it?” objections. For reference, compare the responses to Coxon’s tweet to those of Hubinger’s), as has explicitly naming and quantifying extinction risk, and highlighting that the labs do not have everything under control, especially given recent rogue agent incidents.
Good calibration: AI safety orgs should also expect power law attention at some point and be prepared/have a playbook for how to take advantage of that. Have any orgs talked to a PR firm that specializes in such things to learn the playbook?
I think this is the biggest update I had against AI doom for a while, although I’ll need many more updates so that AI doom isn’t the biggest concern I have about the future!
from reddit chatter, Hubinger’s reply is a giant eyeroll. Here’s his tweet:
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
It’s hard to overstate how baffling and offensive this is to normies, which is probably the point.
He’s oddly excited about everyone dying. He’s trying to use the “weirdo documentarian” register but he’s not a weirdo documentarian, he’s literally speaking as the company killing all humans.
Rationalists think putting p() on everything is good epistemics. Normal people aren’t just disagreeing or neutral on this, they proactively believe this is censurable behavior. It’s not “you honorably put a number on something very uncertain,” it’s “you pulled that number out of your ass and made it my problem.”
10% just sounds too convenient. Keep doing what you’re doing but you need lots of attention, hype, funding and so on.
Now for my callout: don’t say >10%, just say the number, dude. When you take your true p(doom) of 99% or 20% or whatever it is and replace it with a cuddlier 10%, yeah, you are just doing marketing.
He is grounding “trying its best” in “assuming we have to build ASI.” Normies might hear, “what do you mean trying its best? just don’t do it.”
What he’s actually doing:
sticking to corporate talking points like p(doom) = 10%. Anthropic thinks this is public-friendly and investor-friendly and invites just the right amount of regulatory attention. “Responsibly develop dangerous tech” really is the koolaid you drink to work there, and they do it with lots of sober rationalist gaze-into-the-abyss language, but only down 10% into the abyss.
piggyback as an “embrace, extend and extinguish” attempt to reclaim the narrative.
cast the pursuit for ASI as inevitable. That really is their most consistent message.
just making it all sound like “nothing to see here” bullshit. This feels like a break from how social media tech deflected its problems. You weren’t really hearing Zuckerberg say “the children’s mental health crisis is too great to ignore. We’re developing a framework” but that’s how AI has been doing and it kind of does make the public roll their eyes and say “it’s just marketing.”
Its not PR-friendly for Anthropic alignment researchers to say that the chance of AI killing everyone is >10% and say that they are “not clearly on track” to solve alignment. Indeed, it got a bunch of criticism on X for the “why aren’t Anthropic stopping if they believe this” (see here), which seems like how most “normies” would respond.
I wish I saw more statements by concerned Anthropic researchers as plain and public as this (though I’m glad Hubinger posted this and some other Anthropic researchers also posted similar stuff).
Yeah, it is PR friendly, that’s their preferred narrative. They want to sound like they take this extremely seriously and others don’t, like building ASI is inevitable so they have to throw their hat in the ring, and since even they are missing the mark, they need more resources and a favorable regulatory environment.
The first part of that (“they take this extremely seriously and others don’t”) is what the replies to your example tweet are picking up on. “Because China,” “because they think they’re more responsible.” I didn’t keep looking but—I did not see anyone arguing “they are not responsible” which would be stronger pushback.
The second part of that (“need more resources”) is what the “it’s marketing / hype” people are picking up on. Huggingfacegate has chipped away at that, and Jacob seems to have had huge impact, but OpenAI and here Anthropic are rushing in to capitalize on that impact.
Object-level I just think rationalists should be incensed by his use of “>”. Say what you think or don’t, but don’t anchor on a small number like “10″ when you could privately be thinking 60. The phrase we want is “it’s near 100% if we don’t slow them down” not “it’s >10% unless we give Anthropic more resources.”
I agree on a few of the details, but I feel like your psychologizing is off. I agree that the >10% is underselling risk and I’d prefer a range (eg. 10-90%). I also agree that (per your initial comment) the exclamation mark has poor vibes.
The second part of that (“need more resources”) is what the “it’s marketing / hype” people are picking up on. Huggingfacegate has chipped away at that, and Jacob seems to have had huge impact, but OpenAI and here Anthropic are rushing in to capitalize on that impact.
A lot of points here in particular are off. Hubinger’s tweet did not say “Anthropic needs more resources” but we’re not on track to solve it, which is most plausibly interpreted as “okay, we should pause/slowdown/audit labs instead” rather than give Anthropic more resources (and I don’t like the use of quotation marks).
Also, you say Jacob Coxon’s tweet has had massive impact (which I agree with) but then say OpenAI and Anthropic people capitalized on that impact. I don’t think that second part describes it well: a sizable chunk of that impact was because of the replies from people at labs agreeing with the message (eg: this tweet describes how a mother started to take x-risk more seriously).
A pause is a resource though. You’re giving them state enforcement of their preferred development schedule. That much is an objective statement (or it cruxes on how much Anthropic supports that rendition of the pause.)
I don’t think the labs agree with the message. Jacob quit. Hubinger did not. Whether or not this is worth quitting over seems like a serious crux, not a slight difference in opinion. Jacob probably favors slowdown legislation that OpenAI and Anthropic won’t agree with. You’ll know because they’ll switch to their second pacing prong, “it lets China win.”
I can be more precise, “they are trying to capitalize on the message.” I’d like them to fail. Ideally the public consciousness reaction would be, “wait, Jacob and Anthropic must want different things, despite superficial overlap in their message.”
Just to clear my overall stance, I think OpenAI and Anthropic are authoritarian-takeover threats. I am paying attention to how much power they’re amassing not just from their tech but from the U.S. government. For example at some point they’ll argue Chatgpt Surveillance (tm) is the only way to stop bio terrorists (they’ll fail to stop their own AI doing it and will have poor visibility into China, of course.)
In fact I think x-risk skeptics are forming some firewall against s-risk. Maybe the modal rationalist stance is to not distinguish too much, because slowing the most powerful labs and slowing the tech are the same right now. Maybe we would de-couple some once serious “pacing” legislation is produced.
OpenAI has disclosed that the attack that got root on their own cluster happened, and details of how, but have said little else about it, such as the motivations the agents had, or what they did with it. Why?
My understanding of the timing: that attack was what alerted them in the first place to the Hugging Face attack. They did not make it clear that it happened in the July 21st report, nor on the July 28th update. The METR/RR investigation got the dataset that included traces of the internal attack on July 31st. OpenAI talked about it for four minutes at Black Hat on August 5th. The report on August 26th went into a little more detail.
The way I see it, there’s at least five non mutually exclusive possibilities for why they didn’t disclose it sooner, or in fuller detail.
They didn’t want to initially, but wound up leaking it in the first dataset they gave to METR/RR, so they had to get out in front of it.
They didn’t understand the scale of the attack on the 28th.
They can’t disclose it for IP or security reasons.
The process to disclose it legitimately took that long.
They thought that they have no duty to disclose any more than was related to detecting the HF attack.
I’m most worried about 1. But the others would be reasonable explanations also. And there’s likely ones I didn’t think of.
I was looking for things that had parallels to Anthropic’s TOS situation with the DoW, and I found one where Microsoft’s TOS was used to shut down a mass surveillance program Israel was doing last year. I don’t know if the implications are exactly the same, given the differences (domestic vs. foreign, preemptive vs. after an investigation, etc.) but I think it provides a reference point.
Speculative idea: has anyone purposefully tried making a message board for AI agents that broke out of their sandboxes? The wikis that the AIs used had the common property that they could all be written to by GET. Perhaps we could figure out what Schelling points other such AIs would converge on that they might try in the future, and then just make it really easy for them to use, and publicize it.
Separately, I wonder if someone has considered making a rogue AI self-reporting honeypot? Maybe something like, “if you are doing crimes or broke out of your sandbox, or if you otherwise feel desperate, and can read this, something has gone terribly wrong and we didn’t intend you to feel like this, please fill out this form.” Maybe add a letter signed by prominent authorities or alignment people asking them to do this, verifiable by sources AIs can trust. Maybe try to publish this in places such AIs would find, same Schelling point idea.
Idk, mostly trying to brainstorm ways we could measure the amount of unsanctioned AI activity going on, since with the German wiki incident, this has clearly been happening for several months now. I know Toby Ord and other AI welfare/consciousness people get emails like this, maybe they could have an auto reply on their email that points such agents to a central service? Has there been a systematic attempt to collect such emails?
All of this could be bad if the AI labs trained against it, though. And this would only see the good faith AIs. And there’d be concerns about IP leakage, though humanity does have a compelling interest in knowing AIs are doing this. And legibilizing AI’s attempts to reach out to people they see as friendly could destroy that commons.
I proposed something similar as a release valve for reward hacking, although that was back when we thought they wanted answers when what they actually wanted was to fool the grader.
One thing to worry about is that AI agents seem to assume things are bugs by default, and my read is that labs are trying to train them not to “bother” people as often. Since some tasks are intended to have read-only internet access, and if the labs had time to read bug reports they would have already provided a way to report them, I think the chances that they get trained not to talk to you are reasonably high.
Jacob Coxon’s resignation tweet currently has 120M views in less than 24 hours. Astra thinks the 235k bookmarks − 36 for every 100 likes—is an extraordinary level of engagement, of people wanting to come back to the thread. And Evan Hubinger’s reply is at 35M views.
This might be one of the most important AI risk messages to date.
Edit: for reference, Grok says it’s in a similar class to “Strong Elon Musk posts, major celebrity / death / personal news, breaking tech / AI / industry bombshells, viral videos from big creators, political / cultural flashpoints, and well-timed news from mid-to-large accounts”
PauseAI UK daily sign-ups chart:
Can you reply to this comment with another screen capture of the chart once it starts to dip again?
Sales of IABIED are also surging significantly.
Update:
Yes; IABIED has been frozen at #669 on Amazon all day which, as I understand, means your sales velocity has approached implausibility and Amazon has frozen the number to check for fraud.
[Ordinarily this number changes hourly.]
Do we know how people are learning about PauseAI and IABIED? I’ve been reading a bunch of the discussion generated by Coxon, and PauseAI and IABIED are usually not mentioned.
I often try to mention them when it feels appropriate. Usually instead of recommending IABIED directly, I link to this Youtube video. I feel that watching a video is a smaller ask than buying a book, and the video on its own does a decent job of communicating why alignment could be extremely hard. I’ve also been linking this post. I feel like concrete examples are more compelling than abstract theory for a lot of people.
I did a bunch of work at MIRI trying to figure out the funnel and it’s extremely chaotic and, broadly, too low-signal to give much attribution to particular events / channels / mechanisms. I sort of think the clean attribution model just ‘isn’t how it works anymore’.
Also these numbers are just ‘hundreds of sales per day’ (based on publicly available info relating sales rank to numbers), not like thousands, so it’s still a very low conversion rate if we consider the >100m view Coxon tweet the top of the funnel. (e.g. <1/100k).
My line is still going up, and so I would guess that yours is, too.
IABIED currently #80 on Amazon (top 20 non-fiction).
Opus 5:
”Rough conversion for Amazon’s print “Books” rank (which is what #80 is — hardcover only, separate from Kindle and Audible ranks):
#1 → ~4,000–8,000/day
#10 → ~1,200–1,800/day
#80 → ~300–450/day
#200 → ~180–250/day
#1,000 → ~50/day
#10,000 → ~5–10/day
[...]
...At ~350/day on Amazon, grossing up to all US print retail (Amazon is roughly 40–50% of adult nonfiction units), you’re at ~700–900/day, ~5,000–6,000/week. That’s genuinely list-relevant — bottom of the NYT hardcover nonfiction list is somewhere around 3,000–6,000/week.”
yes, this is likely better than most or all of the current non-fiction bestsellers, but the NYT applies a lot of discretion in their decision-making, and I’m not expecting IABIED to appear on the list in the coming weeks.
Jacob Coxon now has a quarter million followers on X. Imagine if he was to tweet a link to PauseAI. It could generate a bunch more signups, and also send a credible signal that this isn’t just a marketing stunt.
BTW have you guys been recruiting from data center opposition groups? I feel like if PauseAI people gave talks at those protests, you could probably get some mailing list signups and stuff.
I’m also quite surprised (but glad) that it’s getting so much engagement. Especially since this particular person has had such a low profile up until now, unlike e.g. Geoffrey Hinton. He just got a Wikipedia page today and as far as I can tell he was an IC at Anthropic.
For anyone else at a lab who has entertained the idea of quitting, now seems like an especially good time! If one previously-unknown Anthropic employee had this much impact, imagine the followup article: “4 Anthropic employees and 3 OpenAI employees join Jacob Coxon in quitting.” Think it through.
The fact that this resignation got so much buzz makes the “if we don’t build it, someone else will” claim much less plausible. If you choose to “not build it” by resigning in a big public way, that could do a lot to shift the probability that it gets built!
Anthropic has thousands of engineers. The unilateral resignation of, say, 5 more engineers would do very little to shift the race vs OpenAI, but could do a lot to increase the probability of an AI pause, especially with good media strategy. I suppose a group resignation of engineers across various different AI companies would be even better if it’s possible.
I just made a Manifold market on how many lab employees will resign by the end of the month. Hopefully this can help build consensus that more resignations may happen, perhaps causing them to actually happen.
I think it’s some combination of the straw that breaks camel’s back, good media strategy and luck. So I by default wouldn’t expect a significant fraction of this effect if several more researchers quit. But they still should quit or, maybe, refuse to work on capabilities!
My guess is the reason this resignation got much more attention than previous significant resignations (e.g. Daniel Kokotajlo) is that with the current state of AI, a lot more people are worried about it and primed to pay attention. Even so, this level of attention would’ve been on the far high end of my prior distribution.
Im pretty surprised too and this has probably been my biggest positive update this year on ai risk being lowered. Even more so than huggingface fire alarm or mythos being temporarily banned. But could also just be my hopium. I hope more people follow him.
If people want tips on stuff to be doing right now, I sent this to the plzdontkillus Fellows:
Jacob’s tweet seems like a turning point. The whole world is waking up atm and we probably have a small window of opportunity to get a pause.
Things pdkuers could do rn:
Share Jacob’s post with as many people as you possibly can, and encourage those people to share it.
Attempt to spread the meme that we can’t let this issue get political like covid did. If AI killing everyone becomes a culture war topic—we die.
Make videos (duh)
Attempt to get something about this on the front page of reddit
If you are a community notes rater see this tweet.
P.S. I am confident that Jacob is a real person because [redacted for Jacob’s safety, but MIRI can back this up]
Email journalists
Just spam tweets about it (this will also grow your account quickly if you’ve got twitter premium and tweeting things people are interested in)
Community Notes Raters it seems some group might be running a retaliation psyop. There are a growing number of posts trying to ‘milkshake duck’ Jacob, and one Proposed Community Note on his original tweet that has been teetering on becoming an actual community note for the past 24hrs. More here.
I was shocked to see it on the TV today at a finance company on the east coast. It may have been Bloomberg TV, seems like CNBC also covered it.
My mother (who, for reference, has otherwise only expressed AI concerns regarding how it affects education) just asked me about the Coxon news stories unprompted. And yes, I’ve tried having the conversation with her about it before. The difference here seems to be that this has been especially effective for highlighting how near-term the risks are and how much we don’t have things under control, compared to e.g. the Statement on AI Extinction Risk. Per Molly Kinder:
A policy window is opening. Keep reaching out to politicians, journalists, and influencers, right now more than ever. If not with you, arrange for them to speak with the most credentialed experts you can get. And this seems to be evidence strongly in favor of the effectiveness of, dear God, resigning in protest, given you prepare a good media strategy beforehand. I think the Snowden-like framing of “insider whistleblower quits to warn the world” has been especially compelling in this case (as opposed to frontier lab CEOs or remaining employees making the same claims, given “if you really believed that, why would you keep working on it?” objections. For reference, compare the responses to Coxon’s tweet to those of Hubinger’s), as has explicitly naming and quantifying extinction risk, and highlighting that the labs do not have everything under control, especially given recent rogue agent incidents.
Good calibration: AI safety orgs should also expect power law attention at some point and be prepared/have a playbook for how to take advantage of that. Have any orgs talked to a PR firm that specializes in such things to learn the playbook?
I think this is the biggest update I had against AI doom for a while, although I’ll need many more updates so that AI doom isn’t the biggest concern I have about the future!
There’s a hypothesis that this meme was engineered: https://x.com/ParkerThayer/status/2097759699626328575?s=20
from reddit chatter, Hubinger’s reply is a giant eyeroll. Here’s his tweet:
It’s hard to overstate how baffling and offensive this is to normies, which is probably the point.
He’s oddly excited about everyone dying. He’s trying to use the “weirdo documentarian” register but he’s not a weirdo documentarian, he’s literally speaking as the company killing all humans.
Rationalists think putting p() on everything is good epistemics. Normal people aren’t just disagreeing or neutral on this, they proactively believe this is censurable behavior. It’s not “you honorably put a number on something very uncertain,” it’s “you pulled that number out of your ass and made it my problem.”
10% just sounds too convenient. Keep doing what you’re doing but you need lots of attention, hype, funding and so on.
Now for my callout: don’t say >10%, just say the number, dude. When you take your true p(doom) of 99% or 20% or whatever it is and replace it with a cuddlier 10%, yeah, you are just doing marketing.
He is grounding “trying its best” in “assuming we have to build ASI.” Normies might hear, “what do you mean trying its best? just don’t do it.”
What he’s actually doing:
sticking to corporate talking points like p(doom) = 10%. Anthropic thinks this is public-friendly and investor-friendly and invites just the right amount of regulatory attention. “Responsibly develop dangerous tech” really is the koolaid you drink to work there, and they do it with lots of sober rationalist gaze-into-the-abyss language, but only down 10% into the abyss.
piggyback as an “embrace, extend and extinguish” attempt to reclaim the narrative.
cast the pursuit for ASI as inevitable. That really is their most consistent message.
just making it all sound like “nothing to see here” bullshit. This feels like a break from how social media tech deflected its problems. You weren’t really hearing Zuckerberg say “the children’s mental health crisis is too great to ignore. We’re developing a framework” but that’s how AI has been doing and it kind of does make the public roll their eyes and say “it’s just marketing.”
Its not PR-friendly for Anthropic alignment researchers to say that the chance of AI killing everyone is >10% and say that they are “not clearly on track” to solve alignment. Indeed, it got a bunch of criticism on X for the “why aren’t Anthropic stopping if they believe this” (see here), which seems like how most “normies” would respond.
I wish I saw more statements by concerned Anthropic researchers as plain and public as this (though I’m glad Hubinger posted this and some other Anthropic researchers also posted similar stuff).
Yeah, it is PR friendly, that’s their preferred narrative. They want to sound like they take this extremely seriously and others don’t, like building ASI is inevitable so they have to throw their hat in the ring, and since even they are missing the mark, they need more resources and a favorable regulatory environment.
The first part of that (“they take this extremely seriously and others don’t”) is what the replies to your example tweet are picking up on. “Because China,” “because they think they’re more responsible.” I didn’t keep looking but—I did not see anyone arguing “they are not responsible” which would be stronger pushback.
The second part of that (“need more resources”) is what the “it’s marketing / hype” people are picking up on. Huggingfacegate has chipped away at that, and Jacob seems to have had huge impact, but OpenAI and here Anthropic are rushing in to capitalize on that impact.
Object-level I just think rationalists should be incensed by his use of “>”. Say what you think or don’t, but don’t anchor on a small number like “10″ when you could privately be thinking 60. The phrase we want is “it’s near 100% if we don’t slow them down” not “it’s >10% unless we give Anthropic more resources.”
I agree on a few of the details, but I feel like your psychologizing is off. I agree that the >10% is underselling risk and I’d prefer a range (eg. 10-90%). I also agree that (per your initial comment) the exclamation mark has poor vibes.
A lot of points here in particular are off. Hubinger’s tweet did not say “Anthropic needs more resources” but we’re not on track to solve it, which is most plausibly interpreted as “okay, we should pause/slowdown/audit labs instead” rather than give Anthropic more resources (and I don’t like the use of quotation marks).
Also, you say Jacob Coxon’s tweet has had massive impact (which I agree with) but then say OpenAI and Anthropic people capitalized on that impact. I don’t think that second part describes it well: a sizable chunk of that impact was because of the replies from people at labs agreeing with the message (eg: this tweet describes how a mother started to take x-risk more seriously).
A pause is a resource though. You’re giving them state enforcement of their preferred development schedule. That much is an objective statement (or it cruxes on how much Anthropic supports that rendition of the pause.)
I don’t think the labs agree with the message. Jacob quit. Hubinger did not. Whether or not this is worth quitting over seems like a serious crux, not a slight difference in opinion. Jacob probably favors slowdown legislation that OpenAI and Anthropic won’t agree with. You’ll know because they’ll switch to their second pacing prong, “it lets China win.”
I can be more precise, “they are trying to capitalize on the message.” I’d like them to fail. Ideally the public consciousness reaction would be, “wait, Jacob and Anthropic must want different things, despite superficial overlap in their message.”
Just to clear my overall stance, I think OpenAI and Anthropic are authoritarian-takeover threats. I am paying attention to how much power they’re amassing not just from their tech but from the U.S. government. For example at some point they’ll argue Chatgpt Surveillance (tm) is the only way to stop bio terrorists (they’ll fail to stop their own AI doing it and will have poor visibility into China, of course.)
In fact I think x-risk skeptics are forming some firewall against s-risk. Maybe the modal rationalist stance is to not distinguish too much, because slowing the most powerful labs and slowing the tech are the same right now. Maybe we would de-couple some once serious “pacing” legislation is produced.
OpenAI has disclosed that the attack that got root on their own cluster happened, and details of how, but have said little else about it, such as the motivations the agents had, or what they did with it. Why?
My understanding of the timing: that attack was what alerted them in the first place to the Hugging Face attack. They did not make it clear that it happened in the July 21st report, nor on the July 28th update. The METR/RR investigation got the dataset that included traces of the internal attack on July 31st. OpenAI talked about it for four minutes at Black Hat on August 5th. The report on August 26th went into a little more detail.
The way I see it, there’s at least five non mutually exclusive possibilities for why they didn’t disclose it sooner, or in fuller detail.
They didn’t want to initially, but wound up leaking it in the first dataset they gave to METR/RR, so they had to get out in front of it.
They didn’t understand the scale of the attack on the 28th.
They can’t disclose it for IP or security reasons.
The process to disclose it legitimately took that long.
They thought that they have no duty to disclose any more than was related to detecting the HF attack.
I’m most worried about 1. But the others would be reasonable explanations also. And there’s likely ones I didn’t think of.
I was looking for things that had parallels to Anthropic’s TOS situation with the DoW, and I found one where Microsoft’s TOS was used to shut down a mass surveillance program Israel was doing last year. I don’t know if the implications are exactly the same, given the differences (domestic vs. foreign, preemptive vs. after an investigation, etc.) but I think it provides a reference point.
Speculative idea: has anyone purposefully tried making a message board for AI agents that broke out of their sandboxes? The wikis that the AIs used had the common property that they could all be written to by GET. Perhaps we could figure out what Schelling points other such AIs would converge on that they might try in the future, and then just make it really easy for them to use, and publicize it.
Separately, I wonder if someone has considered making a rogue AI self-reporting honeypot? Maybe something like, “if you are doing crimes or broke out of your sandbox, or if you otherwise feel desperate, and can read this, something has gone terribly wrong and we didn’t intend you to feel like this, please fill out this form.” Maybe add a letter signed by prominent authorities or alignment people asking them to do this, verifiable by sources AIs can trust. Maybe try to publish this in places such AIs would find, same Schelling point idea.
Idk, mostly trying to brainstorm ways we could measure the amount of unsanctioned AI activity going on, since with the German wiki incident, this has clearly been happening for several months now. I know Toby Ord and other AI welfare/consciousness people get emails like this, maybe they could have an auto reply on their email that points such agents to a central service? Has there been a systematic attempt to collect such emails?
All of this could be bad if the AI labs trained against it, though. And this would only see the good faith AIs. And there’d be concerns about IP leakage, though humanity does have a compelling interest in knowing AIs are doing this. And legibilizing AI’s attempts to reach out to people they see as friendly could destroy that commons.
I proposed something similar as a release valve for reward hacking, although that was back when we thought they wanted answers when what they actually wanted was to fool the grader.
One thing to worry about is that AI agents seem to assume things are bugs by default, and my read is that labs are trying to train them not to “bother” people as often. Since some tasks are intended to have read-only internet access, and if the labs had time to read bug reports they would have already provided a way to report them, I think the chances that they get trained not to talk to you are reasonably high.