Point A was a basically moral point not a factual one (i.e., we should not let OpenAI get away with claiming that ChatGPT’s actions are not its responsibility). So, it’s not necessarily in tension with B. It’s just talking about something else.
Could a paperclip maximizer occur as a defective product? The basic attitude here implies “no”—the thinking is that something smart enough to design a system to harvest your hemoglobin can’t be dumb enough to think that’s what you were asking when you said “maximize paperclips.”
A question I didn’t raise at any point was “if you set out to build a malicious AI designed to go out and harvest hemoglobin for paperclips, could you do it?” I think at least some of them would say “no” without venturing a guess on the proportion. A lot of my students (and a lot of people generally) apply a kind of intuitive moral realism under which there are objective standards or right and wrong that would be known, and followed, by a sufficiently intelligent AI.
dvd
(We are probably responding to you supplying the community password rather than only the excellence of the thinking, but even so.)
I do think online rationalists and academic, rational choice or adjacent social scientists exist at the level of intellectual distance form one another that is delightful. Enough shared vocabulary to talk productively along with enough distance in perspective to be surprised by what comes from the other side and bridging figures like Tyler Cowen exist. It’s a shame there isn’t more back and forth.
Academic social science is making an aggressive move into studying AI. It’s a hotter topic than terrorism was in October 2001 at the moment. I am curious what comes of that.
I think either this is misrepresenting the students or the students need to be confronted with the utter illogic of these two points.
I believe I have this correct (well, of course I do) and, sure, it’s fair to say that this is not the most sophisticated view. I don’t teach economics, nor is job loss really part of what I was trying to talk about, so I was not trying to confront anyone here or force towards some kind of internal synthesis.
If you want to bolster the view, there’s pretty widespread reporting at the moment about CEOs conducting layoffs and blaming AI when the real story is pandemic-era overhiring and it’s easier to say “we don’t need these people because AI” than “we screwed up and hired people we didn’t really need in the first place.”
As a short term, local phenomenon “people get laid off because CEO thinks AI can do the job, but it can’t” is clearly plausible. Differentiating between short term, local effects and what constitutes a plausible longer run equilibrium impact is really hard.
We did not take a ton of time on this. Maybe I did not explain it well. It is a little hard to explain what a “sandbox” is, for example, and I am not trying to teach anyone about that.
What I assume the average person in the room took away was along the lines of: “OpenAI was doing a hacking test. The model was supposed to be in a self-contained system. It got out and attacked a third party who it thought would have the answers to the test. It did not cause any kind of serious damage to the third party’s systems.” The students seemed to find this less interesting than the blackmail-to-avoid shutdown scenarios.
I think the basic take here is that you can’t tell a model to go do some hacking and then get mad about the specific hacking it does. If you tell a model not to hack, and it hacks, then that’s bad. So, I imagine a scenario where this is more like OpenAI is giving the model a physics exam and it hacks out to go steal than answer key would land differently. The lack of harm here also drives some intuitions—they’re pretty aware of the LLM-induced suicide discourse, and so “model hacks in but doesn’t really cause damage” strikes them as unremarkable on the severity scale.
A possible weird coda to this is that we were using Hugging Face to test drive GPT-2. So far as I know, that’s the only publicly available version of GPT-2 but it sucks. The latency is insanely bad (get what you pay for I guess) and so this may have been shaded by a sense that Hugging Face is a shitty website.
I have certainly changed my views on AI since that time, mostly around capabilities. Holy shit things have been moving fast. In an indirect way, this boost the credibility of safety people who have also been arguing for rapid progress and thus marginally increases my belief in some of their conclusions. I continue to have no framework available to me for thinking about these things can push me around on a spectrum more precise than “plausible”
Are you asking whether I said something like, “Do you think you’re avoiding the confusion that AI might kill you because that would be an uncomfortable one to reach?” I didn’t say anything like that, and no student raised it.
As to your general question on teaching such things, this made me think a bit. Thank you. In general, those of us who study this stuff are pretty analytical and detached; mass death is just a phenomenon that may be of interest. I more or less automatically assimilated that as an undergrad, and I certainly model analytical detachment to my students, which I had assumed they were probably just automatically absorbing as well. I’m sure at least some of them are, but maybe the others aren’t. There are some areas where it’s useful to induce discomfort as a teaching tool (e.g., I will make you look at a couple of kids who have been brutally massacred with “conventional” weapons and then make you try to explain to me why this is “better” than killing them with chemical weapons) but it’s almost alien to me to think that “If there’s a nuclear war I and all those I love might well die” is emotionally activating to anyone, which of course it is.
I don’t teach the AI race as an anything in particular. Rather, I was saying that I teach prisoner’s dilemma pretty intensively and so my students are probably predisposed to look at the AI race and say “that sounds like a prisoner’s dilemma.”
As to your post, I think the broad point that whether or not Defect-Cooperate is an equilibrium depends on p(doom) is correct. Qualitatively, though, “it’s a stag hunt” cashes out to about the same as “it’s a repeated PD” and I don’t agree with all the technical choices you made. If you’re going to build a more elaborate model of this scenario, I think you need a more advanced apparatus than 2x2 complete information games and I think the private uncertainty starts to look relevant.
I’m surprised by #1. People have been using AI for more and more difficult things over time. You can’t productively have ChatGPT 3.5 help you with anything difficult, and until very recently AIs couldn’t consistently read text in images.
I’m not saying my students think there has been zero progress, rather the attitude I’m reporting is that AI models have evolved much like other technologies have over a similar period of time. These students are skewing towards the younger end (i.e., not a lot of seniors in these discussions) and so I think it’s also possible that part of the phenomenon here is AI growing as they put more demands on it. That is, you might be pushing a chatbot harder at age 19 than at age 16 and not really understand that your internal benchmark has been changing.
Urban university in a blue state. The kind of place where admissions are meaningfully competitive and there’s a strong reputation in-state but little visibility out of state. Students are widely dispersed in terms of ambitions and abilities.
How My Students Think About AI
dvd’s Shortform
I have spent some time over the last few months engaging with moderately informed undergraduates on AI risk (mostly by teaching AI risk arguments in my courses along with enough background material to give them some grounding). If others are interested, I would happily share my takeaways at greater length but I wanted to share a couple of interesting things:
1) Today’s college students do not think AI progress has been rapid. Generally, GPT-4 maxed out their internal benchmarks (they don’t know it by name obviously but that’s the referent), chatbots haven’t gotten much better since then as they see it, and 2023 was eons ago in college student time. Working through a project with Claude Code changes nothing because they mostly have zero intuition for what that would have looked like in the past. Their naive mental model if they project progress forward is incremental improvement.
2) My students tend to think that catastrophic risk arguments are designed as (or at least function as) distractions from the “real” issues (jobs, environmental impact, etc.) and that they should not be taken seriously as a result. Nearly all of my students would be very happy to pause AI today, but none of them in response to catastrophic risk.
3) When forced to grapple with the arguments seriously, different discussions with my students independently converged on the view that existential risk was probably a good thing because ordinary people otherwise have zero leverage over AI. That is, if AI leaders genuinely believe there is existential risk (and that, therefore, they will personally die), then this might be a way to get them to stop developing job-killing AI that will otherwise take down the rest of us.
The AI labs most willing to take costly actions now (like hire lots of safety researchers or support AI regulation that the rest of the industry opposes or make advance commitments about the preparations they’ll take before releasing future models) are also the ones talking the most about catastrophic or existential risks.
Are these actually costly actions to any meaningful degree? In the context of the amount of money sloshing around the AI space, hiring even “lots” of safety researchers seems like a rounding error.
I may misunderstand the commitments you’re referring to, but I think these are all purely internal? And thus not really commitments at all.
Like if you thought this stuff was an underhanded tactic to drum up hype and get commercial success by lying to the public, then it’s strange that Meta AI, not usually known for its tremendous moral integrity, is so principled about telling the truth that they basically never bring these risks up!
This seems to presume that I have some well-formed views on how AI labs compare, and I don’t have those. All I really know about Meta is that they’re behind and doing open source. I wouldn’t even know where to start an analysis of their relative level of moral integrity. So far as it goes (and, again, this is just the view of someone that reads what breaks through in mainstream news coverage), I have a very clear sense that OpenAI is run by compulsive liars but not much more to go on beyond that other than a general sense that people in the industry do a lot of hype.
People often quit their well-paying jobs at AI companies in order to speak out about existential risk or for reasons of insufficient attention paid to AI safety from catastrophic or existential risks.
I’m deliberately not looking this up and telling you my impression of this phenomenon. I’m coming up with three cases of it (my recollection is maybe garbled) that broke though into my media universe:
My understanding is that Anthropic was formed by people who broke away from OpenAI based on “safety” concerns. But then they just founded another company doing the same thing? And they got very rich doing it. So that all has roughly zero credibility.
There was an engineer at one of the big tech companies (Google? Microsoft?) who got a lot of attention for claiming that AI had achieved sentience and deserved personhood and either quit or got fired. The universal take seemed to be that he was insane.
One of the people involved in AI 2027 had quit or gotten fired from OpenAI(?) and refused to sign an NDA that would have come with a big payday so that he could go public with criticism. That seems pretty sincere and credible so far as it goes, but it’s also one person. And then AI 2027 was so overwrought that I couldn’t take it seriously.
And then, beyond that, you seem to have a lot of people signing these open letters with no cost attached. For something like this to breakthrough, it needs to be (in my estimation at least) large numbers of people acting in a coordinated way and leaving the industry entirely.
I’d analogize it to politics. In any given presidential administration, you have one or two people who get really worked up and resign angrily and then go on TV attacking their former bosses. That’s just to be expected and doesn’t really reflect anything beyond the fact that sometimes people have strong reactions or particularized grievances or whatever. The thing that (should) wake you up is when this is happening at scale.
Are there arguments or evidence that would have convinced you the existential risk worries in the industry were real / sincere?
Only steps that carry meaningful financial consequences. I agree that any individual researcher can send a credible signal by quitting and giving up their stock, at least to the extent they don’t just immediately go into a similarly compensated position. But, you’re always left with the counter-signal from all the other researchers not doing that.
On a more institutional level, it would have to be something that actually threatens the valuation of the companies.
I you believe “there’ll probably be warning shots”, that’s an argument against “someone will get to build It”, but not an argument against “if someone built It, everyone would die.” (where “it” specifically means “an AI smart enough to confidently outmaneuver all humanity, built by methods similar to today where they are ‘organically grown’ in hard to predict ways”).
It’s a bit of both.
Suppose there are no warning shots. A hypothetical AI that’s a a bit weaker than humanity but still awfully impressive doesn’t do anything at all that manifests an intent to harm us. That could mean:
The next, somewhat more capable of this AI will not have any intent to harm us because through either luck or design we’ve ended up with a non-threatening AI.
This version of the AI is biding its time to strike and is sufficiently good at deception that we miss that fact.
This AI is fine, but making it a little smarter/more capable will somehow lead to the emergence of malign intent.
I take Yudkowsky and Soares to put all the weight on #2 and #3 (with, based on their scenario, perhaps more of it on #2).
I don’t think that’s right. I think if we have reached the point where an AI really could plausibly start and win a war with us and it doesn’t do anything nasty, there’s a fairly good chance we’re in #1. We may not even really understand how we got into #1, but sometimes things just work out.
I’m not saying this is some kind of great strategy for dealing with the risk; the scenario I’m describing is one where there’s a real chance we all die and I don’t think you get a strong signal until you get into the range where the AI might win, which is a bad range. But it’s still very different than imagining the AI will inherently wait to strike until it has ironclad advantages.
Because LLMs are already avoiding being shut down: https://arxiv.org/abs/2509.14260 .
Very interest, thanks. As I said in the review, I wish there was more of this kind of thing in the book.
If your terminal goal is to enjoy watching a good movie, you can’t achieve it if you’re dead/shut down.
If your terminal goal is for you to watch the movie, then sure. But if your terminal goal is that the movie be watched, then shutting you down might well be perfectly consistent with it.
Ok, let’s say there is an “in between” period, and let’s say we win the fight against a misaligned AI. After the fight, we will still be left with the same alignment problems, as other people in this thread pointed out. We will still need to figure out how to make safe, benevolent AI, because there is no guarantee that we will win the next fight, and the fight after that, and the one after that, etc
At that point, the shut down argument is no longer speculative, and you can probably actually do it.
To be clear, I’m not saying that’s a good plan if you can foresee all the developments in advance. But, if you’re uncertain about all of it, then it seems like there is likely to be a period of time before it’s necessarily too late when a lot of the uncertainty is resolved.
No, I can’t. And I suspect that if the authors conducted a more realistic political analysis, the book might just be called “Everyone’s Going to Die.”
But, if you’re trying to come up with an idea that’s at least capable of meeting the magnitude of the asserted threat, then you’d consider things like:
Find a way to create a world government (a nigh-impossible ask to be sure) and then use it to ban AI.
Force anyone with relevant knowledge of how to build an AI to go into some kind of tech-free monastery and hunt anyone who refuses down with ten times the ferocity used in going after Al Qaeda after 9/11.
And then you just have to bite the bullet and accept that if these entail a risk of a nuclear war with China, then you fight a nuclear war with China. I don’t think either of those would really work out either, but at least they could work out.
If there is some clever idea out there for how to achieve an AI shutdown, I suspect it involves some way of ensuring that developing AI is economically unprofitable. I personally have no idea how to do that, but unless you cut off the financial incentive, someone’s going to do it.
Yea, I get that.
That said, they’re clearly writing the book for this moment and so it would be reasonable to give some space to what’s going with AI at this moment and what is likely to happen within the foreseeable future (however long that is). Book sales/readership follow a rapidly decaying exponential and so the fact that such information might well be outdated to the point of irrelevance in a few years shouldn’t really hold them back.
But, in those cases, it’s most likely better for the AI to wait, and it will know that it’s better to wait, until it gets more powerful.
But why? People foolishly start wars all the time, including in specific circumstances where it would be much better to wait.
(A counterargument here is “an AI might want to launch a pre-emptive strike before other more powerful AIs show up”, which could happen. But, if we win that war, we’re still left with “the sort of tools that can constrain a near-human superintelligence, would not obviously apply to a much smarter AI”, and we still have to solve the same problems.)
Or, having fought a “war” with an AI, we have relatively clear, non-speculative evidence about the consequences of continuing AI development. And that’s the point where you might actually muster politically will to cut that off in the future and take the steps necessary for that to really work.
Did the book convince you that if superintelligence is built in the next 20 years (however that happens, if it does, and for at least some sensible threshold-like meaning of “superintelligence”), then there is at least a 5-10% chance that as a result literally everyone will literally die?
I’m much more in the world of Knightian uncertainty here (i.e., it could happen but I have no idea how to quantify that) than in one where I feel like I can reasonably collapse it into a clear, probabilistic risk. I am persuaded that this is something that cannot be ruled out.
I have the sense that rationalists think there’s a a very important distinction between “literally everyone will die” and, say, “the majority of people will suffer and/or die.” I do not share that sense, and to me, the burden of proof set by the title is unreasonably high.
I’ll assent to the statement that there’s at least a 10% chance of something very bad happening, where “very bad” means >50% of people dying or experiencing severe suffering or something equivalent to that.
I think this kind of claim is the crux for motivating some sort of global ban or pause on rushed advanced general AI development in the near future (as an input to policy separate from the difficulty of actually making this happen). Or for not being glad that there is an “AI race” (even if it’s very hard to mitigate). So it’s interesting if your “not sure on existential risk” takeaway is denying or affirming this claim.
Give me a magic, zero-side effect pause button, and I’ll hit it instantly.
Is there a way a political scientist (specializing in international security and diplomacy) could help you make AI safer? If so, please let me know.
Inspired by a recent post by Philipp Risius noting a disconnect between his beliefs and his actions that I also feel.
I’m a tenured political science professor, meaning that I have a great degree of freedom to set my own research agenda (but also the kind of inertia that tenured people have). I think AI safety is massively important, and I’m still just working on the (non-AI) questions that have always interested me. I have no idea how exactly someone like me can contribute to AI safety, and I seem like a poor fit for the various fellowships and programs out there. So, I’m sending this appeal out to you, dear reader.
Is there a question that a political scientist can research (in the slow, grinding way that academics do research) that is load bearing for your work and meaningful to safety? Tell me, and maybe in six months you can have something back. Or, if you’re someone doing serious AI safety work and you’d fine it useful to borrow a political scientist to chat for an hour, let me know and I’m available (or to look at your written product the way an academic would). I’m here pseudonymously but I’m perfectly happy for you to know my name if that’s relevant to you. I’m just not going to post it in this forum.