So I want to be clear on my own position: I think joining and staying at an AGI company is prima facie strong evidence of poor character. The argument is really straightforward:
Having a high chance of killing many people nonconsensually is evil.
People in institutions doing very evil things are often bad people.
Thus, if you’re in an institution doing very evil things, this provides strong evidence you’re a bad person.
This strong evidence is not absolute, and can be overruled by the specifics of your situation. For example, Andrei Sakharov and Stanislav Petrov are both good people (likely far better people than me). This is true even though the Soviet nuclear program and Soviet missile command likely greatly increased the risks to humanity in general and to many millions or even billions of specific humans specifically.
As another example, Jeffrey Wigand joined a tobacco company as a scientist partially for financial reasons and partially because they recruited his scientific expertise to help them make a “safer cigarette,” reducing carcinogens and other health costs of smoking. His safer cigarettes research unsurprisingly didn’t lead to much, but he eventually becoming a whistleblower for the tobacco industry’s deceptive practices.
So it is not impossible for specific people to do good things in an evil situation, but it’s rare and takes a great deal of courage and luck. Most people in the Soviet military are not Petrov, and most tobacco scientists are not Wigand.
The base rate is very much against any specific person in an AGI lab being or doing good by dint of their work, nor do I see sufficiently strong evidence among individuals I know at AGI companies to override this prior (and indeed probably more evidence in favor of the original hypothesis, on balance).
Similarly, I’m not saying that being in an AI company necessarily means you’re interpersonally a bad person (though the manifold in character is real, and I do expect a correlation). But even granting that, I think the balance of evidence and reasons is clear: I’m sure many people in AGI companies are polite and nice to their friends, recycle, tip waiters well, don’t cheat on their partners, don’t reply-all to emails, and use the right pronouns. However, this does not make you a good person. Analogously, a tobacco marketer can be nice to his romantic partner, recycle, tip well, say “please” and “thank you”, return the shopping carts to the right locations, try to avoid micro-aggressions, stand to the right on escalators, like and share substack posts they enjoy, refuse to jaywalk, etc – while spending the majority of his waking hours on figuring out more ways to get teenagers addicted to cancer sticks. That marketer is not a good person. Interpersonal niceness is not sufficient to be good.
I want to be clear and unambiguous with my analysis because I think in these parts people have strong economic and social incentives to be wishy-washy and oblique with their critiques of people at AI companies, or even to not be critical at all. I think these incentives are corrosive, and I want to make it easier for others to resist them.
If you’re currently my friend or acquaintance and you work at an AGI lab, please understand that from my perspective I’ve already priced in your occupation in our relationship, for better or for worse. From my perspective I see no significant reason to re-evaluate our relationship, though you’re welcome to do so if my views here are news to you.
However, if you’re thinking about a career change, I’m happy to discuss it with you and maintain confidentiality.
Suppose that the only AGI labs were Anthropic who has a marginally better chance to align its AI and xAI-like labs (e.g. xAI itself, Chinese labs) who don’t care about alignment at all. In order to rule out the appearance of a misaligned ASI, the world would have to destroy any sufficiently resourced lab. If the world keeps both Anthropic and xAI-like labs, then attempting to slow Anthropic down won’t make mankind more likely to survive. I find it unlikely that anyone here has any ways to attack xAI without involving the USG.
As for attempts to claim that “People in institutions doing very evil things are often bad people” and citing the Soviet military, but not Sakharov and Petrov as examples, I struggle to understand it. Where exactly does the boundary lay where most good people don’t enter or end up self-modified à-la inhabitants of moral mazes? How is the same argument applicable or inapplicable to American military? If it’s inapplicable, then what difference is legible to the soldiers themselves?
How is the same argument applicable or inapplicable to American military?
Thanks, coming up with appropriate analogies was hard. I did consider Hugh Thompson instead, but thought that it’d confuse the message more. (There’s also an issue of his bravery in resisting massacres by his allies during the heat of battle probably being structurally very different than the type of bravery that scientists and engineers in an AI company need to do the right thing).
I wanted examples that underlines the scope and moral seriousness of the charges (I think understating it would undercut my own message), but didn’t want to give example that are mind-killed or people in my likely audience would pattern-match to “cartoonishly evil”.
At the same time I didn’t want to get into debates about which non-AI institutions are “actually evil.”
If you take the view that the AI companies as institutions are evil and so everyone who works for them is bad, sure, fine. But do you apply this principle generally, or is this special pleading? Just about all large institutions collaborate with the state, and plenty of people work directly for it. Do you make the same judgments about their moral character as well? USAID might be a particularly fun example to consider, but it should work with any, say, $1B+ market cap company.
Besides the points made above (that the activities of major labs are more dangerous/reprehensible than those of current world governments), I would also point out that AI labs are much less subject to confusion/diffusion of responsibility than states are. The US government has ~2.68m non-military employees whose activities span the gamut from “make sure old people don’t starve to death” to “torturing prisoners”. While people do continue to try, it has proven very difficult to extricate the bad activities from the good, and so it is very difficult to say that the US government, as a whole, is evil, even if many parts of it, and many of its activities, certainly may be. So an employee could easily say “I’m not working on the bad stuff, and it’s not like me succeeding here will make the bad stuff more successful” and be correct about that.
But with labs like OpenAI and Anthropic, this isn’t the case. There are a few thousand people working there, of which a smaller subset are working directly on research/software/infrastructure. These companies essentially do one thing, which is build/train/serve increasingly capable models and products that leverage them. While the direct impacts of these product offerings do vary, all profits made thereby will feed into training larger models, and the success of these product offerings largely depends on those larger models. The only people that come close to escaping this logic are those working on safety, but to a large extent that seems to operate more like harm reduction than an actual counterbalance to the bad activities of the broader organization; especially considering how their efforts often just serve to enable further development of model capabilities. An analogy would be if the military were to invade another country without provocation, destroying its infrastructure in the process, but thereafter deploying the Army Corps of Engineers to rebuild key portions like bridges and the water supply in order to facilitate a long-term occupation. We could agree that it’s better for the occupied population to have water, but that it would have been better for the invasion to have never happened in the first place, thus making participation in the ACoE morally dubious there. So altogether, it is much easier to assign moral responsibility to people working at the labs than people working for the government: just about everyone involved is fundamentally working to enable something very-bad, they mostly understand that this is very bad, and yet they freely choose to do so, in no small part thanks to how they perceive it as affording personal gain.
I’m less negative about the state in general and maybe USAID specifically than you are. Something that’s implicit here is that judgments of evil are agent-relative, if not in absolute terms (eg because you’re a realist), then at minimum epistemically.
For example, if somebody thinks working at AI companies is great (and this may be plausible for a variety of both prosaic and cosmic reasons), they can then reasonably infer that my comment is predictably evil and this speaks poorly of my character.
I basically buy that this is plausible. Definitely some people, including some people that external observers would think are reasonable, would hold it. My main meta-level defense against my comment being evidence-free is that I think human motivations interests etc are not 100% diametrically opposed, such that if everybody speaks clearly about their motivations and reasons and judgments I do not expect it to net out to zero (or even very close).
Well, the state isn’t planning to kill me personally, and anthropic is planning to at least give the ol russian roulette wheel a spin in that regard. They expect that if they don’t someone else will. Or maybe the state is, depending on how stable we think the current MAD standoffs with Russia and China are. The analogy to the soviet nuclear program is pretty good, in that they were also explicitly gambling with everyone’s lives for complicated game theory reasons.
I think this mostly checks out, but I think this corollary follows if you replace bullet point 3 and 4 in the argument:
3. Paying money to evil institutions lets them do more evil things. 4. If you’re paying money to these institutions, it’s strong evidence you’re a bad person
If you claim that the net benefit of using their products is more than the net evil your additional dollar causes, then AI Co employees can claim the same thing! Maybe they’re offsetting their evil by donating most of their salary to reducing x-risk or something.
You’re right that the base rate is low, and I predict that anyone claiming they would use AI to reduce x-risk is similarly low to them being a Andrei Sakharov or a Stanislav Petrov.
I think FWIW the argument makes sense, but you can also straightforwardly apply it for anyone with a Claude or ChatGPT subscription too.
Some of my friends think my public criticisms of AGI companies and AI researchers, especially in satire, are too oblique. There has also been some confusion in the LessWrong discussions on AGI companies and character.
So I want to be clear on my own position: I think joining and staying at an AGI company is prima facie strong evidence of poor character. The argument is really straightforward:
The AGI companies are on the path to building systems that are powerful, agentic, destructive, and have a high chance of killing many people.
Having a high chance of killing many people nonconsensually is evil.
People in institutions doing very evil things are often bad people.
Thus, if you’re in an institution doing very evil things, this provides strong evidence you’re a bad person.
This strong evidence is not absolute, and can be overruled by the specifics of your situation. For example, Andrei Sakharov and Stanislav Petrov are both good people (likely far better people than me). This is true even though the Soviet nuclear program and Soviet missile command likely greatly increased the risks to humanity in general and to many millions or even billions of specific humans specifically.
As another example, Jeffrey Wigand joined a tobacco company as a scientist partially for financial reasons and partially because they recruited his scientific expertise to help them make a “safer cigarette,” reducing carcinogens and other health costs of smoking. His safer cigarettes research unsurprisingly didn’t lead to much, but he eventually becoming a whistleblower for the tobacco industry’s deceptive practices.
So it is not impossible for specific people to do good things in an evil situation, but it’s rare and takes a great deal of courage and luck. Most people in the Soviet military are not Petrov, and most tobacco scientists are not Wigand.
The base rate is very much against any specific person in an AGI lab being or doing good by dint of their work, nor do I see sufficiently strong evidence among individuals I know at AGI companies to override this prior (and indeed probably more evidence in favor of the original hypothesis, on balance).
Similarly, I’m not saying that being in an AI company necessarily means you’re interpersonally a bad person (though the manifold in character is real, and I do expect a correlation). But even granting that, I think the balance of evidence and reasons is clear: I’m sure many people in AGI companies are polite and nice to their friends, recycle, tip waiters well, don’t cheat on their partners, don’t reply-all to emails, and use the right pronouns. However, this does not make you a good person. Analogously, a tobacco marketer can be nice to his romantic partner, recycle, tip well, say “please” and “thank you”, return the shopping carts to the right locations, try to avoid micro-aggressions, stand to the right on escalators, like and share substack posts they enjoy, refuse to jaywalk, etc – while spending the majority of his waking hours on figuring out more ways to get teenagers addicted to cancer sticks. That marketer is not a good person. Interpersonal niceness is not sufficient to be good.
I want to be clear and unambiguous with my analysis because I think in these parts people have strong economic and social incentives to be wishy-washy and oblique with their critiques of people at AI companies, or even to not be critical at all. I think these incentives are corrosive, and I want to make it easier for others to resist them.
If you’re currently my friend or acquaintance and you work at an AGI lab, please understand that from my perspective I’ve already priced in your occupation in our relationship, for better or for worse. From my perspective I see no significant reason to re-evaluate our relationship, though you’re welcome to do so if my views here are news to you.
However, if you’re thinking about a career change, I’m happy to discuss it with you and maintain confidentiality.
I don’t work at any AGI lab, but I am confused for reasons similar to The Counterfactual Quiet AGI Timeline, which I brought up in comments when @Richard_Ngo devoted an entire sequence to proving a similar thesis. I made a whole post as a response to Ngo’s thesis.
Suppose that the only AGI labs were Anthropic who has a marginally better chance to align its AI and xAI-like labs (e.g. xAI itself, Chinese labs) who don’t care about alignment at all. In order to rule out the appearance of a misaligned ASI, the world would have to destroy any sufficiently resourced lab. If the world keeps both Anthropic and xAI-like labs, then attempting to slow Anthropic down won’t make mankind more likely to survive. I find it unlikely that anyone here has any ways to attack xAI without involving the USG.
As for attempts to claim that “People in institutions doing very evil things are often bad people” and citing the Soviet military, but not Sakharov and Petrov as examples, I struggle to understand it. Where exactly does the boundary lay where most good people don’t enter or end up self-modified à-la inhabitants of moral mazes? How is the same argument applicable or inapplicable to American military? If it’s inapplicable, then what difference is legible to the soldiers themselves?
Thanks, this is a good critique.
Thanks, coming up with appropriate analogies was hard. I did consider Hugh Thompson instead, but thought that it’d confuse the message more. (There’s also an issue of his bravery in resisting massacres by his allies during the heat of battle probably being structurally very different than the type of bravery that scientists and engineers in an AI company need to do the right thing).
I wanted examples that underlines the scope and moral seriousness of the charges (I think understating it would undercut my own message), but didn’t want to give example that are mind-killed or people in my likely audience would pattern-match to “cartoonishly evil”.
At the same time I didn’t want to get into debates about which non-AI institutions are “actually evil.”
If you take the view that the AI companies as institutions are evil and so everyone who works for them is bad, sure, fine. But do you apply this principle generally, or is this special pleading? Just about all large institutions collaborate with the state, and plenty of people work directly for it. Do you make the same judgments about their moral character as well? USAID might be a particularly fun example to consider, but it should work with any, say, $1B+ market cap company.
Besides the points made above (that the activities of major labs are more dangerous/reprehensible than those of current world governments), I would also point out that AI labs are much less subject to confusion/diffusion of responsibility than states are. The US government has ~2.68m non-military employees whose activities span the gamut from “make sure old people don’t starve to death” to “torturing prisoners”. While people do continue to try, it has proven very difficult to extricate the bad activities from the good, and so it is very difficult to say that the US government, as a whole, is evil, even if many parts of it, and many of its activities, certainly may be. So an employee could easily say “I’m not working on the bad stuff, and it’s not like me succeeding here will make the bad stuff more successful” and be correct about that.
But with labs like OpenAI and Anthropic, this isn’t the case. There are a few thousand people working there, of which a smaller subset are working directly on research/software/infrastructure. These companies essentially do one thing, which is build/train/serve increasingly capable models and products that leverage them. While the direct impacts of these product offerings do vary, all profits made thereby will feed into training larger models, and the success of these product offerings largely depends on those larger models. The only people that come close to escaping this logic are those working on safety, but to a large extent that seems to operate more like harm reduction than an actual counterbalance to the bad activities of the broader organization; especially considering how their efforts often just serve to enable further development of model capabilities. An analogy would be if the military were to invade another country without provocation, destroying its infrastructure in the process, but thereafter deploying the Army Corps of Engineers to rebuild key portions like bridges and the water supply in order to facilitate a long-term occupation. We could agree that it’s better for the occupied population to have water, but that it would have been better for the invasion to have never happened in the first place, thus making participation in the ACoE morally dubious there. So altogether, it is much easier to assign moral responsibility to people working at the labs than people working for the government: just about everyone involved is fundamentally working to enable something very-bad, they mostly understand that this is very bad, and yet they freely choose to do so, in no small part thanks to how they perceive it as affording personal gain.
I’m less negative about the state in general and maybe USAID specifically than you are. Something that’s implicit here is that judgments of evil are agent-relative, if not in absolute terms (eg because you’re a realist), then at minimum epistemically.
For example, if somebody thinks working at AI companies is great (and this may be plausible for a variety of both prosaic and cosmic reasons), they can then reasonably infer that my comment is predictably evil and this speaks poorly of my character.
I basically buy that this is plausible. Definitely some people, including some people that external observers would think are reasonable, would hold it. My main meta-level defense against my comment being evidence-free is that I think human motivations interests etc are not 100% diametrically opposed, such that if everybody speaks clearly about their motivations and reasons and judgments I do not expect it to net out to zero (or even very close).
Well, the state isn’t planning to kill me personally, and anthropic is planning to at least give the ol russian roulette wheel a spin in that regard. They expect that if they don’t someone else will. Or maybe the state is, depending on how stable we think the current MAD standoffs with Russia and China are. The analogy to the soviet nuclear program is pretty good, in that they were also explicitly gambling with everyone’s lives for complicated game theory reasons.
I agree basically in full, and endorse this post.
I think this mostly checks out, but I think this corollary follows if you replace bullet point 3 and 4 in the argument:
3. Paying money to evil institutions lets them do more evil things.
4. If you’re paying money to these institutions, it’s strong evidence you’re a bad person
If you claim that the net benefit of using their products is more than the net evil your additional dollar causes, then AI Co employees can claim the same thing! Maybe they’re offsetting their evil by donating most of their salary to reducing x-risk or something.
You’re right that the base rate is low, and I predict that anyone claiming they would use AI to reduce x-risk is similarly low to them being a Andrei Sakharov or a Stanislav Petrov.
I think FWIW the argument makes sense, but you can also straightforwardly apply it for anyone with a Claude or ChatGPT subscription too.