There’s a loop that consequentialists can get into that’s something like:
“I’m altruistic, and altruistic people should have more power”
“Now that I’ve gained power, the world is better off because I’ll do good things with it”
“Now that the world is better off, I have even more evidence for my altruism”
(return to step 1)
The problem here is that you’re not cashing out “the world being improved” in anything foundational, but rather in large part in the fact that you believe in yourself. When I think about Anthropic considering how they should orient to a pause, for example, I imagine them thinking “but actually we’re so altruistic that it’s bad for the world if we lose ground, so we shouldn’t pay significant costs in supporting a pause”. And so the impact gets pushed off for another cycle...
What is foundational enough that we can trust it to ground claims about who should have a lot of power? For instance, if Anthropic had made some scientific breakthroughs regarding alignment, then I’d feel significantly better about their existence. This is why I talk so much about science in this post, because I think that if you use a scientific lens then there’s very little “high-quality alignment research” coming out of Anthropic. They have done a bunch of stuff to affect current Claudes, but the whole question is how much it will generalize, and my current sense is that Anthropic is too focused on iterative empirical work to actually do the kinds of things that might generalize a long way (aside from their mechinterp team, but even they pivoted away from fundamental stuff towards being short-term useful with SAEs).
Another kind of foundation is “these people are so high-integrity and honest that we can actually believe that, when they make the world locally worse (e.g. via racing) in exchange for more power, they’re not deceiving themselves or us about why they’re doing so”. But again this would require pretty different evidence than Anthropic are currently willing to provide (e.g. like credible evidence about the process by which they changed their strategy over time vis a vis racing).
at least more than would have been done by the counterfactual frontier AI company
This kind of reasoning is extremely vulnerable to adversarial dynamics. If my friend is about to mug you, and I come in and say “hey, I’m going to mug you instead, but take less of your money, so I’m helping you”, then only a schmuck says “thanks”. (Unfortunately, the AI safety community is consistently a schmuck in this sense.) The correct approach is to instead hold agents accountable at multiple scales at once. For example, you might evaluate both the counterfactual where only Anthropic doesn’t exist, but then also the counterfactual where neither Anthropic nor OpenAI exist, and also the counterfactual where no companies are racing towards AGI, and so on.
Of course it’s hard to reason about the details of the evaluation, but I think people intuitively do this on a social level, and this lets them do credit assignment in a much better way than when doing explicit consequentialist reasoning (which typically bakes in individualism in how it chooses counterfactuals). Another way of putting this which I found to spark some insight: you should sometimes view “competitors” as “collaborators in making the competition between them grow to larger scales”. More here.
So I would be interested to hear more about why you think things are worse than they would be in a counterfactual where AI safety people did not push forward AI capabilities and join frontier AI companies.
And yeah, my impression is that not much has come out of agent foundations, which feeds into the counterfactual
This is covered in much more detail in my next post, so ask me again after that if it’s still unclear.
Re the “agent foundations” thing, one of the points I was trying to convey in the latter half of this post is that agent foundations never really lost the debate, they instead lost the prestige and influence game against OpenPhil and OpenAI. The more I’ve looked into agent foundations, the more MIRI’s results seem like important scientific progress—which unfortunately is very hard to reason about in terms of expected impact (for reasons related to this reply to Leo). I mostly don’t intend to try to convey this in writing, because the inferential gap is too high, but here’s one relevant post about how I came to appreciate Garrabrant induction much more.
There’s a loop that consequentialists can get into that’s something like:
“I’m altruistic, and altruistic people should have more power”
Is this a problem with consequentialism, or more of a problem with overconfidence/self-deception in one’s altruism? I feel like a lot of non-consequentialists who are very confident in their own goodness can also get into a similar loop, where they use their ingroup’s goodness as reason/excuse to gain more and more power. Various religions and political movements (such as wokeness) seem to be examples of this.
To paraphrase Weinberg: With or without consequentialism, good people can support good things and bad people can support evil things; but for good people to support evil things—that takes consequentialism.
Obviously this is a very blunt and stylized idea, but I do think that consequentialism is a particularly effective subversion of people who otherwise would have been guided in good directions by their moral instincts.
If you looked at non-consequentialist ideologies that was contemporaneous with Bentham, so many of them advocated for policy positions that we now consider evil.
Nor are non-ideological people immune. Plenty of “ordinary men” do evil things when people around them do.
I reject the paraphrased claim in at least a couple ways:
One, deontological or virtue ethics induce people to make bad choices when it comes to trolley type problems, compared to consequentialists who make the better choice(s).
Two, it begs the question by presupposing evil people instead of fallible humans. I expect this is downstream of broader mistake v conflict theory stuff, and don’t expect us to resolve ethics in this comment section, but seems worth calling out foundational disagreements as being sufficient to explain differing conclusions.
I think there’s two things going on in different people.
The “I’m altruistic, and altruistic people should have more power” thing feels very much like the kind of generic self-justification thing (“I’m $INGROUP so therefore I’m good, and good people should have more power”) that will make anyone feel their power-grab is justified. Many secular Westerners have been taught that “I’m of the True Religion” or “I’m of the Best Nationality” aren’t acceptable justifications for claiming your own superiority over others, but “I’m altruistic” sounds non-ingroupy so it bypasses that immune response.
Knowing about consequentialism can help further rationalize that, but without the self-serving instinct one wouldn’t be motivated to generate the “this is why I should have more power” story and accompanying rationalization in the first place.
But then, if someone is mostly altruistic and someone else has already generated an elaborate “this is why we should grab power” rationalization, the altruistic person may not be able to notice the flaws in the logic, and their consequentialism will then drive them to adopt that reasoning.
So you get two kinds of people: ones whose moral instincts would have been overpowered by their self-serving ones anyway, and ones whose moral instincts would have held firm if not for the story that the former group sold them on.
“There’s a loop that consequentialists can get into that’s something like: … “Now that I’ve gained power, ”
But I am not at Anthropic, and neither are any of my friends. So the dynamic does not apply to me.
And at any rate, the argument is not that ppl at Anthropic or OpenAI are not deluding themselves. It’s rather whether their effect on the world is better/worse than the counterfactual AI company. You correctly point towards the need for the “safety-consciousness” to play out in anything meaningful to count as evidence, but then you only list as a criterion that their current alignment techniques should actually generalize.
To me it is pretty clear that we have solid evidence: with their taking a stand on autonomous weapons Antropic have behaved quite clearly differently than the counterfactual AI company would have. In general, comparing frontier AI companies in general to tobacco or oil companies, the differences in ethical behavior seems very clear to me. You might of course still argue that this is in fact worse for the world, because chumps take the safety postures of frontier companies seriously, thus preventing actually meaningful political actions.
There’s a loop that consequentialists can get into that’s something like:
“I’m altruistic, and altruistic people should have more power”
“Now that I’ve gained power, the world is better off because I’ll do good things with it”
“Now that the world is better off, I have even more evidence for my altruism”
(return to step 1)
The problem here is that you’re not cashing out “the world being improved” in anything foundational, but rather in large part in the fact that you believe in yourself. When I think about Anthropic considering how they should orient to a pause, for example, I imagine them thinking “but actually we’re so altruistic that it’s bad for the world if we lose ground, so we shouldn’t pay significant costs in supporting a pause”. And so the impact gets pushed off for another cycle...
What is foundational enough that we can trust it to ground claims about who should have a lot of power? For instance, if Anthropic had made some scientific breakthroughs regarding alignment, then I’d feel significantly better about their existence. This is why I talk so much about science in this post, because I think that if you use a scientific lens then there’s very little “high-quality alignment research” coming out of Anthropic. They have done a bunch of stuff to affect current Claudes, but the whole question is how much it will generalize, and my current sense is that Anthropic is too focused on iterative empirical work to actually do the kinds of things that might generalize a long way (aside from their mechinterp team, but even they pivoted away from fundamental stuff towards being short-term useful with SAEs).
Another kind of foundation is “these people are so high-integrity and honest that we can actually believe that, when they make the world locally worse (e.g. via racing) in exchange for more power, they’re not deceiving themselves or us about why they’re doing so”. But again this would require pretty different evidence than Anthropic are currently willing to provide (e.g. like credible evidence about the process by which they changed their strategy over time vis a vis racing).
This kind of reasoning is extremely vulnerable to adversarial dynamics. If my friend is about to mug you, and I come in and say “hey, I’m going to mug you instead, but take less of your money, so I’m helping you”, then only a schmuck says “thanks”. (Unfortunately, the AI safety community is consistently a schmuck in this sense.) The correct approach is to instead hold agents accountable at multiple scales at once. For example, you might evaluate both the counterfactual where only Anthropic doesn’t exist, but then also the counterfactual where neither Anthropic nor OpenAI exist, and also the counterfactual where no companies are racing towards AGI, and so on.
Of course it’s hard to reason about the details of the evaluation, but I think people intuitively do this on a social level, and this lets them do credit assignment in a much better way than when doing explicit consequentialist reasoning (which typically bakes in individualism in how it chooses counterfactuals). Another way of putting this which I found to spark some insight: you should sometimes view “competitors” as “collaborators in making the competition between them grow to larger scales”. More here.
This is covered in much more detail in my next post, so ask me again after that if it’s still unclear.
Re the “agent foundations” thing, one of the points I was trying to convey in the latter half of this post is that agent foundations never really lost the debate, they instead lost the prestige and influence game against OpenPhil and OpenAI. The more I’ve looked into agent foundations, the more MIRI’s results seem like important scientific progress—which unfortunately is very hard to reason about in terms of expected impact (for reasons related to this reply to Leo). I mostly don’t intend to try to convey this in writing, because the inferential gap is too high, but here’s one relevant post about how I came to appreciate Garrabrant induction much more.
Is this a problem with consequentialism, or more of a problem with overconfidence/self-deception in one’s altruism? I feel like a lot of non-consequentialists who are very confident in their own goodness can also get into a similar loop, where they use their ingroup’s goodness as reason/excuse to gain more and more power. Various religions and political movements (such as wokeness) seem to be examples of this.
To paraphrase Weinberg: With or without consequentialism, good people can support good things and bad people can support evil things; but for good people to support evil things—that takes consequentialism.
Obviously this is a very blunt and stylized idea, but I do think that consequentialism is a particularly effective subversion of people who otherwise would have been guided in good directions by their moral instincts.
If you looked at non-consequentialist ideologies that was contemporaneous with Bentham, so many of them advocated for policy positions that we now consider evil.
Nor are non-ideological people immune. Plenty of “ordinary men” do evil things when people around them do.
Thomas Carlyle coined the term “dismal science” to argue in favor of slavery, against people like Mill and other economics-minded utilitarians.
https://en.wikipedia.org/wiki/The_dismal_science
I reject the paraphrased claim in at least a couple ways:
One, deontological or virtue ethics induce people to make bad choices when it comes to trolley type problems, compared to consequentialists who make the better choice(s).
Two, it begs the question by presupposing evil people instead of fallible humans. I expect this is downstream of broader mistake v conflict theory stuff, and don’t expect us to resolve ethics in this comment section, but seems worth calling out foundational disagreements as being sufficient to explain differing conclusions.
I think there’s two things going on in different people.
The “I’m altruistic, and altruistic people should have more power” thing feels very much like the kind of generic self-justification thing (“I’m $INGROUP so therefore I’m good, and good people should have more power”) that will make anyone feel their power-grab is justified. Many secular Westerners have been taught that “I’m of the True Religion” or “I’m of the Best Nationality” aren’t acceptable justifications for claiming your own superiority over others, but “I’m altruistic” sounds non-ingroupy so it bypasses that immune response.
Knowing about consequentialism can help further rationalize that, but without the self-serving instinct one wouldn’t be motivated to generate the “this is why I should have more power” story and accompanying rationalization in the first place.
But then, if someone is mostly altruistic and someone else has already generated an elaborate “this is why we should grab power” rationalization, the altruistic person may not be able to notice the flaws in the logic, and their consequentialism will then drive them to adopt that reasoning.
So you get two kinds of people: ones whose moral instincts would have been overpowered by their self-serving ones anyway, and ones whose moral instincts would have held firm if not for the story that the former group sold them on.
“There’s a loop that consequentialists can get into that’s something like: … “Now that I’ve gained power, ”
But I am not at Anthropic, and neither are any of my friends. So the dynamic does not apply to me.
And at any rate, the argument is not that ppl at Anthropic or OpenAI are not deluding themselves. It’s rather whether their effect on the world is better/worse than the counterfactual AI company. You correctly point towards the need for the “safety-consciousness” to play out in anything meaningful to count as evidence, but then you only list as a criterion that their current alignment techniques should actually generalize.
To me it is pretty clear that we have solid evidence: with their taking a stand on autonomous weapons Antropic have behaved quite clearly differently than the counterfactual AI company would have. In general, comparing frontier AI companies in general to tobacco or oil companies, the differences in ethical behavior seems very clear to me. You might of course still argue that this is in fact worse for the world, because chumps take the safety postures of frontier companies seriously, thus preventing actually meaningful political actions.