Had a nice conversation with Richard a few days ago which I will now try to summarize parts of, for those interested, including especially my future self who might want to pick up the thread later:
In theory-land, expected utility maximization seems to have some problems. Start with Pascal’s wager and move on to the commitment races problem.
In practice, there’s this common observation (which seems plausible to me though not certain) that people/communities who identify as expected utility maximizers and/or consequentialists seem to end up getting surprisingly bad outcomes surprisingly often.
Are these things related? Maybe, but maybe not; here are two other, more mundane explanations for the latter observation:
First, community norms. The EA community grew out of a bunch of Oxford philosophers who had decided that what really matters is maximizing expected utility, and the philosophy literature typically contrasts norm-following with utility-maximization (understandably since it’s intrinsically interested in the contrast) which perhaps accidentally led to the mistaken impression amongst many community members that it just isn’t that important to think about, promote, and enforce norms. The thoughtful consequentialist position of course would be that norms may be very important indeed, but alas, due to the vibes emanating from the philosophy literature, the result was a community that spent unusually little time thinking about what norms to create and promote and enforce. Alas, perhaps it turns out that this is an important path to impact for social movements and that movements which neglect this aspect of their impact can more easily end up being manipulated or coopted or otherwise messed up.
Second, rationalization. Consider how, if you are forecasting the future of some trend (e.g. AI company revenue) you have a range of options ranging from very simple “choose either exponential or linear, and then fit a trend and extrapolate it” to very complex e.g. “consider all the headwinds and tailwinds and datacenters coming online and new models they might train and total addressable market for those models and...” In theory, adding more bells and whistles to your model should enable you to eke out more predictive accuracy gains. In practice, there’s this common wisdom that you should try to keep things simple. Why? Perhaps because you are yourself biased, and it’s easier for your bias to creep in to your work if you have more degrees of freedom and judgment calls for that bias to operate on. Analogously, perhaps consequentialism / EU-maximization is systematically more bias-prone than other moral philosophies, for similar reasons. The future is hard to predict; the consequences of actions are hard to predict; in this uncertain world, it’s easy to convince yourself that almost anything is the EU-maximizing move, if you consider enough arguments for and against and then bias the weighting towards the arguments in favor. Whereas if you have some deontological constraints to obey, you can try to rationalize why some loophole doesn’t REALLY count as violating the constraint, but maybe it’s generally harder to do this? More speculatively, another way of making decisions is via predictions about the approval of others e.g. “what would my friends and family think.” Maybe it’s harder to rationalize approval because it feels more near-mode to your brain; your brain is good at making near-term predictions like this and has learned not to make bad predictions that’ll be falsified immediately. And so maybe there are moral systems that kinda hijack this machinery and use hypothetical versions e.g. “what would Jesus do.”
Whereas if you have some deontological constraints to obey, you can try to rationalize why some loophole doesn’t REALLY count as violating the constraint, but maybe it’s generally harder to do this?
You can rationalize by finding loopholes, but you can also rationalize by choosing different deontological constraints or virtues, or different interpretation of them. This seems like a somewhat serious problem? I think it’s easy to find a group of people who will all say that it’s really important to be high-integrity, but then who will strongly disagree about what being high-integrity means and maybe think that the other people in the group are acting in a low-integrity way. There’s questions about honesty binds you to something like never lying, or to something more like always giving people an accurate impression of all important facts, or something else. People disagree a lot about whether there’s deontological reasons to not advance AI capabilities on the current margin. Or whether there’s deontological reasons to not eat meat.
It’s reasonably easy to pick a convenient position here, because the whole debate about how to pick virtues or deontological principles is pretty hard to ground? (Like the consequentialist debate feels more grounded in empirical disagreements that are easier to operationalize, although in practice I agree that there’s enough hard calls there that it’s easy to rationalize things.)
I think making decisions via predictions about others’ approval has less of this problem, because you can just check. Though it has other problems, like ideally you want to be able to outperform your friends and family. (But it seems useful to at least check whether an action would horrify your friends or family. Though even then, e.g. ‘religious deconversion’ may well horrify some people even when it’s the right thing to do.)
good points. But a counterpoint: If a community forms around a certain set of virtues and rules at time T, then that set becomes sticky and harder to change and therefore harder to twist into letting you do the power-seeking or cowardly or status-seeking thing later. For people who don’t have any community or ethical identity, yes, they can choose from the menu the one that most helps them get status power etc. But then once they’ve chosen it, there are switching costs. By contrast if you are a consequentialist, you can simply convince yourself that actually you need to JOIN the AGI company, or that actually you need to ACCELERATE chip production, or whatever. The evidence about what’ll have good vs. bad consequences is constantly changing, so if you change your mind for rationalizing reasons it can be disguised (to yourself and others) as a change of mind based on good evidence.
I guess this could be different on a community level.
Maybe it’s more tractable and/or desirable to run a community that enforces conformance to the same rules/virtues (including interpretations and case-law about them) than it is to run a community that enforces conformance to the same beliefs about what actions have very good vs. very harmful consequences.
But even the former seems tough for the cases where it’s hard to conclusively argue for any particular set of rules/virtues. It’s rough to exile some contingent of people who just have some reasonable disagreements; and if you do you might just get splinter groups which still interact a lot in practice.
In practice, there’s this common wisdom that you should try to keep things simple. Why?
I would have said the most important reason is that you (and people you talk with) can hold all of the simple model in your head and learn the right lessons from it. Whereas if you add too much complexity, you can’t track what’s going on anymore, so it’s hard to do anything with the model other than just deferring to it, and the model probably isn’t so good that you should just defer to it.
Had a nice conversation with Richard a few days ago which I will now try to summarize parts of, for those interested, including especially my future self who might want to pick up the thread later:
In theory-land, expected utility maximization seems to have some problems. Start with Pascal’s wager and move on to the commitment races problem.
In practice, there’s this common observation (which seems plausible to me though not certain) that people/communities who identify as expected utility maximizers and/or consequentialists seem to end up getting surprisingly bad outcomes surprisingly often.
Are these things related? Maybe, but maybe not; here are two other, more mundane explanations for the latter observation:
First, community norms. The EA community grew out of a bunch of Oxford philosophers who had decided that what really matters is maximizing expected utility, and the philosophy literature typically contrasts norm-following with utility-maximization (understandably since it’s intrinsically interested in the contrast) which perhaps accidentally led to the mistaken impression amongst many community members that it just isn’t that important to think about, promote, and enforce norms. The thoughtful consequentialist position of course would be that norms may be very important indeed, but alas, due to the vibes emanating from the philosophy literature, the result was a community that spent unusually little time thinking about what norms to create and promote and enforce. Alas, perhaps it turns out that this is an important path to impact for social movements and that movements which neglect this aspect of their impact can more easily end up being manipulated or coopted or otherwise messed up.
Second, rationalization. Consider how, if you are forecasting the future of some trend (e.g. AI company revenue) you have a range of options ranging from very simple “choose either exponential or linear, and then fit a trend and extrapolate it” to very complex e.g. “consider all the headwinds and tailwinds and datacenters coming online and new models they might train and total addressable market for those models and...” In theory, adding more bells and whistles to your model should enable you to eke out more predictive accuracy gains. In practice, there’s this common wisdom that you should try to keep things simple. Why? Perhaps because you are yourself biased, and it’s easier for your bias to creep in to your work if you have more degrees of freedom and judgment calls for that bias to operate on. Analogously, perhaps consequentialism / EU-maximization is systematically more bias-prone than other moral philosophies, for similar reasons. The future is hard to predict; the consequences of actions are hard to predict; in this uncertain world, it’s easy to convince yourself that almost anything is the EU-maximizing move, if you consider enough arguments for and against and then bias the weighting towards the arguments in favor. Whereas if you have some deontological constraints to obey, you can try to rationalize why some loophole doesn’t REALLY count as violating the constraint, but maybe it’s generally harder to do this? More speculatively, another way of making decisions is via predictions about the approval of others e.g. “what would my friends and family think.” Maybe it’s harder to rationalize approval because it feels more near-mode to your brain; your brain is good at making near-term predictions like this and has learned not to make bad predictions that’ll be falsified immediately. And so maybe there are moral systems that kinda hijack this machinery and use hypothetical versions e.g. “what would Jesus do.”
You can rationalize by finding loopholes, but you can also rationalize by choosing different deontological constraints or virtues, or different interpretation of them. This seems like a somewhat serious problem? I think it’s easy to find a group of people who will all say that it’s really important to be high-integrity, but then who will strongly disagree about what being high-integrity means and maybe think that the other people in the group are acting in a low-integrity way. There’s questions about honesty binds you to something like never lying, or to something more like always giving people an accurate impression of all important facts, or something else. People disagree a lot about whether there’s deontological reasons to not advance AI capabilities on the current margin. Or whether there’s deontological reasons to not eat meat.
It’s reasonably easy to pick a convenient position here, because the whole debate about how to pick virtues or deontological principles is pretty hard to ground? (Like the consequentialist debate feels more grounded in empirical disagreements that are easier to operationalize, although in practice I agree that there’s enough hard calls there that it’s easy to rationalize things.)
I think making decisions via predictions about others’ approval has less of this problem, because you can just check. Though it has other problems, like ideally you want to be able to outperform your friends and family. (But it seems useful to at least check whether an action would horrify your friends or family. Though even then, e.g. ‘religious deconversion’ may well horrify some people even when it’s the right thing to do.)
good points. But a counterpoint: If a community forms around a certain set of virtues and rules at time T, then that set becomes sticky and harder to change and therefore harder to twist into letting you do the power-seeking or cowardly or status-seeking thing later. For people who don’t have any community or ethical identity, yes, they can choose from the menu the one that most helps them get status power etc. But then once they’ve chosen it, there are switching costs. By contrast if you are a consequentialist, you can simply convince yourself that actually you need to JOIN the AGI company, or that actually you need to ACCELERATE chip production, or whatever. The evidence about what’ll have good vs. bad consequences is constantly changing, so if you change your mind for rationalizing reasons it can be disguised (to yourself and others) as a change of mind based on good evidence.
simple solution: if you’re consequentialist, you should be dogmatically attached to your beliefs and refuse to change them regardless of the evidence.
I guess this could be different on a community level.
Maybe it’s more tractable and/or desirable to run a community that enforces conformance to the same rules/virtues (including interpretations and case-law about them) than it is to run a community that enforces conformance to the same beliefs about what actions have very good vs. very harmful consequences.
But even the former seems tough for the cases where it’s hard to conclusively argue for any particular set of rules/virtues. It’s rough to exile some contingent of people who just have some reasonable disagreements; and if you do you might just get splinter groups which still interact a lot in practice.
I would have said the most important reason is that you (and people you talk with) can hold all of the simple model in your head and learn the right lessons from it. Whereas if you add too much complexity, you can’t track what’s going on anymore, so it’s hard to do anything with the model other than just deferring to it, and the model probably isn’t so good that you should just defer to it.
(Richard thinks these things are related, whereas I am more skeptical and think the mundane explanations are probably sufficient)