That’s great. So now I’ll construct a homunculus of an argument for a conclusion that neither of us agrees with, but which follows the structure, form, and abstract principles you have used. Notice how the principles feel less sound this time around.
“You shouldn’t hug your loved ones”, says the homunculus, “because it encourages close contact between people, which leads to the spread of infectious disease, including catastrophic pandemics. You should instead practice social distancing, which discourages close contact.
“Pandemics pose substantial, catastrophic risks, including loss of life. So encouraging or promoting close physical contact is something one should view as inherently negative. One should default to not hugging, and only hug when the benefits clearly outweigh the harms.
“In this case, the benefit is pretty much zero. I doubt you had your wealth substantially enhanced by this hug. The benefit is, what, some momentary comfort for you and your loved one? Increasing pandemic risk for your own comfort is antisocial behavior. If everyone did that, the effects would add up. It has negative externalities, and you shouldn’t do it.
“For someone who supposedly supports the idea of public health, you’re remarkably enthusiastic about supporting such disease-promoting behavior. Eyebrow-raising, to say the least.”
Obviously we don’t agree with the homunculus’ conclusion, so something is wrong with its argument. Maybe one or more of its general principles is flawed, or has unaddressed limitations?
Pandemics pose substantial, catastrophic risks, including loss of life. So encouraging or promoting close physical contact is something one should view as inherently negative.
I do not believe the net effect of human contact is negative.
Human-human contact does not promote aggregate risk in the case of catastrophic disease in the same way in our current environment.
If I believed that the net expected utility of human-human contact was negative, then yes I would be skeptical of human-human contact.
I do not believe that.
Increasing pandemic risk for your own comfort is antisocial behavior. If everyone did that, the effects would add up.
Yes, when the net utility from human-human physical contact is negative, I do consider increasing that risk as anti-social behavior and yes the risk do add up. In normal circumstances, the net utility is not negative in some special circumstances (i.e., during a pandemic) the net utility can be negative, in which case yes you should seek to avoid human-human physical contact. Not doing so for your own comfort (e.g., throwing parties or social events involving close physical contact during a pandemic) is antisocial behavior.
“It does promote pandemic risk”, says the homunculus. “You’re making physical contact look more appealing. It shifts the public sentiment. You’re going around promoting the greatness of physical contact, undermining your public-health messaging and contributing to a positive-feedback cycle of increasing physical contact and positive sentiment for the same. There’s no amount that doesn’t contribute.
“The expected harm of promoting a deadly pandemic even a little bit is pretty substantial. Anything that increases that risk should be difficult to justify. You would need to see a pretty large benefit to outweigh the risk, and as I said the benefit in this case is pretty much zero.”
I don’t know what case you think you are responding to, but it isn’t one I made.
When I think the net expected utility of (marginal) actions within a set of actions is negative (in the current world), then I am skeptical of actions in that set. I do not think the net expected utility of all human contact is negative. If you believe the expected outcome of AI is development is negative, I cannot imagine you do not also think the net expected utility of all additional contributions to AI is also negative. If you think that, then you should be skeptical towards any contribution.
That is exactly the attitude I take currently with carbon emissions, for example. I think the net expected utility of all (marginal) anthropogenic carbon emissions are negative. Therefore, I try to avoid actions that increase carbon emissions and default to viewing carbon emitting activities as negative. As I have repeated, that isn’t to say actions that increase carbon emissions are always bad, sometimes they may be justified by greater social utility and sometimes we may be willing to do them for our personal benefits—we are only human after all.
In certain periods, I have taken this attidude with outside-human contact as well. When I expect the net expected utility of all additional contact between people is negative (e.g., during a major pandemic), then I am skeptical towards participating in any activity that would promote or encourage that kind of contact.
Ah, I hadn’t appreciated how central that principle was in your thinking. Attempting to say it in my own words: “If there’s a set of actions such that taking all the actions would be worse than refraining from all the actions, because of the net harm, then we should presume that taking any particular action in that set is worse than not taking it, absent a really good reason.”
I’ll note this principle is quite sensitive to how you group actions into families.
Take the set of all carbon emissions starting now. If we suddenly refrained from all of them, we’d face famine pretty soon, which would be worse than global warming. So the principle doesn’t apply to marginal carbon emissions (in the sense of emissions we’re about to make but haven’t made yet).
If instead you take the set of future carbon emissions minus the ones we were “going to do anyways”, in the course of our daily lives, then you probably think it’s better to refrain from all these “extra” emissions, so by your principle one should by default refrain from any particular marginal emission. (“Marginal” here means going out of one’s way, doing something non-ordinary.)
(But what if one was planning on going into a career in oil drilling? Is that an “extra” action?)
The homunculus says, “We already had SARS, MERS, and COVID. We’ll surely get a fourth coronavirus epidemic soon, and the next one could be deadly enough to outweigh the benefit of the preceding decade of human contact. Better to refrain from all but medically necessary human contact until we develop broad-spectrum vaccines against all coronaviruses. Therefore any particular hug should be avoided by default.”
I’d be happy to give up the entire deep learning revolution to avoid existential risk from AI, so your principle would have me refrain from praising Google Translate in 2017.
If there’s a set of actions such that taking all the actions would be worse than refraining from all the actions, because of the net harm, then we should presume that taking any particular action in that set is worse than not taking it, absent a really good reason
I would drop the ‘absent a really good reason.’ It isn’t necessarily a strong recommendation—it is approximately a default assumption. I previously formalized it as: If I know the net expected utility of set is negative (i.e. where ). Without additional information about expected utility of any given element , I can say .
Absent further information and evaluation of an action, the default expectation should be it is negative (in proportion to how negative you view the aggregate and a weighting of the factor as part of the whole).
Take the set of all carbon emissions starting now. If we suddenly refrained from all of them, we’d face famine pretty soon, which would be worse than global warming. So the principle doesn’t apply to marginal carbon emissions (in the sense of emissions we’re about to make but haven’t made yet).
I think you are misunderstanding what I the expected utility function would work out to, the counter-factual isn’t “stop all carbon emissions,” it is a probability distribution of what we expect carbon emissions to be. It would be:
.
Where j represents each possible amount of future emissions, represents the probability of j emissions, represents the expected utility of the consequences of j level of emissions.
We can also say:
(that is, the expected utility resulting from j level emissions is the sum-product of the probability of each outcome at j level of emissions times the utility of each of those outcomes).
Because the probability distribution of carbon emissions is well beyond a level that is safe and the bottom-end of the distribution hardly touches on a level where we would expect more harm from famine, the expected utility of future carbon emissions works out to be negative.
Our actions/inactions nudge the probability curve slightly one way or another, they do not flip it on and off. IF we got to the point where the probability distribution was such that it was well below catastrophic risk, then the EU would become positive.
We already had SARS, MERS, and COVID. We’ll surely get a fourth coronavirus epidemic soon, and the next one could be deadly enough to outweigh the benefit of the preceding decade of human contact.
If you believe that risk is so substantial and human contact is part of that risk, then yes I would say you should try to minimize human contact. I do not believe the risk is that substantial, I believe that in net marginal human-human contact is a social neutral-to-positive. There could be something deadly enough to outweigh the benefit, but I would estimate the risk at a substantially lower-level.
Thanks for clarifying. Here’s my main criticism of your principle.
Without additional information about expected utility of any given element , I can say [its expected utility is negative.]
I basically agree with this. If all you know about an action is that it belongs to a set of actions with presumed independent effects that add up to negative value, then the expected value of that action is negative.
However, you usually have additional information about the action, so the argument doesn’t apply. So your principle is very weak.
I think this is the answer to your question at the top of this thread: Why do people concerned about catastrophic risks from AI sometimes do things in the direction of supporting AI? Because they know the details of the particular actions they’re taking, and based on those details they judge those actions to have positive expected value. (Ideally! Many people are nincompoops who do things for bad reasons. But I think the Unslop contest was fine.) That’s why I said “you have to disagree on the numbers”.
I think you tend to use a stronger version of the principle which is something like, “Come on, there’s no way you have enough information to know that it’s positive expected value, something funny is going on.” To which I would say, “No really, let’s talk about the numbers, it checks out.”
I’m running out of energy to spend on this thread, though.
Also, I should say, that when you plug in the numbers you get big negatives. Let’s say i estimate at the current probilibility distribution of AI development has a net expected utility of say −1,000,000,000,000 ( say, my p(doom) is ~1/8, i think 8 billion humans will die and i value each human at 1,000).
If i think my actions (including positive feedback!) are contributing 1⁄100,000,000th of AI development, then my expectated utility from the AI development I contribute should be ~ −10,000 (about the same as the value I would place on 10 human lives). So unless I think i am saving 10 lives (or equivelent) in contributing 1 hundredth of a millionth of AI development, I should find my actions negative.
Current total investment in AI is under 2 trillion. Being conservative (a dollar now means more than a dollar later!) you should value contributing $2,000 to AI development as being roughly as negative as killing a stranger (under a simplified framework, that just looks at dollars).
Now we’re getting somewhere. You’re assuming that providing one billionth of this year’s investment in AI means decreasing the expected utility from human extinction by one billionth. But that’s not right. If you invest in AI, that means the AI companies will choose to raise less money from other investors. The net effect is that AI companies will raise a little more money than if you hadn’t invested (but less than the amount you invest), get a better interest rate, and spend less effort on fundraising this year. And this will have some effect on the probability of an extinction-level AI catastrophe this generation, but the effect isn’t linear in the amount you invest.
In a parallel thread, your theory of impact was that divestment has an effect on regulation. That’s also nonlinear.
Investing in AI is something I’m likely to do. (Arguably >10% of the S&P 500 by market cap use marginal investment to try to build frontier models.) Here’s how I justify it: The amount of stock I can buy won’t move the stock price perceptibly. If I estimated the price movement it would be tiny. That would translate into AI companies expending a tiny bit less effort on selling stock this year to fund datacenters, which accelerates timelines a little bit, which increases existential risk a very tiny amount. On the other hand, if I invest my savings in a broad-market index fund that includes AI, that’ll increase my savings and I’ll have more time to spend doing AI safety research, calling my Congressperson, etc., which are more directly impactful. Plus, I have lots of personal uses for more money.
I’m basically satisfied with where this thread ended up. (I just wanted to convey that it’s in general reasonable to do things that seem individually worth it, even if they seem to be promoting AI in some way.) I want to flag that I won’t have the energy to find a crux in our disagreement over the Unslop contest or whether I should buy index funds. Though if you have a good argument about why not to buy index funds I haven’t thought of, I’ll be interested to hear that.
You seem to be confusing some things, probably because I was not as precise with definitions of terms as I should have been.
In a parallel thread, your theory of impact was that divestment has an effect on regulation. That’s also nonlinear.
Divestment was an example of a common behavior activists engage in (which I would broadly speaking endorse). The impact from divestment on a industry is different from the negative impact of funding, advertising them, and paying an industry. They not direct inversions of each other, divesting is a very different activity.
Now we’re getting somewhere. You’re assuming that providing one billionth of this year’s investment in AI means decreasing the expected utility from human extinction by one billionth.
I said “contributing to AI development” here, not investing. What is meant by “investing” is more nuanced (e.g. buying stocks second-hand on the open market is not actually ‘investing’ in the economic sense).
Here, I am including something as simple as buying AI products (a substantial portion of money sent will contribute to AI development).
But that’s not right. If you invest in AI, that means the AI companies will choose to raise less money from other investors.
This is (generally) the exact opposite of how investment works. Me investing a dollar, encourages Steve to invest a dollar which provides collateral to borrow another 2 dollars. Me not investing a dollar encourages Steve not to invest a dollar. Me spending a dollar a product boosts that company’s profits, encouraging others to invest more dollars and moving outward the company’s expected demand, encouraging them to invest internally in expanding that product.
Investments follow trends, both of other investments and of consumer behavior, moving investment towards one thing moves more investment towards it.
which increases existential risk a very tiny amount.
A very tiny increase in an extreme risk is very negative. And remember, these are aggregate risks that cover some probability distribution of outcomes, not just flat +/- x risk. It is accurate, as such, to estimate the EU as a portion of the whole, not just the absolute movement in risk.
Investing in AI is something I’m likely to do. (Arguably >10% of the S&P 500 by market cap use marginal investment to try to build frontier models.)
Assuming you aren’t a private equity firm or buying directly in IPOs, your personal finances have minimal direct impact on the economic investment which is the investment one cares about. Divesting from exposed firms is a signaling behavior and really most useful when the divestment is coming from funds that are actually providing investment, not just trading on the open market (divestment movements generally focus on capital funds and large endowments).
However, you usually have additional information about the action, so the argument doesn’t apply. So your principle is very weak.
Yes, it is a weak default case, but it generalizes. If rational, the people who find the aggregate risk the most serious should also have the highest threshold for engaging in any behavior increasing that risk.
This largely holds when it comes to the public actions of e.g. climate activists and even what i called “pop” AI criticis, but the opposite is true for some of the people who proclaim to be the most critical of AI.
To which I would say, “No really, let’s talk about the numbers, it checks out.”
Then give the numbers. The only concrete numbers you gave me was gwern thinking there was negative expected utility to the production of new fiction, which runs directly counter to your claimed rationale.
I disagree with it as well, but it is the estimated utility in the post from gwern you directed me to.
Hugging a loved one, holding the door for a stranger, giving a dollar to charity, etc.
That’s great. So now I’ll construct a homunculus of an argument for a conclusion that neither of us agrees with, but which follows the structure, form, and abstract principles you have used. Notice how the principles feel less sound this time around.
“You shouldn’t hug your loved ones”, says the homunculus, “because it encourages close contact between people, which leads to the spread of infectious disease, including catastrophic pandemics. You should instead practice social distancing, which discourages close contact.
“Pandemics pose substantial, catastrophic risks, including loss of life. So encouraging or promoting close physical contact is something one should view as inherently negative. One should default to not hugging, and only hug when the benefits clearly outweigh the harms.
“In this case, the benefit is pretty much zero. I doubt you had your wealth substantially enhanced by this hug. The benefit is, what, some momentary comfort for you and your loved one? Increasing pandemic risk for your own comfort is antisocial behavior. If everyone did that, the effects would add up. It has negative externalities, and you shouldn’t do it.
“For someone who supposedly supports the idea of public health, you’re remarkably enthusiastic about supporting such disease-promoting behavior. Eyebrow-raising, to say the least.”
Obviously we don’t agree with the homunculus’ conclusion, so something is wrong with its argument. Maybe one or more of its general principles is flawed, or has unaddressed limitations?
I do not believe the net effect of human contact is negative.
Human-human contact does not promote aggregate risk in the case of catastrophic disease in the same way in our current environment.
If I believed that the net expected utility of human-human contact was negative, then yes I would be skeptical of human-human contact.
I do not believe that.
Yes, when the net utility from human-human physical contact is negative, I do consider increasing that risk as anti-social behavior and yes the risk do add up. In normal circumstances, the net utility is not negative in some special circumstances (i.e., during a pandemic) the net utility can be negative, in which case yes you should seek to avoid human-human physical contact. Not doing so for your own comfort (e.g., throwing parties or social events involving close physical contact during a pandemic) is antisocial behavior.
“It does promote pandemic risk”, says the homunculus. “You’re making physical contact look more appealing. It shifts the public sentiment. You’re going around promoting the greatness of physical contact, undermining your public-health messaging and contributing to a positive-feedback cycle of increasing physical contact and positive sentiment for the same. There’s no amount that doesn’t contribute.
“The expected harm of promoting a deadly pandemic even a little bit is pretty substantial. Anything that increases that risk should be difficult to justify. You would need to see a pretty large benefit to outweigh the risk, and as I said the benefit in this case is pretty much zero.”
I don’t know what case you think you are responding to, but it isn’t one I made.
When I think the net expected utility of (marginal) actions within a set of actions is negative (in the current world), then I am skeptical of actions in that set. I do not think the net expected utility of all human contact is negative. If you believe the expected outcome of AI is development is negative, I cannot imagine you do not also think the net expected utility of all additional contributions to AI is also negative. If you think that, then you should be skeptical towards any contribution.
That is exactly the attitude I take currently with carbon emissions, for example. I think the net expected utility of all (marginal) anthropogenic carbon emissions are negative. Therefore, I try to avoid actions that increase carbon emissions and default to viewing carbon emitting activities as negative. As I have repeated, that isn’t to say actions that increase carbon emissions are always bad, sometimes they may be justified by greater social utility and sometimes we may be willing to do them for our personal benefits—we are only human after all.
In certain periods, I have taken this attidude with outside-human contact as well. When I expect the net expected utility of all additional contact between people is negative (e.g., during a major pandemic), then I am skeptical towards participating in any activity that would promote or encourage that kind of contact.
Ah, I hadn’t appreciated how central that principle was in your thinking. Attempting to say it in my own words: “If there’s a set of actions such that taking all the actions would be worse than refraining from all the actions, because of the net harm, then we should presume that taking any particular action in that set is worse than not taking it, absent a really good reason.”
I’ll note this principle is quite sensitive to how you group actions into families.
Take the set of all carbon emissions starting now. If we suddenly refrained from all of them, we’d face famine pretty soon, which would be worse than global warming. So the principle doesn’t apply to marginal carbon emissions (in the sense of emissions we’re about to make but haven’t made yet).
If instead you take the set of future carbon emissions minus the ones we were “going to do anyways”, in the course of our daily lives, then you probably think it’s better to refrain from all these “extra” emissions, so by your principle one should by default refrain from any particular marginal emission. (“Marginal” here means going out of one’s way, doing something non-ordinary.)
(But what if one was planning on going into a career in oil drilling? Is that an “extra” action?)
The homunculus says, “We already had SARS, MERS, and COVID. We’ll surely get a fourth coronavirus epidemic soon, and the next one could be deadly enough to outweigh the benefit of the preceding decade of human contact. Better to refrain from all but medically necessary human contact until we develop broad-spectrum vaccines against all coronaviruses. Therefore any particular hug should be avoided by default.”
I’d be happy to give up the entire deep learning revolution to avoid existential risk from AI, so your principle would have me refrain from praising Google Translate in 2017.
I would drop the ‘absent a really good reason.’ It isn’t necessarily a strong recommendation—it is approximately a default assumption. I previously formalized it as: If I know the net expected utility of set is negative (i.e. where ). Without additional information about expected utility of any given element , I can say .
Absent further information and evaluation of an action, the default expectation should be it is negative (in proportion to how negative you view the aggregate and a weighting of the factor as part of the whole).
I think you are misunderstanding what I the expected utility function would work out to, the counter-factual isn’t “stop all carbon emissions,” it is a probability distribution of what we expect carbon emissions to be. It would be:
Where j represents each possible amount of future emissions, represents the probability of j emissions, represents the expected utility of the consequences of j level of emissions.
We can also say:
Because the probability distribution of carbon emissions is well beyond a level that is safe and the bottom-end of the distribution hardly touches on a level where we would expect more harm from famine, the expected utility of future carbon emissions works out to be negative.
Our actions/inactions nudge the probability curve slightly one way or another, they do not flip it on and off. IF we got to the point where the probability distribution was such that it was well below catastrophic risk, then the EU would become positive.
If you believe that risk is so substantial and human contact is part of that risk, then yes I would say you should try to minimize human contact. I do not believe the risk is that substantial, I believe that in net marginal human-human contact is a social neutral-to-positive. There could be something deadly enough to outweigh the benefit, but I would estimate the risk at a substantially lower-level.
Thanks for clarifying. Here’s my main criticism of your principle.
I basically agree with this. If all you know about an action is that it belongs to a set of actions with presumed independent effects that add up to negative value, then the expected value of that action is negative.
However, you usually have additional information about the action, so the argument doesn’t apply. So your principle is very weak.
I think this is the answer to your question at the top of this thread: Why do people concerned about catastrophic risks from AI sometimes do things in the direction of supporting AI? Because they know the details of the particular actions they’re taking, and based on those details they judge those actions to have positive expected value. (Ideally! Many people are nincompoops who do things for bad reasons. But I think the Unslop contest was fine.) That’s why I said “you have to disagree on the numbers”.
I think you tend to use a stronger version of the principle which is something like, “Come on, there’s no way you have enough information to know that it’s positive expected value, something funny is going on.” To which I would say, “No really, let’s talk about the numbers, it checks out.”
I’m running out of energy to spend on this thread, though.
Also, I should say, that when you plug in the numbers you get big negatives. Let’s say i estimate at the current probilibility distribution of AI development has a net expected utility of say −1,000,000,000,000 ( say, my p(doom) is ~1/8, i think 8 billion humans will die and i value each human at 1,000).
If i think my actions (including positive feedback!) are contributing 1⁄100,000,000th of AI development, then my expectated utility from the AI development I contribute should be ~ −10,000 (about the same as the value I would place on 10 human lives). So unless I think i am saving 10 lives (or equivelent) in contributing 1 hundredth of a millionth of AI development, I should find my actions negative.
Current total investment in AI is under 2 trillion. Being conservative (a dollar now means more than a dollar later!) you should value contributing $2,000 to AI development as being roughly as negative as killing a stranger (under a simplified framework, that just looks at dollars).
Now we’re getting somewhere. You’re assuming that providing one billionth of this year’s investment in AI means decreasing the expected utility from human extinction by one billionth. But that’s not right. If you invest in AI, that means the AI companies will choose to raise less money from other investors. The net effect is that AI companies will raise a little more money than if you hadn’t invested (but less than the amount you invest), get a better interest rate, and spend less effort on fundraising this year. And this will have some effect on the probability of an extinction-level AI catastrophe this generation, but the effect isn’t linear in the amount you invest.
In a parallel thread, your theory of impact was that divestment has an effect on regulation. That’s also nonlinear.
Investing in AI is something I’m likely to do. (Arguably >10% of the S&P 500 by market cap use marginal investment to try to build frontier models.) Here’s how I justify it: The amount of stock I can buy won’t move the stock price perceptibly. If I estimated the price movement it would be tiny. That would translate into AI companies expending a tiny bit less effort on selling stock this year to fund datacenters, which accelerates timelines a little bit, which increases existential risk a very tiny amount. On the other hand, if I invest my savings in a broad-market index fund that includes AI, that’ll increase my savings and I’ll have more time to spend doing AI safety research, calling my Congressperson, etc., which are more directly impactful. Plus, I have lots of personal uses for more money.
I’m basically satisfied with where this thread ended up. (I just wanted to convey that it’s in general reasonable to do things that seem individually worth it, even if they seem to be promoting AI in some way.) I want to flag that I won’t have the energy to find a crux in our disagreement over the Unslop contest or whether I should buy index funds. Though if you have a good argument about why not to buy index funds I haven’t thought of, I’ll be interested to hear that.
You seem to be confusing some things, probably because I was not as precise with definitions of terms as I should have been.
Divestment was an example of a common behavior activists engage in (which I would broadly speaking endorse). The impact from divestment on a industry is different from the negative impact of funding, advertising them, and paying an industry. They not direct inversions of each other, divesting is a very different activity.
I said “contributing to AI development” here, not investing. What is meant by “investing” is more nuanced (e.g. buying stocks second-hand on the open market is not actually ‘investing’ in the economic sense).
Here, I am including something as simple as buying AI products (a substantial portion of money sent will contribute to AI development).
This is (generally) the exact opposite of how investment works. Me investing a dollar, encourages Steve to invest a dollar which provides collateral to borrow another 2 dollars. Me not investing a dollar encourages Steve not to invest a dollar. Me spending a dollar a product boosts that company’s profits, encouraging others to invest more dollars and moving outward the company’s expected demand, encouraging them to invest internally in expanding that product.
Investments follow trends, both of other investments and of consumer behavior, moving investment towards one thing moves more investment towards it.
A very tiny increase in an extreme risk is very negative. And remember, these are aggregate risks that cover some probability distribution of outcomes, not just flat +/- x risk. It is accurate, as such, to estimate the EU as a portion of the whole, not just the absolute movement in risk.
Assuming you aren’t a private equity firm or buying directly in IPOs, your personal finances have minimal direct impact on the economic investment which is the investment one cares about. Divesting from exposed firms is a signaling behavior and really most useful when the divestment is coming from funds that are actually providing investment, not just trading on the open market (divestment movements generally focus on capital funds and large endowments).
Yes, it is a weak default case, but it generalizes. If rational, the people who find the aggregate risk the most serious should also have the highest threshold for engaging in any behavior increasing that risk.
This largely holds when it comes to the public actions of e.g. climate activists and even what i called “pop” AI criticis, but the opposite is true for some of the people who proclaim to be the most critical of AI.
Then give the numbers. The only concrete numbers you gave me was gwern thinking there was negative expected utility to the production of new fiction, which runs directly counter to your claimed rationale.