The difference is that you’re applying a deontological principle to existential risk in a way that imho doesn’t make sense. I was hoping you’d accept my invitation to try out the principle in a domain where your intuitions are reversed — where the benefit feels real and central, where the harm feels theoretical, trivial, unserious. That would be a better test of your principles. It is an uncomfortable thing to do, and you don’t have to. You’ve been a good sport. Thanks for the conversation!
I was hoping you’d accept my invitation to try out the principle in a domain where your intuitions are reversed — where the benefit feels real and central, where the harm feels theoretical, trivial, unserious.
Must have missed this, but i do this every day. E.g., i think compassion is an average good, and default to being less skeptical of the utility of any given act of compassion.
The difference is that you’re applying a deontological principle to existential risk in a way that imho doesn’t make sens
I think there is still some misunderstanding of the other on one or both our parts.. My claims were entirely based in utilitarian heuristics, I am not sure what deontological principle you are referring to.
Let me put it another way, just on the case of promotion in general (not even including the additional negative utility of weakening the message of activists)
If in the aggregate of X has substantial enough risk to result in a negative expected utility, then the average case of X has a negative expected utility.
If the average case has a negative utility, then any given case should be viewed skeptically (and the greater the average negative utility the more critical one should be).
If you view the aggregate of the current tragetory of AI as having enough harm/risk to be of a negative expexted utility (e.g., if you believe there is substantial risk of current trends leading to develops that cause the extinction of all humans), then under that view the average contribution to AI has a negative expected utility.
If the average contribution to AI has a negative expected utility, then any given contribution should be viewed skeptically (and the greater the risk, the more one should default to opposition)
Much “pop” criticism of AI views it as a net negative not because of existential risk but rather job loss and amorphous negative views of AI products, among other factors.
These “pop” critics default to viewing AI developments critically (but will generally view AI as at least acceptable when it has what they view as an outsized positive effect).
These “pop” critics, thus, appear consistent with (4).
Some of the people proclaiming to view the aggregate of the current tragetory of AI as the most negative (i.e., they place the most credence in estimations of a very high extinction risk) seem among the most enthusiastic hen it comes to promoting AI use cases and development.
To be clear, I don’t think there’s anything wrong with having deontological principles. I try to follow some deontological principles myself! Here are some things you’ve said that make me think you’re not using entirely utilitarian reasoning (bolding mine):
not withstanding benefits that are greater than the risks
something one should view as inherently negative and only acceptable when the benefits clearly outweigh the harms
One should be somewhat risk adverse about the calculations
There is no amount that doesn’t contribute
one should default to avoiding
There is a pretty massive difference between actively endorsing and encouraging the use of AI and inadvertently wearing a shirt with a logo because you’re lazy.
so long as it has no negative externalities
any given contribution should be viewed skeptically
Utilitarianism would ask: Do the benefits outweigh the harms? But a deontology also cares about the bolded terms: Is the act inherently negative? Should it be avoided by default? Is there an unacceptable harm that cannot be balanced by any amount of benefit? Is the act intentional or inadvertent? Is there an extenuating motivation?
In fact my first comment in this thread recommended that you do the utilitarian thing and estimate the benefits and harms. Instead you responded with arguments that, to my eye, would let you stick to your conclusion even if the benefits turned out to technically outweigh the harms.
What’s a concrete example of an ordinary, everyday act of compassion you’d wholeheartedly endorse?
Utilitarianism would ask: Do the benefits outweigh the harms?
Yes, and I am saying you can (under the described beliefs) use carbon emissions or promoting AI as heuristics for the harms outweighing the benefits. If the only information I know about an action is that it increases carbon emissions or it promotes AI then if I believe those things are—in aggregate—bad then my expectation for any given action should be that it is bad.
If I know the net expected utility of set is negative (i.e. where ). Without additional information about expected utility of any given element , I can say .
We never have perfect knowledge of utility expectations.
In fact my first comment in this thread recommended that you do the utilitarian thing and estimate the benefits and harms.
The actual utility calculation is unclear and depends on one’s personal values which I wasn’t working on. But going by what you gave me for gwern’s expected utility, I should be able to say:
That’s great. So now I’ll construct a homunculus of an argument for a conclusion that neither of us agrees with, but which follows the structure, form, and abstract principles you have used. Notice how the principles feel less sound this time around.
“You shouldn’t hug your loved ones”, says the homunculus, “because it encourages close contact between people, which leads to the spread of infectious disease, including catastrophic pandemics. You should instead practice social distancing, which discourages close contact.
“Pandemics pose substantial, catastrophic risks, including loss of life. So encouraging or promoting close physical contact is something one should view as inherently negative. One should default to not hugging, and only hug when the benefits clearly outweigh the harms.
“In this case, the benefit is pretty much zero. I doubt you had your wealth substantially enhanced by this hug. The benefit is, what, some momentary comfort for you and your loved one? Increasing pandemic risk for your own comfort is antisocial behavior. If everyone did that, the effects would add up. It has negative externalities, and you shouldn’t do it.
“For someone who supposedly supports the idea of public health, you’re remarkably enthusiastic about supporting such disease-promoting behavior. Eyebrow-raising, to say the least.”
Obviously we don’t agree with the homunculus’ conclusion, so something is wrong with its argument. Maybe one or more of its general principles is flawed, or has unaddressed limitations?
Pandemics pose substantial, catastrophic risks, including loss of life. So encouraging or promoting close physical contact is something one should view as inherently negative.
I do not believe the net effect of human contact is negative.
Human-human contact does not promote aggregate risk in the case of catastrophic disease in the same way in our current environment.
If I believed that the net expected utility of human-human contact was negative, then yes I would be skeptical of human-human contact.
I do not believe that.
Increasing pandemic risk for your own comfort is antisocial behavior. If everyone did that, the effects would add up.
Yes, when the net utility from human-human physical contact is negative, I do consider increasing that risk as anti-social behavior and yes the risk do add up. In normal circumstances, the net utility is not negative in some special circumstances (i.e., during a pandemic) the net utility can be negative, in which case yes you should seek to avoid human-human physical contact. Not doing so for your own comfort (e.g., throwing parties or social events involving close physical contact during a pandemic) is antisocial behavior.
“It does promote pandemic risk”, says the homunculus. “You’re making physical contact look more appealing. It shifts the public sentiment. You’re going around promoting the greatness of physical contact, undermining your public-health messaging and contributing to a positive-feedback cycle of increasing physical contact and positive sentiment for the same. There’s no amount that doesn’t contribute.
“The expected harm of promoting a deadly pandemic even a little bit is pretty substantial. Anything that increases that risk should be difficult to justify. You would need to see a pretty large benefit to outweigh the risk, and as I said the benefit in this case is pretty much zero.”
I don’t know what case you think you are responding to, but it isn’t one I made.
When I think the net expected utility of (marginal) actions within a set of actions is negative (in the current world), then I am skeptical of actions in that set. I do not think the net expected utility of all human contact is negative. If you believe the expected outcome of AI is development is negative, I cannot imagine you do not also think the net expected utility of all additional contributions to AI is also negative. If you think that, then you should be skeptical towards any contribution.
That is exactly the attitude I take currently with carbon emissions, for example. I think the net expected utility of all (marginal) anthropogenic carbon emissions are negative. Therefore, I try to avoid actions that increase carbon emissions and default to viewing carbon emitting activities as negative. As I have repeated, that isn’t to say actions that increase carbon emissions are always bad, sometimes they may be justified by greater social utility and sometimes we may be willing to do them for our personal benefits—we are only human after all.
In certain periods, I have taken this attidude with outside-human contact as well. When I expect the net expected utility of all additional contact between people is negative (e.g., during a major pandemic), then I am skeptical towards participating in any activity that would promote or encourage that kind of contact.
Ah, I hadn’t appreciated how central that principle was in your thinking. Attempting to say it in my own words: “If there’s a set of actions such that taking all the actions would be worse than refraining from all the actions, because of the net harm, then we should presume that taking any particular action in that set is worse than not taking it, absent a really good reason.”
I’ll note this principle is quite sensitive to how you group actions into families.
Take the set of all carbon emissions starting now. If we suddenly refrained from all of them, we’d face famine pretty soon, which would be worse than global warming. So the principle doesn’t apply to marginal carbon emissions (in the sense of emissions we’re about to make but haven’t made yet).
If instead you take the set of future carbon emissions minus the ones we were “going to do anyways”, in the course of our daily lives, then you probably think it’s better to refrain from all these “extra” emissions, so by your principle one should by default refrain from any particular marginal emission. (“Marginal” here means going out of one’s way, doing something non-ordinary.)
(But what if one was planning on going into a career in oil drilling? Is that an “extra” action?)
The homunculus says, “We already had SARS, MERS, and COVID. We’ll surely get a fourth coronavirus epidemic soon, and the next one could be deadly enough to outweigh the benefit of the preceding decade of human contact. Better to refrain from all but medically necessary human contact until we develop broad-spectrum vaccines against all coronaviruses. Therefore any particular hug should be avoided by default.”
I’d be happy to give up the entire deep learning revolution to avoid existential risk from AI, so your principle would have me refrain from praising Google Translate in 2017.
If there’s a set of actions such that taking all the actions would be worse than refraining from all the actions, because of the net harm, then we should presume that taking any particular action in that set is worse than not taking it, absent a really good reason
I would drop the ‘absent a really good reason.’ It isn’t necessarily a strong recommendation—it is approximately a default assumption. I previously formalized it as: If I know the net expected utility of set is negative (i.e. where ). Without additional information about expected utility of any given element , I can say .
Absent further information and evaluation of an action, the default expectation should be it is negative (in proportion to how negative you view the aggregate and a weighting of the factor as part of the whole).
Take the set of all carbon emissions starting now. If we suddenly refrained from all of them, we’d face famine pretty soon, which would be worse than global warming. So the principle doesn’t apply to marginal carbon emissions (in the sense of emissions we’re about to make but haven’t made yet).
I think you are misunderstanding what I the expected utility function would work out to, the counter-factual isn’t “stop all carbon emissions,” it is a probability distribution of what we expect carbon emissions to be. It would be:
.
Where j represents each possible amount of future emissions, represents the probability of j emissions, represents the expected utility of the consequences of j level of emissions.
We can also say:
(that is, the expected utility resulting from j level emissions is the sum-product of the probability of each outcome at j level of emissions times the utility of each of those outcomes).
Because the probability distribution of carbon emissions is well beyond a level that is safe and the bottom-end of the distribution hardly touches on a level where we would expect more harm from famine, the expected utility of future carbon emissions works out to be negative.
Our actions/inactions nudge the probability curve slightly one way or another, they do not flip it on and off. IF we got to the point where the probability distribution was such that it was well below catastrophic risk, then the EU would become positive.
We already had SARS, MERS, and COVID. We’ll surely get a fourth coronavirus epidemic soon, and the next one could be deadly enough to outweigh the benefit of the preceding decade of human contact.
If you believe that risk is so substantial and human contact is part of that risk, then yes I would say you should try to minimize human contact. I do not believe the risk is that substantial, I believe that in net marginal human-human contact is a social neutral-to-positive. There could be something deadly enough to outweigh the benefit, but I would estimate the risk at a substantially lower-level.
Thanks for clarifying. Here’s my main criticism of your principle.
Without additional information about expected utility of any given element , I can say [its expected utility is negative.]
I basically agree with this. If all you know about an action is that it belongs to a set of actions with presumed independent effects that add up to negative value, then the expected value of that action is negative.
However, you usually have additional information about the action, so the argument doesn’t apply. So your principle is very weak.
I think this is the answer to your question at the top of this thread: Why do people concerned about catastrophic risks from AI sometimes do things in the direction of supporting AI? Because they know the details of the particular actions they’re taking, and based on those details they judge those actions to have positive expected value. (Ideally! Many people are nincompoops who do things for bad reasons. But I think the Unslop contest was fine.) That’s why I said “you have to disagree on the numbers”.
I think you tend to use a stronger version of the principle which is something like, “Come on, there’s no way you have enough information to know that it’s positive expected value, something funny is going on.” To which I would say, “No really, let’s talk about the numbers, it checks out.”
I’m running out of energy to spend on this thread, though.
Also, I should say, that when you plug in the numbers you get big negatives. Let’s say i estimate at the current probilibility distribution of AI development has a net expected utility of say −1,000,000,000,000 ( say, my p(doom) is ~1/8, i think 8 billion humans will die and i value each human at 1,000).
If i think my actions (including positive feedback!) are contributing 1⁄100,000,000th of AI development, then my expectated utility from the AI development I contribute should be ~ −10,000 (about the same as the value I would place on 10 human lives). So unless I think i am saving 10 lives (or equivelent) in contributing 1 hundredth of a millionth of AI development, I should find my actions negative.
Current total investment in AI is under 2 trillion. Being conservative (a dollar now means more than a dollar later!) you should value contributing $2,000 to AI development as being roughly as negative as killing a stranger (under a simplified framework, that just looks at dollars).
Now we’re getting somewhere. You’re assuming that providing one billionth of this year’s investment in AI means decreasing the expected utility from human extinction by one billionth. But that’s not right. If you invest in AI, that means the AI companies will choose to raise less money from other investors. The net effect is that AI companies will raise a little more money than if you hadn’t invested (but less than the amount you invest), get a better interest rate, and spend less effort on fundraising this year. And this will have some effect on the probability of an extinction-level AI catastrophe this generation, but the effect isn’t linear in the amount you invest.
In a parallel thread, your theory of impact was that divestment has an effect on regulation. That’s also nonlinear.
Investing in AI is something I’m likely to do. (Arguably >10% of the S&P 500 by market cap use marginal investment to try to build frontier models.) Here’s how I justify it: The amount of stock I can buy won’t move the stock price perceptibly. If I estimated the price movement it would be tiny. That would translate into AI companies expending a tiny bit less effort on selling stock this year to fund datacenters, which accelerates timelines a little bit, which increases existential risk a very tiny amount. On the other hand, if I invest my savings in a broad-market index fund that includes AI, that’ll increase my savings and I’ll have more time to spend doing AI safety research, calling my Congressperson, etc., which are more directly impactful. Plus, I have lots of personal uses for more money.
I’m basically satisfied with where this thread ended up. (I just wanted to convey that it’s in general reasonable to do things that seem individually worth it, even if they seem to be promoting AI in some way.) I want to flag that I won’t have the energy to find a crux in our disagreement over the Unslop contest or whether I should buy index funds. Though if you have a good argument about why not to buy index funds I haven’t thought of, I’ll be interested to hear that.
You seem to be confusing some things, probably because I was not as precise with definitions of terms as I should have been.
In a parallel thread, your theory of impact was that divestment has an effect on regulation. That’s also nonlinear.
Divestment was an example of a common behavior activists engage in (which I would broadly speaking endorse). The impact from divestment on a industry is different from the negative impact of funding, advertising them, and paying an industry. They not direct inversions of each other, divesting is a very different activity.
Now we’re getting somewhere. You’re assuming that providing one billionth of this year’s investment in AI means decreasing the expected utility from human extinction by one billionth.
I said “contributing to AI development” here, not investing. What is meant by “investing” is more nuanced (e.g. buying stocks second-hand on the open market is not actually ‘investing’ in the economic sense).
Here, I am including something as simple as buying AI products (a substantial portion of money sent will contribute to AI development).
But that’s not right. If you invest in AI, that means the AI companies will choose to raise less money from other investors.
This is (generally) the exact opposite of how investment works. Me investing a dollar, encourages Steve to invest a dollar which provides collateral to borrow another 2 dollars. Me not investing a dollar encourages Steve not to invest a dollar. Me spending a dollar a product boosts that company’s profits, encouraging others to invest more dollars and moving outward the company’s expected demand, encouraging them to invest internally in expanding that product.
Investments follow trends, both of other investments and of consumer behavior, moving investment towards one thing moves more investment towards it.
which increases existential risk a very tiny amount.
A very tiny increase in an extreme risk is very negative. And remember, these are aggregate risks that cover some probability distribution of outcomes, not just flat +/- x risk. It is accurate, as such, to estimate the EU as a portion of the whole, not just the absolute movement in risk.
Investing in AI is something I’m likely to do. (Arguably >10% of the S&P 500 by market cap use marginal investment to try to build frontier models.)
Assuming you aren’t a private equity firm or buying directly in IPOs, your personal finances have minimal direct impact on the economic investment which is the investment one cares about. Divesting from exposed firms is a signaling behavior and really most useful when the divestment is coming from funds that are actually providing investment, not just trading on the open market (divestment movements generally focus on capital funds and large endowments).
However, you usually have additional information about the action, so the argument doesn’t apply. So your principle is very weak.
Yes, it is a weak default case, but it generalizes. If rational, the people who find the aggregate risk the most serious should also have the highest threshold for engaging in any behavior increasing that risk.
This largely holds when it comes to the public actions of e.g. climate activists and even what i called “pop” AI criticis, but the opposite is true for some of the people who proclaim to be the most critical of AI.
To which I would say, “No really, let’s talk about the numbers, it checks out.”
Then give the numbers. The only concrete numbers you gave me was gwern thinking there was negative expected utility to the production of new fiction, which runs directly counter to your claimed rationale.
The difference is that you’re applying a deontological principle to existential risk in a way that imho doesn’t make sense. I was hoping you’d accept my invitation to try out the principle in a domain where your intuitions are reversed — where the benefit feels real and central, where the harm feels theoretical, trivial, unserious. That would be a better test of your principles. It is an uncomfortable thing to do, and you don’t have to. You’ve been a good sport. Thanks for the conversation!
Must have missed this, but i do this every day. E.g., i think compassion is an average good, and default to being less skeptical of the utility of any given act of compassion.
I think there is still some misunderstanding of the other on one or both our parts.. My claims were entirely based in utilitarian heuristics, I am not sure what deontological principle you are referring to.
Let me put it another way, just on the case of promotion in general (not even including the additional negative utility of weakening the message of activists)
If in the aggregate of X has substantial enough risk to result in a negative expected utility, then the average case of X has a negative expected utility.
If the average case has a negative utility, then any given case should be viewed skeptically (and the greater the average negative utility the more critical one should be).
If you view the aggregate of the current tragetory of AI as having enough harm/risk to be of a negative expexted utility (e.g., if you believe there is substantial risk of current trends leading to develops that cause the extinction of all humans), then under that view the average contribution to AI has a negative expected utility.
If the average contribution to AI has a negative expected utility, then any given contribution should be viewed skeptically (and the greater the risk, the more one should default to opposition)
Much “pop” criticism of AI views it as a net negative not because of existential risk but rather job loss and amorphous negative views of AI products, among other factors.
These “pop” critics default to viewing AI developments critically (but will generally view AI as at least acceptable when it has what they view as an outsized positive effect).
These “pop” critics, thus, appear consistent with (4).
Some of the people proclaiming to view the aggregate of the current tragetory of AI as the most negative (i.e., they place the most credence in estimations of a very high extinction risk) seem among the most enthusiastic hen it comes to promoting AI use cases and development.
This, facially, appears contradictory with (4).
To be clear, I don’t think there’s anything wrong with having deontological principles. I try to follow some deontological principles myself! Here are some things you’ve said that make me think you’re not using entirely utilitarian reasoning (bolding mine):
Utilitarianism would ask: Do the benefits outweigh the harms? But a deontology also cares about the bolded terms: Is the act inherently negative? Should it be avoided by default? Is there an unacceptable harm that cannot be balanced by any amount of benefit? Is the act intentional or inadvertent? Is there an extenuating motivation?
In fact my first comment in this thread recommended that you do the utilitarian thing and estimate the benefits and harms. Instead you responded with arguments that, to my eye, would let you stick to your conclusion even if the benefits turned out to technically outweigh the harms.
What’s a concrete example of an ordinary, everyday act of compassion you’d wholeheartedly endorse?
Yes, and I am saying you can (under the described beliefs) use carbon emissions or promoting AI as heuristics for the harms outweighing the benefits. If the only information I know about an action is that it increases carbon emissions or it promotes AI then if I believe those things are—in aggregate—bad then my expectation for any given action should be that it is bad.
If I know the net expected utility of set is negative (i.e. where ). Without additional information about expected utility of any given element , I can say .
We never have perfect knowledge of utility expectations.
The actual utility calculation is unclear and depends on one’s personal values which I wasn’t working on. But going by what you gave me for gwern’s expected utility, I should be able to say:
I disagree with that. So, that’s my utilitarian counterargument.
What do you think of the (generally reasonable) deontological language I highlighted in the quotes above?
And do you want to choose a concrete example of an ordinary, everyday act of compassion you wholeheartedly endorse?
I disagree with it as well, but it is the estimated utility in the post from gwern you directed me to.
Hugging a loved one, holding the door for a stranger, giving a dollar to charity, etc.
That’s great. So now I’ll construct a homunculus of an argument for a conclusion that neither of us agrees with, but which follows the structure, form, and abstract principles you have used. Notice how the principles feel less sound this time around.
“You shouldn’t hug your loved ones”, says the homunculus, “because it encourages close contact between people, which leads to the spread of infectious disease, including catastrophic pandemics. You should instead practice social distancing, which discourages close contact.
“Pandemics pose substantial, catastrophic risks, including loss of life. So encouraging or promoting close physical contact is something one should view as inherently negative. One should default to not hugging, and only hug when the benefits clearly outweigh the harms.
“In this case, the benefit is pretty much zero. I doubt you had your wealth substantially enhanced by this hug. The benefit is, what, some momentary comfort for you and your loved one? Increasing pandemic risk for your own comfort is antisocial behavior. If everyone did that, the effects would add up. It has negative externalities, and you shouldn’t do it.
“For someone who supposedly supports the idea of public health, you’re remarkably enthusiastic about supporting such disease-promoting behavior. Eyebrow-raising, to say the least.”
Obviously we don’t agree with the homunculus’ conclusion, so something is wrong with its argument. Maybe one or more of its general principles is flawed, or has unaddressed limitations?
I do not believe the net effect of human contact is negative.
Human-human contact does not promote aggregate risk in the case of catastrophic disease in the same way in our current environment.
If I believed that the net expected utility of human-human contact was negative, then yes I would be skeptical of human-human contact.
I do not believe that.
Yes, when the net utility from human-human physical contact is negative, I do consider increasing that risk as anti-social behavior and yes the risk do add up. In normal circumstances, the net utility is not negative in some special circumstances (i.e., during a pandemic) the net utility can be negative, in which case yes you should seek to avoid human-human physical contact. Not doing so for your own comfort (e.g., throwing parties or social events involving close physical contact during a pandemic) is antisocial behavior.
“It does promote pandemic risk”, says the homunculus. “You’re making physical contact look more appealing. It shifts the public sentiment. You’re going around promoting the greatness of physical contact, undermining your public-health messaging and contributing to a positive-feedback cycle of increasing physical contact and positive sentiment for the same. There’s no amount that doesn’t contribute.
“The expected harm of promoting a deadly pandemic even a little bit is pretty substantial. Anything that increases that risk should be difficult to justify. You would need to see a pretty large benefit to outweigh the risk, and as I said the benefit in this case is pretty much zero.”
I don’t know what case you think you are responding to, but it isn’t one I made.
When I think the net expected utility of (marginal) actions within a set of actions is negative (in the current world), then I am skeptical of actions in that set. I do not think the net expected utility of all human contact is negative. If you believe the expected outcome of AI is development is negative, I cannot imagine you do not also think the net expected utility of all additional contributions to AI is also negative. If you think that, then you should be skeptical towards any contribution.
That is exactly the attitude I take currently with carbon emissions, for example. I think the net expected utility of all (marginal) anthropogenic carbon emissions are negative. Therefore, I try to avoid actions that increase carbon emissions and default to viewing carbon emitting activities as negative. As I have repeated, that isn’t to say actions that increase carbon emissions are always bad, sometimes they may be justified by greater social utility and sometimes we may be willing to do them for our personal benefits—we are only human after all.
In certain periods, I have taken this attidude with outside-human contact as well. When I expect the net expected utility of all additional contact between people is negative (e.g., during a major pandemic), then I am skeptical towards participating in any activity that would promote or encourage that kind of contact.
Ah, I hadn’t appreciated how central that principle was in your thinking. Attempting to say it in my own words: “If there’s a set of actions such that taking all the actions would be worse than refraining from all the actions, because of the net harm, then we should presume that taking any particular action in that set is worse than not taking it, absent a really good reason.”
I’ll note this principle is quite sensitive to how you group actions into families.
Take the set of all carbon emissions starting now. If we suddenly refrained from all of them, we’d face famine pretty soon, which would be worse than global warming. So the principle doesn’t apply to marginal carbon emissions (in the sense of emissions we’re about to make but haven’t made yet).
If instead you take the set of future carbon emissions minus the ones we were “going to do anyways”, in the course of our daily lives, then you probably think it’s better to refrain from all these “extra” emissions, so by your principle one should by default refrain from any particular marginal emission. (“Marginal” here means going out of one’s way, doing something non-ordinary.)
(But what if one was planning on going into a career in oil drilling? Is that an “extra” action?)
The homunculus says, “We already had SARS, MERS, and COVID. We’ll surely get a fourth coronavirus epidemic soon, and the next one could be deadly enough to outweigh the benefit of the preceding decade of human contact. Better to refrain from all but medically necessary human contact until we develop broad-spectrum vaccines against all coronaviruses. Therefore any particular hug should be avoided by default.”
I’d be happy to give up the entire deep learning revolution to avoid existential risk from AI, so your principle would have me refrain from praising Google Translate in 2017.
I would drop the ‘absent a really good reason.’ It isn’t necessarily a strong recommendation—it is approximately a default assumption. I previously formalized it as: If I know the net expected utility of set is negative (i.e. where ). Without additional information about expected utility of any given element , I can say .
Absent further information and evaluation of an action, the default expectation should be it is negative (in proportion to how negative you view the aggregate and a weighting of the factor as part of the whole).
I think you are misunderstanding what I the expected utility function would work out to, the counter-factual isn’t “stop all carbon emissions,” it is a probability distribution of what we expect carbon emissions to be. It would be:
Where j represents each possible amount of future emissions, represents the probability of j emissions, represents the expected utility of the consequences of j level of emissions.
We can also say:
Because the probability distribution of carbon emissions is well beyond a level that is safe and the bottom-end of the distribution hardly touches on a level where we would expect more harm from famine, the expected utility of future carbon emissions works out to be negative.
Our actions/inactions nudge the probability curve slightly one way or another, they do not flip it on and off. IF we got to the point where the probability distribution was such that it was well below catastrophic risk, then the EU would become positive.
If you believe that risk is so substantial and human contact is part of that risk, then yes I would say you should try to minimize human contact. I do not believe the risk is that substantial, I believe that in net marginal human-human contact is a social neutral-to-positive. There could be something deadly enough to outweigh the benefit, but I would estimate the risk at a substantially lower-level.
Thanks for clarifying. Here’s my main criticism of your principle.
I basically agree with this. If all you know about an action is that it belongs to a set of actions with presumed independent effects that add up to negative value, then the expected value of that action is negative.
However, you usually have additional information about the action, so the argument doesn’t apply. So your principle is very weak.
I think this is the answer to your question at the top of this thread: Why do people concerned about catastrophic risks from AI sometimes do things in the direction of supporting AI? Because they know the details of the particular actions they’re taking, and based on those details they judge those actions to have positive expected value. (Ideally! Many people are nincompoops who do things for bad reasons. But I think the Unslop contest was fine.) That’s why I said “you have to disagree on the numbers”.
I think you tend to use a stronger version of the principle which is something like, “Come on, there’s no way you have enough information to know that it’s positive expected value, something funny is going on.” To which I would say, “No really, let’s talk about the numbers, it checks out.”
I’m running out of energy to spend on this thread, though.
Also, I should say, that when you plug in the numbers you get big negatives. Let’s say i estimate at the current probilibility distribution of AI development has a net expected utility of say −1,000,000,000,000 ( say, my p(doom) is ~1/8, i think 8 billion humans will die and i value each human at 1,000).
If i think my actions (including positive feedback!) are contributing 1⁄100,000,000th of AI development, then my expectated utility from the AI development I contribute should be ~ −10,000 (about the same as the value I would place on 10 human lives). So unless I think i am saving 10 lives (or equivelent) in contributing 1 hundredth of a millionth of AI development, I should find my actions negative.
Current total investment in AI is under 2 trillion. Being conservative (a dollar now means more than a dollar later!) you should value contributing $2,000 to AI development as being roughly as negative as killing a stranger (under a simplified framework, that just looks at dollars).
Now we’re getting somewhere. You’re assuming that providing one billionth of this year’s investment in AI means decreasing the expected utility from human extinction by one billionth. But that’s not right. If you invest in AI, that means the AI companies will choose to raise less money from other investors. The net effect is that AI companies will raise a little more money than if you hadn’t invested (but less than the amount you invest), get a better interest rate, and spend less effort on fundraising this year. And this will have some effect on the probability of an extinction-level AI catastrophe this generation, but the effect isn’t linear in the amount you invest.
In a parallel thread, your theory of impact was that divestment has an effect on regulation. That’s also nonlinear.
Investing in AI is something I’m likely to do. (Arguably >10% of the S&P 500 by market cap use marginal investment to try to build frontier models.) Here’s how I justify it: The amount of stock I can buy won’t move the stock price perceptibly. If I estimated the price movement it would be tiny. That would translate into AI companies expending a tiny bit less effort on selling stock this year to fund datacenters, which accelerates timelines a little bit, which increases existential risk a very tiny amount. On the other hand, if I invest my savings in a broad-market index fund that includes AI, that’ll increase my savings and I’ll have more time to spend doing AI safety research, calling my Congressperson, etc., which are more directly impactful. Plus, I have lots of personal uses for more money.
I’m basically satisfied with where this thread ended up. (I just wanted to convey that it’s in general reasonable to do things that seem individually worth it, even if they seem to be promoting AI in some way.) I want to flag that I won’t have the energy to find a crux in our disagreement over the Unslop contest or whether I should buy index funds. Though if you have a good argument about why not to buy index funds I haven’t thought of, I’ll be interested to hear that.
You seem to be confusing some things, probably because I was not as precise with definitions of terms as I should have been.
Divestment was an example of a common behavior activists engage in (which I would broadly speaking endorse). The impact from divestment on a industry is different from the negative impact of funding, advertising them, and paying an industry. They not direct inversions of each other, divesting is a very different activity.
I said “contributing to AI development” here, not investing. What is meant by “investing” is more nuanced (e.g. buying stocks second-hand on the open market is not actually ‘investing’ in the economic sense).
Here, I am including something as simple as buying AI products (a substantial portion of money sent will contribute to AI development).
This is (generally) the exact opposite of how investment works. Me investing a dollar, encourages Steve to invest a dollar which provides collateral to borrow another 2 dollars. Me not investing a dollar encourages Steve not to invest a dollar. Me spending a dollar a product boosts that company’s profits, encouraging others to invest more dollars and moving outward the company’s expected demand, encouraging them to invest internally in expanding that product.
Investments follow trends, both of other investments and of consumer behavior, moving investment towards one thing moves more investment towards it.
A very tiny increase in an extreme risk is very negative. And remember, these are aggregate risks that cover some probability distribution of outcomes, not just flat +/- x risk. It is accurate, as such, to estimate the EU as a portion of the whole, not just the absolute movement in risk.
Assuming you aren’t a private equity firm or buying directly in IPOs, your personal finances have minimal direct impact on the economic investment which is the investment one cares about. Divesting from exposed firms is a signaling behavior and really most useful when the divestment is coming from funds that are actually providing investment, not just trading on the open market (divestment movements generally focus on capital funds and large endowments).
Yes, it is a weak default case, but it generalizes. If rational, the people who find the aggregate risk the most serious should also have the highest threshold for engaging in any behavior increasing that risk.
This largely holds when it comes to the public actions of e.g. climate activists and even what i called “pop” AI criticis, but the opposite is true for some of the people who proclaim to be the most critical of AI.
Then give the numbers. The only concrete numbers you gave me was gwern thinking there was negative expected utility to the production of new fiction, which runs directly counter to your claimed rationale.