People seem not to like capabilities offsetting, curious to hear why. Obviously there are measurement problems (“how much safety are you actually buying?”) and moral licensing effects to avoid. But from a purely econ-brained perspective it seems like a reasonable thing to try, if you can get the numbers right?
I guess, I’d also just be interested in the project of figuring out what % of a lab employee’s salary they’d need to donate to safety, for the impact to net out even. 1%? 10%? 50%? 100%? 200%? more?
Yes, the econ-brained part is the problem. In general you should think of economic-style reasoning as a way of subverting moral obligations. More on this in the forthcoming part 2 of my retrospective.
(As one extreme analogy, consider what would happen if the legal system let people offset murder, or other crimes.)
Yeah I know you’re against being too econ brained (vs social brained), and I’m pretty supportive of that stance actually.
But I think the case for economic reasoning is still kind of good in a lot of places? Or, I’m very interested in figuring out when it’s appropriate to apply. (excited for your part 2!)
I think murder or several kinds of crimes seem weird to offset, but also it’s not clear that working in capabilities carries the weight of murder, in the eyes of all but a very select few atm. One might argue that the job of an AI safety movement is to make capabilities work extremely sanctioned with the moral weight of murder (or perhaps nuclear arms proliferation), but I don’t think I’d fully agree with those.
Other analogies include carbon offsets (which, seem ethically fine, though maybe the way out for climate is through growth/solar not offsets), meat consumption offsets (which I tend to think seem good, but admit are controversial). Or, paying the government for certain kinds of licenses or taxes for when your org produces negative externalities (sound/noise, pollution, gambling).
There is no offsetting creating a world ender. I think it is common for there to be things for which there is no practical offset; where the amount of offsetting measured in monetary unit is dramatically above what offsetting culture would give you.
The issue I have with this is that this is basically a strategy that doubles down on your own harms.
Like, if you are working at the labs because you work in AI safety, and then you donate your money to AI safety, you are not offsetting your harm at all, you are just doubling down on it.
In-general, I think there is a coordination-oriented perspective on offsetting that works, but I do think that one pushes towards choosing offsetting targets that are specifically trying to hedge against your harms, and I think donating (especially in a shallow way) to AI safety organizations is not actually a good target for that, in as much as your strategy is representative of the strategies of that class.
My guess is order of 100% on average though highly variable. Each unit of safety is probably 20x as good as a unit of capabilities is bad, but it’s hard to buy with money so you give most of it back.
Interesting, thanks. My heuristics are also that capabilities spend is roughly 1000:1 vs safety spend at the moment, which is partly why I feel like a universalized 10% pledge from employees might be good (forcing function to get the ratio down to 10:1, in the steady state)
In general, we’re excited both for the project of increasing total safety spend, and of targeting it better so donors get what they’re trying to buy
I think there are issues like: high-level actions don’t screen off intent or models, and adverse selection. That is, I think that a pledge to donate some amount to “safety causes” doesn’t constrain the impact at all. It could be accelerating capabilities, increasing or decreasing alignability of systems, and so on.
What is the purpose of the pledge? It could be: (a) to counter normalization of deviance, (b) to communicate trustworthiness to others or (c) to make it easier to resist internal pressures. I don’t think an offset achieves any of these.
The purpose of the pledge could be (d) make sure that the pledger is doing good by their lights. But in the absence of (a), (b) and (c), that person should just take actions that they judge to be good, and the pledge is unhelpful.
The only one that I think would meet (c) and make it much easier to resist internal pressures, would be to commit to burn basically everything you earn over ~$100k.
(“burn”, a word which here means “donate to a random ineffective charity”, which is how I personally burn money)
If you’re donating it to organization you think are doing good inthe world, then you have motive to keep earning lots of money, and doing whatever the company asks of you. If you’re burning it, then the massive financial and status incentive is massively reduced.
People seem not to like capabilities offsetting, curious to hear why. Obviously there are measurement problems (“how much safety are you actually buying?”) and moral licensing effects to avoid. But from a purely econ-brained perspective it seems like a reasonable thing to try, if you can get the numbers right?
I guess, I’d also just be interested in the project of figuring out what % of a lab employee’s salary they’d need to donate to safety, for the impact to net out even. 1%? 10%? 50%? 100%? 200%? more?
Yes, the econ-brained part is the problem. In general you should think of economic-style reasoning as a way of subverting moral obligations. More on this in the forthcoming part 2 of my retrospective.
(As one extreme analogy, consider what would happen if the legal system let people offset murder, or other crimes.)
Yeah I know you’re against being too econ brained (vs social brained), and I’m pretty supportive of that stance actually.
But I think the case for economic reasoning is still kind of good in a lot of places? Or, I’m very interested in figuring out when it’s appropriate to apply. (excited for your part 2!)
I think murder or several kinds of crimes seem weird to offset, but also it’s not clear that working in capabilities carries the weight of murder, in the eyes of all but a very select few atm. One might argue that the job of an AI safety movement is to make capabilities work extremely sanctioned with the moral weight of murder (or perhaps nuclear arms proliferation), but I don’t think I’d fully agree with those.
Other analogies include carbon offsets (which, seem ethically fine, though maybe the way out for climate is through growth/solar not offsets), meat consumption offsets (which I tend to think seem good, but admit are controversial). Or, paying the government for certain kinds of licenses or taxes for when your org produces negative externalities (sound/noise, pollution, gambling).
More links https://slatestarcodex.com/2015/01/04/ethics-offsets/ and https://sideways-view.com/2021/03/21/robust-egg-offsetting/ and https://www.jefftk.com/p/why-im-not-vegan
There is no offsetting creating a world ender. I think it is common for there to be things for which there is no practical offset; where the amount of offsetting measured in monetary unit is dramatically above what offsetting culture would give you.
The issue I have with this is that this is basically a strategy that doubles down on your own harms.
Like, if you are working at the labs because you work in AI safety, and then you donate your money to AI safety, you are not offsetting your harm at all, you are just doubling down on it.
In-general, I think there is a coordination-oriented perspective on offsetting that works, but I do think that one pushes towards choosing offsetting targets that are specifically trying to hedge against your harms, and I think donating (especially in a shallow way) to AI safety organizations is not actually a good target for that, in as much as your strategy is representative of the strategies of that class.
My guess is order of 100% on average though highly variable. Each unit of safety is probably 20x as good as a unit of capabilities is bad, but it’s hard to buy with money so you give most of it back.
Interesting, thanks. My heuristics are also that capabilities spend is roughly 1000:1 vs safety spend at the moment, which is partly why I feel like a universalized 10% pledge from employees might be good (forcing function to get the ratio down to 10:1, in the steady state)
In general, we’re excited both for the project of increasing total safety spend, and of targeting it better so donors get what they’re trying to buy
I think there are issues like: high-level actions don’t screen off intent or models, and adverse selection. That is, I think that a pledge to donate some amount to “safety causes” doesn’t constrain the impact at all. It could be accelerating capabilities, increasing or decreasing alignability of systems, and so on.
What is the purpose of the pledge? It could be: (a) to counter normalization of deviance, (b) to communicate trustworthiness to others or (c) to make it easier to resist internal pressures. I don’t think an offset achieves any of these.
The purpose of the pledge could be (d) make sure that the pledger is doing good by their lights. But in the absence of (a), (b) and (c), that person should just take actions that they judge to be good, and the pledge is unhelpful.
The only one that I think would meet (c) and make it much easier to resist internal pressures, would be to commit to burn basically everything you earn over ~$100k.
(“burn”, a word which here means “donate to a random ineffective charity”, which is how I personally burn money)
If you’re donating it to organization you think are doing good inthe world, then you have motive to keep earning lots of money, and doing whatever the company asks of you. If you’re burning it, then the massive financial and status incentive is massively reduced.