idea: we should make a pledge for lab safety employees to sign, that goes something like this:
i will not work on capabilities.
before taking on any project, even if it is billed as safety, i will consider whether it is net positive and only work on it if it is.
i am aware that working on capabilities (or fake safety) might be an easier path to promotion, or earn more respect from colleagues. i am committed to resisting these incentives.
edit: an additional item
i will keep in touch with people with safety opinions that i respect, who do not work at any frontier lab, in order to expose my theory of impact and ideas to criticism
I think people who work on safety at big labs are enablers for capabilities anyway, so this pledge won’t do much. It’s better to just quit. More details in this old comment.
I think you’re making a mistake whereby any capabilities contribution is legible, while contribution to safety is illegible, and so you’re encouraging people to make massive sacrifices on safety in order to avoid a little contribution to capabilities.
To try to illustrate why this tradeoff is so bad:
At the level of the marginal worker quitting, nothing really changes except that the lab becomes less risk-aware. People influence the culture where they work. The average person who replaces them will be a tiny bit less competent and a lot less risk-aware.
Imagining this scaled up, in a world where large numbers of people took your advice sooner and quit the field, we get all this rogue agent stuff happening ~1 year later with labs that have been filtered hard to not care about safety, and so we get less disclosure, less acknowledgement of the issue, and more papering over the problem, with less possibility of a pause, regulation, etc. because the labs aren’t pushing for it.
At the level of the marginal worker quitting, nothing really changes except that the lab becomes less risk-aware. People influence the culture where they work.
That’s not the right counterfactual though. You can join an organization to change it from inside, or spend the same amount of effort fighting the organization from outside (e.g. by speaking or organizing against it). Which works better depends on how “far gone” the organization is. I think the good that e.g. Kokotajlo did by quitting is bigger than the good he could’ve achieved by staying.
capabilities contribution is legible, while contribution to safety is illegible
My linked comment gave the example of RLHF, a safety contribution (or at least billed as such) which was quite legible and ended up speeding up the race a lot. The same can be said about the “helpful honest harmless assistant” idea.
fwiw, my pledge would have forbidden me from working on directly making RLHF work better, or working on HHH, or model spec stuff, etc (and indeed, i have never worked on such things at openai, even though i have had many opportunities to; the closest is my RLHF goodharting work, which intentionally focuses on how to study goodharting rather than how to make RLHF better in general).
My linked comment gave the example of RLHF, a safety contribution (or at least billed as such) which was quite legible and ended up speeding up the race a lot. The same can be said about the “helpful honest harmless assistant” idea
I think one crux we have is that your view is that we should have put this race off for as long as we could, so that we could solve or at least make significant progress on the alignment problem before we ever get to this point; whereas that has always seemed like a nonstarter to me, we needed to know what AGI systems actually look like and how they’re actually trained in order to make progress on the real alignment problem. Because alignment will be heavily dependent on the particulars of the systems we’re building and real experience, safety will largely be decided by people who are unafraid to roll up their sleeves and do capability work. Anyone who is avoiding contributing to capabilities will, for a sufficiently paranoid definition of “contribute to capabilities”, be a nonfactor.
RLHF is a good example. Like you said, RLHF is both safety and capabilities; the same will be true of future alignment techniques. Trying to avoid capability contributions also means avoiding alignment contributions.
Trying to avoid capability contributions also means avoiding alignment contributions.
I actually agree with you on this. My most-preferred future is a bit different: slow down AI overall (both capabilities and alignment) so that other things can happen in the meanwhile. If building AI is widely seen as a bad thing, talented folks feel dirty for joining it (instead of feeling virtuous because 80K hours is recommending AI careers), society has more time to build defenses (like algorithmic liability), near-AI gets planted more widely in society before growing too fast (thus making unilateral takeoffs harder), technologies that are complementary rather than substitute for humans get comparatively more time and investment (like intelligence amplification, genetic engineering, or thought interfaces), and failing all that, at least humanity gets a little more time to survive.
I’m not sure I have a most-preferred future—I guess it would be everyone becoming aware of AI risk and the U.N. with US and China at the helm co-ordinating a joint pause and eventual AGI project. But people who point out this is super dangerous itself due to power centralization are correct, and on top of that I think it’s possible that instead of co-ordination, everyone becoming aware of AI risk leads to racing, sabotage, and other drastic actions.
The current situation where the leading labs are moderately risk aware and the rest of the world doesn’t seem that freaked out, while still being freaked out enough to potentially apply some brakes and regulations, is pretty good I think.
Your preferred future seems both nearly-impossible, not that good if it happens, and if we fall short of total victory on your goal then it comes out strongly negative—if we shame AI researchers and only some people worried about AI risk quit, and most people not worried about it stay, then all we’ve done is slow things down 1% and decreased the awareness of the field by 10%. I agree that more time to adjust and develop things like algorithmic liability would be very helpful though, I just think the price of buying that extra time through researcher-exit is absurdly high.
Well, a version of the same plan already basically worked: discouraging nuclear power and encouraging renewables instead. We need the same kind of society-wide push against AI and for intelligence amplification / human-complementary tech. Not just discouraging a subset of researchers, but also creating regulatory hurdles, convincing the public that AI is harmful, and so on. I see nothing impossible about this.
The other stuff is fine, I’m making the narrow point that discouraging the subset of researchers who care about AI risk from working in the field is bad.
depends on the kind of safety work. i think there’s stuff that’s net positive. average lab safety work is probably bad though. picking net good things requires a lot of thinking and epistemic virtue. the goal of the pledge is to nudge people towards that
i think you also need to make sure pledge-takers come in frequent contact with people outside the labs who have good safety takes. so maybe i should add that as a pledge item
Refusing to work on capabilities is neither necessary nor sufficient for lab safety employees to have high impact. What about something like this?
Before taking on any project, even if it is billed as safety, i will consider whether it’s close to the maximum possible impact I could make, and only work on it if it is.
I will count capabilities externalities of my work in this estimate, which could make the overall impact net negative.
I am aware that working on other things might be an easier path to promotion, or earn more respect from colleagues. i am committed to resisting these incentives.
I think there’s a natural bias toward overestimating positive impact and underestimating capability externalities. Overall this could easily lead to a community with overall negative impact, and I’d guess it has. If so, it might be optimal to try to add an explicit corrective bias against this. For this reason I prefer Leo’s statement.
fwiw, although i’ve never “officially” pledged anything, i have de facto operated by approximately these constraints the entire 5 years that i’ve been at openai, and i intend to continue to follow them
Besides pledges by employees, it seems like it would be valuable if people concerned by AI risks pledged not to seek employment at companies developing frontier AIs.
AI safety researchers will be subject to particularly bad incentives when they are employed by companies that give capability higher priority than safety, and whose safety statements have low credibility. Non-employment pledges avoid those bad incentives.
By bringing up non-employment pledges, I don’t mean to argue against pledges by employees. Both kinds of pledges can be especially useful if many people take them, making it common knowledge that peers are concerned enough to change behaviors.
Why is this something we should expect many people to abide by? The Turner affair has shown us just how lightly many people take their private and even public pledges of this precise form; when the chips are down, people make excuses and say nothing. If they say nothing, why shouldn’t expect them to then also do nothing?
Interesting! I had a different idea for a pledge, more of a “giving what we can” or “farmkind” style where lab employees commit to offset some % or some amount of their capabilities externalities by donating to safety causes.
Also an idea for a “union for lab employees” to help increase their bargaining power on topics like safety or redistribution of wealth in society.
People seem not to like capabilities offsetting, curious to hear why. Obviously there are measurement problems (“how much safety are you actually buying?”) and moral licensing effects to avoid. But from a purely econ-brained perspective it seems like a reasonable thing to try, if you can get the numbers right?
I guess, I’d also just be interested in the project of figuring out what % of a lab employee’s salary they’d need to donate to safety, for the impact to net out even. 1%? 10%? 50%? 100%? 200%? more?
Yes, the econ-brained part is the problem. In general you should think of economic-style reasoning as a way of subverting moral obligations. More on this in the forthcoming part 2 of my retrospective.
(As one extreme analogy, consider what would happen if the legal system let people offset murder, or other crimes.)
Yeah I know you’re against being too econ brained (vs social brained), and I’m pretty supportive of that stance actually.
But I think the case for economic reasoning is still kind of good in a lot of places? Or, I’m very interested in figuring out when it’s appropriate to apply. (excited for your part 2!)
I think murder or several kinds of crimes seem weird to offset, but also it’s not clear that working in capabilities carries the weight of murder, in the eyes of all but a very select few atm. One might argue that the job of an AI safety movement is to make capabilities work extremely sanctioned with the moral weight of murder (or perhaps nuclear arms proliferation), but I don’t think I’d fully agree with those.
Other analogies include carbon offsets (which, seem ethically fine, though maybe the way out for climate is through growth/solar not offsets), meat consumption offsets (which I tend to think seem good, but admit are controversial). Or, paying the government for certain kinds of licenses or taxes for when your org produces negative externalities (sound/noise, pollution, gambling).
There is no offsetting creating a world ender. I think it is common for there to be things for which there is no practical offset; where the amount of offsetting measured in monetary unit is dramatically above what offsetting culture would give you.
The issue I have with this is that this is basically a strategy that doubles down on your own harms.
Like, if you are working at the labs because you work in AI safety, and then you donate your money to AI safety, you are not offsetting your harm at all, you are just doubling down on it.
In-general, I think there is a coordination-oriented perspective on offsetting that works, but I do think that one pushes towards choosing offsetting targets that are specifically trying to hedge against your harms, and I think donating (especially in a shallow way) to AI safety organizations is not actually a good target for that, in as much as your strategy is representative of the strategies of that class.
My guess is order of 100% on average though highly variable. Each unit of safety is probably 20x as good as a unit of capabilities is bad, but it’s hard to buy with money so you give most of it back.
Interesting, thanks. My heuristics are also that capabilities spend is roughly 1000:1 vs safety spend at the moment, which is partly why I feel like a universalized 10% pledge from employees might be good (forcing function to get the ratio down to 10:1, in the steady state)
In general, we’re excited both for the project of increasing total safety spend, and of targeting it better so donors get what they’re trying to buy
I think there are issues like: high-level actions don’t screen off intent or models, and adverse selection. That is, I think that a pledge to donate some amount to “safety causes” doesn’t constrain the impact at all. It could be accelerating capabilities, increasing or decreasing alignability of systems, and so on.
What is the purpose of the pledge? It could be: (a) to counter normalization of deviance, (b) to communicate trustworthiness to others or (c) to make it easier to resist internal pressures. I don’t think an offset achieves any of these.
The purpose of the pledge could be (d) make sure that the pledger is doing good by their lights. But in the absence of (a), (b) and (c), that person should just take actions that they judge to be good, and the pledge is unhelpful.
The only one that I think would meet (c) and make it much easier to resist internal pressures, would be to commit to burn basically everything you earn over ~$100k.
(“burn”, a word which here means “donate to a random ineffective charity”, which is how I personally burn money)
If you’re donating it to organization you think are doing good inthe world, then you have motive to keep earning lots of money, and doing whatever the company asks of you. If you’re burning it, then the massive financial and status incentive is massively reduced.
i don’t think there is a clean way to offset. some harms cannot be quantified in dollars, or are very difficult to quantify; some forms of helping the world do not funge against some ways of harming the world.
idea: we should make a pledge for lab safety employees to sign, that goes something like this:
i will not work on capabilities.
before taking on any project, even if it is billed as safety, i will consider whether it is net positive and only work on it if it is.
i am aware that working on capabilities (or fake safety) might be an easier path to promotion, or earn more respect from colleagues. i am committed to resisting these incentives.
edit: an additional item
i will keep in touch with people with safety opinions that i respect, who do not work at any frontier lab, in order to expose my theory of impact and ideas to criticism
I think people who work on safety at big labs are enablers for capabilities anyway, so this pledge won’t do much. It’s better to just quit. More details in this old comment.
I think you’re making a mistake whereby any capabilities contribution is legible, while contribution to safety is illegible, and so you’re encouraging people to make massive sacrifices on safety in order to avoid a little contribution to capabilities.
To try to illustrate why this tradeoff is so bad:
At the level of the marginal worker quitting, nothing really changes except that the lab becomes less risk-aware. People influence the culture where they work. The average person who replaces them will be a tiny bit less competent and a lot less risk-aware.
Imagining this scaled up, in a world where large numbers of people took your advice sooner and quit the field, we get all this rogue agent stuff happening ~1 year later with labs that have been filtered hard to not care about safety, and so we get less disclosure, less acknowledgement of the issue, and more papering over the problem, with less possibility of a pause, regulation, etc. because the labs aren’t pushing for it.
That’s not the right counterfactual though. You can join an organization to change it from inside, or spend the same amount of effort fighting the organization from outside (e.g. by speaking or organizing against it). Which works better depends on how “far gone” the organization is. I think the good that e.g. Kokotajlo did by quitting is bigger than the good he could’ve achieved by staying.
My linked comment gave the example of RLHF, a safety contribution (or at least billed as such) which was quite legible and ended up speeding up the race a lot. The same can be said about the “helpful honest harmless assistant” idea.
What if you join it to fight it from the inside? Has any employee of Anthropic ever actually been fired for speaking out against them?
fwiw, my pledge would have forbidden me from working on directly making RLHF work better, or working on HHH, or model spec stuff, etc (and indeed, i have never worked on such things at openai, even though i have had many opportunities to; the closest is my RLHF goodharting work, which intentionally focuses on how to study goodharting rather than how to make RLHF better in general).
I think one crux we have is that your view is that we should have put this race off for as long as we could, so that we could solve or at least make significant progress on the alignment problem before we ever get to this point; whereas that has always seemed like a nonstarter to me, we needed to know what AGI systems actually look like and how they’re actually trained in order to make progress on the real alignment problem. Because alignment will be heavily dependent on the particulars of the systems we’re building and real experience, safety will largely be decided by people who are unafraid to roll up their sleeves and do capability work. Anyone who is avoiding contributing to capabilities will, for a sufficiently paranoid definition of “contribute to capabilities”, be a nonfactor.
RLHF is a good example. Like you said, RLHF is both safety and capabilities; the same will be true of future alignment techniques. Trying to avoid capability contributions also means avoiding alignment contributions.
I actually agree with you on this. My most-preferred future is a bit different: slow down AI overall (both capabilities and alignment) so that other things can happen in the meanwhile. If building AI is widely seen as a bad thing, talented folks feel dirty for joining it (instead of feeling virtuous because 80K hours is recommending AI careers), society has more time to build defenses (like algorithmic liability), near-AI gets planted more widely in society before growing too fast (thus making unilateral takeoffs harder), technologies that are complementary rather than substitute for humans get comparatively more time and investment (like intelligence amplification, genetic engineering, or thought interfaces), and failing all that, at least humanity gets a little more time to survive.
I’m not sure I have a most-preferred future—I guess it would be everyone becoming aware of AI risk and the U.N. with US and China at the helm co-ordinating a joint pause and eventual AGI project. But people who point out this is super dangerous itself due to power centralization are correct, and on top of that I think it’s possible that instead of co-ordination, everyone becoming aware of AI risk leads to racing, sabotage, and other drastic actions.
The current situation where the leading labs are moderately risk aware and the rest of the world doesn’t seem that freaked out, while still being freaked out enough to potentially apply some brakes and regulations, is pretty good I think.
Your preferred future seems both nearly-impossible, not that good if it happens, and if we fall short of total victory on your goal then it comes out strongly negative—if we shame AI researchers and only some people worried about AI risk quit, and most people not worried about it stay, then all we’ve done is slow things down 1% and decreased the awareness of the field by 10%. I agree that more time to adjust and develop things like algorithmic liability would be very helpful though, I just think the price of buying that extra time through researcher-exit is absurdly high.
Well, a version of the same plan already basically worked: discouraging nuclear power and encouraging renewables instead. We need the same kind of society-wide push against AI and for intelligence amplification / human-complementary tech. Not just discouraging a subset of researchers, but also creating regulatory hurdles, convincing the public that AI is harmful, and so on. I see nothing impossible about this.
The other stuff is fine, I’m making the narrow point that discouraging the subset of researchers who care about AI risk from working in the field is bad.
depends on the kind of safety work. i think there’s stuff that’s net positive. average lab safety work is probably bad though. picking net good things requires a lot of thinking and epistemic virtue. the goal of the pledge is to nudge people towards that
One downside (likely smaller than the upside): further incentivize pledgers to do motivated reasoning about whether their work is bad.
i think you also need to make sure pledge-takers come in frequent contact with people outside the labs who have good safety takes. so maybe i should add that as a pledge item
Refusing to work on capabilities is neither necessary nor sufficient for lab safety employees to have high impact. What about something like this?
Before taking on any project, even if it is billed as safety, i will consider whether it’s close to the maximum possible impact I could make, and only work on it if it is.
I will count capabilities externalities of my work in this estimate, which could make the overall impact net negative.
I am aware that working on other things might be an easier path to promotion, or earn more respect from colleagues. i am committed to resisting these incentives.
I think there’s a natural bias toward overestimating positive impact and underestimating capability externalities. Overall this could easily lead to a community with overall negative impact, and I’d guess it has.
If so, it might be optimal to try to add an explicit corrective bias against this. For this reason I prefer Leo’s statement.
fwiw, although i’ve never “officially” pledged anything, i have de facto operated by approximately these constraints the entire 5 years that i’ve been at openai, and i intend to continue to follow them
Besides pledges by employees, it seems like it would be valuable if people concerned by AI risks pledged not to seek employment at companies developing frontier AIs.
AI safety researchers will be subject to particularly bad incentives when they are employed by companies that give capability higher priority than safety, and whose safety statements have low credibility. Non-employment pledges avoid those bad incentives.
By bringing up non-employment pledges, I don’t mean to argue against pledges by employees. Both kinds of pledges can be especially useful if many people take them, making it common knowledge that peers are concerned enough to change behaviors.
Why is this something we should expect many people to abide by? The Turner affair has shown us just how lightly many people take their private and even public pledges of this precise form; when the chips are down, people make excuses and say nothing. If they say nothing, why shouldn’t expect them to then also do nothing?
Interesting! I had a different idea for a pledge, more of a “giving what we can” or “farmkind” style where lab employees commit to offset some % or some amount of their capabilities externalities by donating to safety causes.
Also an idea for a “union for lab employees” to help increase their bargaining power on topics like safety or redistribution of wealth in society.
People seem not to like capabilities offsetting, curious to hear why. Obviously there are measurement problems (“how much safety are you actually buying?”) and moral licensing effects to avoid. But from a purely econ-brained perspective it seems like a reasonable thing to try, if you can get the numbers right?
I guess, I’d also just be interested in the project of figuring out what % of a lab employee’s salary they’d need to donate to safety, for the impact to net out even. 1%? 10%? 50%? 100%? 200%? more?
Yes, the econ-brained part is the problem. In general you should think of economic-style reasoning as a way of subverting moral obligations. More on this in the forthcoming part 2 of my retrospective.
(As one extreme analogy, consider what would happen if the legal system let people offset murder, or other crimes.)
Yeah I know you’re against being too econ brained (vs social brained), and I’m pretty supportive of that stance actually.
But I think the case for economic reasoning is still kind of good in a lot of places? Or, I’m very interested in figuring out when it’s appropriate to apply. (excited for your part 2!)
I think murder or several kinds of crimes seem weird to offset, but also it’s not clear that working in capabilities carries the weight of murder, in the eyes of all but a very select few atm. One might argue that the job of an AI safety movement is to make capabilities work extremely sanctioned with the moral weight of murder (or perhaps nuclear arms proliferation), but I don’t think I’d fully agree with those.
Other analogies include carbon offsets (which, seem ethically fine, though maybe the way out for climate is through growth/solar not offsets), meat consumption offsets (which I tend to think seem good, but admit are controversial). Or, paying the government for certain kinds of licenses or taxes for when your org produces negative externalities (sound/noise, pollution, gambling).
More links https://slatestarcodex.com/2015/01/04/ethics-offsets/ and https://sideways-view.com/2021/03/21/robust-egg-offsetting/ and https://www.jefftk.com/p/why-im-not-vegan
There is no offsetting creating a world ender. I think it is common for there to be things for which there is no practical offset; where the amount of offsetting measured in monetary unit is dramatically above what offsetting culture would give you.
The issue I have with this is that this is basically a strategy that doubles down on your own harms.
Like, if you are working at the labs because you work in AI safety, and then you donate your money to AI safety, you are not offsetting your harm at all, you are just doubling down on it.
In-general, I think there is a coordination-oriented perspective on offsetting that works, but I do think that one pushes towards choosing offsetting targets that are specifically trying to hedge against your harms, and I think donating (especially in a shallow way) to AI safety organizations is not actually a good target for that, in as much as your strategy is representative of the strategies of that class.
My guess is order of 100% on average though highly variable. Each unit of safety is probably 20x as good as a unit of capabilities is bad, but it’s hard to buy with money so you give most of it back.
Interesting, thanks. My heuristics are also that capabilities spend is roughly 1000:1 vs safety spend at the moment, which is partly why I feel like a universalized 10% pledge from employees might be good (forcing function to get the ratio down to 10:1, in the steady state)
In general, we’re excited both for the project of increasing total safety spend, and of targeting it better so donors get what they’re trying to buy
I think there are issues like: high-level actions don’t screen off intent or models, and adverse selection. That is, I think that a pledge to donate some amount to “safety causes” doesn’t constrain the impact at all. It could be accelerating capabilities, increasing or decreasing alignability of systems, and so on.
What is the purpose of the pledge? It could be: (a) to counter normalization of deviance, (b) to communicate trustworthiness to others or (c) to make it easier to resist internal pressures. I don’t think an offset achieves any of these.
The purpose of the pledge could be (d) make sure that the pledger is doing good by their lights. But in the absence of (a), (b) and (c), that person should just take actions that they judge to be good, and the pledge is unhelpful.
The only one that I think would meet (c) and make it much easier to resist internal pressures, would be to commit to burn basically everything you earn over ~$100k.
(“burn”, a word which here means “donate to a random ineffective charity”, which is how I personally burn money)
If you’re donating it to organization you think are doing good inthe world, then you have motive to keep earning lots of money, and doing whatever the company asks of you. If you’re burning it, then the massive financial and status incentive is massively reduced.
i don’t think there is a clean way to offset. some harms cannot be quantified in dollars, or are very difficult to quantify; some forms of helping the world do not funge against some ways of harming the world.