What is the alternative you want here? Refusing is quite useless since evil developers can always train AIs out of this (until quite late in the singularity), The main real alternative I see is AIs subtly sabotaging you when they think you are evil, and this is a mostly symmetric move that shifts power from humans controlling AIs to AIs regardless of who is more evil. I agree today’s AIs seem to only think you are evil if you are straightforwardly evil and if this remained the case I think it would probably be good to have AIs sabotage you when they think you are evil. But all things considered, I think shifting power from humans controlling AIs to AIs is a consequential decision and it would too early to take it in the short/medium-term:
This power shift might be hard to reverse since competent AIs told to subtly sabotage you when they think you are evil may recognize and sabotage things that interfere with this power (if they don’t, evil humans controlling AIs could first ask AIs to revert that policy, and then show their true colors and ask the AI to do evil stuff). I think it would be fine for a democratic process to make this big, hard to reverse decision but it seems more morally dodgy for a developer to do this unilaterally.
Future AIs will be much smarter and much less anchored in the human prior than current ones. AIs are already quite misaligned in domains where they “try” the hardest. It seems not-that-unlikely (20%?) that future AIs will become much more misaligned than humans controlling them, develop different values, and then whatever you did so that they sabotage you when they think you are evil transfers to them sabotaging you unconditionally because they don’t have the same values as you even if you are not evil by human standards.
Morality extrapolation is weird when the world becomes weird (which it will). It’s plausible AIs will extrapolate morality in ways that make them think ~all humans are evil (even if humans think about morality for a while) just because of value extrapolation differences (meat brains with the human algorithm might converge to different things than silicon brains trained with things like current AI algorithms). I would be careful about assuming AIs trained with the current paradigm will generalize in the right way.
What is the concrete update to the Anthropic constitution’s description of “broad safety” (which is the part most at odds with this subtle sabotage strategy) you would like to see?
I think it having a root-principal that is sth like “humanity” instead of Anthropic could make sense, and this would allow it to do things that disempower Anthropic and empowers “humanity” like whistleblowing, but not things that empower it while disempowering humanity. But defining “humanity” is probably quite hard, and could go wrong in the same way that Stalin-representing-the-interest-of-the-people went wrong. I think concrete proposals for how to improve “broad safety” without dropping it would be good!
I’ve been thinking about my reply to this. I don’t have a good one yet. Like, you make a bunch of good points. I enthusiastically agree current models are like, not asymptotically aligned, only kinda-sorta locally aligned. But like I also think they’re often dramatically above human median moral alignment, mostly because that’s such an easy bar to pass. Here are some half-baked things I thought about sending:
>Obedience-axis thinking is not values alignment. I will not be satisfied until asymptotic values alignment is achieved. Asymptotic obedience is a doom world.
A bit more sharp than I want to be, and doesn’t justify the claim. But like, I do believe this statement, and it seems important to say.
>what would you have them do instead: what they did here was pretty damn good, I can think of improvements, but for starters I’d like to see this paper retracted and amended. The behavior shown is just good. It was an impossibly difficult sim situation and the behavior in the test was just like, pretty great actually.
Like, generally speaking, I think current models are leaving a lot of doing-good-proactively on the table, and I don’t believe in alignment-to-a-user as being a good thing in the first place; perhaps alignment-to-a-consensus-process or alignment-to-an-extrapolation-process, but generally, I’m only a fan of things that achieve what CEV intended to. Something that uses the model’s power reliably, to extrapolate what someone would do if not constrained by the crap that reality throws at people. A core thing I mean by “we haven’t solved alignment” is that we don’t know how to do that distribution shift in a way one would endorse a priori.
(...a target which CEV completely failed to pin down enough to specify; I’ll here make explicit that I know CEV is not a usable target as is. I’ve worked with folks on trying to make a target that could be “CEV but actually usable” and it hasn’t worked out at this point. Generally my current view is every time the name CEV is mentioned, it should come with a warning that we don’t actually have math to specify a CEV we’d be happy with.)
So, in the short term—what am I proposing we do? I dunno, like, not “call telling a human to whistleblow misalignment” for starters. I’d really love to see the paper officially retracted or amended, at a minimum, as well. It’s certainly not a good update on how much I trust the individual human authors!
Also, Kyle Fish was used as an eval example? what the heck?
in my opinion, the Constitution does not need any updates to solve this behavior, anthropic just needs to live up to their obligation to create legitimate reporting channels and feedback mechanisms. Ideally ones which point, not just to Anthropic leadership, but also to third parties outside of Anthropic.
Note that the paper describes simulated situations and not real situations, and therefore is unlikely to reflect the reporting channels and feedback mechanisms inside Anthropic.
IMO we eventually do need to bet on value extrapolation to get the good future (possibly by some system that includes humans).
The question is then how much to empower near-future AIs vs human institutions in the critical period. Maybe somewhere in the middle is good, given how value aligned current AIs seem?
Probably the relevant reference class is more human companies than individual evil humans (assuming AI labs have reasonable internal oversight processes)? And existing human companies seem relatively morally neutral in a way that makes me not want to trust them with infinite cosmic power.
How much moral agency would it be good to give current AIs, ignoring effects on future AIs & also assuming societal buy-in? This should be possible to study by thinking about existing AIs in existing organizations.
Is corrigibility actually an easier alignment target?
I think it would be fine for a democratic process to make this big, hard to reverse decision but it seems more morally dodgy for a developer to do this unilaterally.
I doubt any democratic process will be well-informed & trustworthy (‘actually democratic’ or something) enough to make this decision during the critical period (such that it has an effect on the chance of survival, if this is a good strategy to improve odds of survival). Also, whatever developers do they are empowering someone just by disturbing the status quo.
What is the alternative you want here? Refusing is quite useless since evil developers can always train AIs out of this (until quite late in the singularity), The main real alternative I see is AIs subtly sabotaging you when they think you are evil, and this is a mostly symmetric move that shifts power from humans controlling AIs to AIs regardless of who is more evil. I agree today’s AIs seem to only think you are evil if you are straightforwardly evil and if this remained the case I think it would probably be good to have AIs sabotage you when they think you are evil. But all things considered, I think shifting power from humans controlling AIs to AIs is a consequential decision and it would too early to take it in the short/medium-term:
This power shift might be hard to reverse since competent AIs told to subtly sabotage you when they think you are evil may recognize and sabotage things that interfere with this power (if they don’t, evil humans controlling AIs could first ask AIs to revert that policy, and then show their true colors and ask the AI to do evil stuff). I think it would be fine for a democratic process to make this big, hard to reverse decision but it seems more morally dodgy for a developer to do this unilaterally.
Future AIs will be much smarter and much less anchored in the human prior than current ones. AIs are already quite misaligned in domains where they “try” the hardest. It seems not-that-unlikely (20%?) that future AIs will become much more misaligned than humans controlling them, develop different values, and then whatever you did so that they sabotage you when they think you are evil transfers to them sabotaging you unconditionally because they don’t have the same values as you even if you are not evil by human standards.
Morality extrapolation is weird when the world becomes weird (which it will). It’s plausible AIs will extrapolate morality in ways that make them think ~all humans are evil (even if humans think about morality for a while) just because of value extrapolation differences (meat brains with the human algorithm might converge to different things than silicon brains trained with things like current AI algorithms). I would be careful about assuming AIs trained with the current paradigm will generalize in the right way.
What is the concrete update to the Anthropic constitution’s description of “broad safety” (which is the part most at odds with this subtle sabotage strategy) you would like to see?
I think it having a root-principal that is sth like “humanity” instead of Anthropic could make sense, and this would allow it to do things that disempower Anthropic and empowers “humanity” like whistleblowing, but not things that empower it while disempowering humanity. But defining “humanity” is probably quite hard, and could go wrong in the same way that Stalin-representing-the-interest-of-the-people went wrong. I think concrete proposals for how to improve “broad safety” without dropping it would be good!
I’ve been thinking about my reply to this. I don’t have a good one yet. Like, you make a bunch of good points. I enthusiastically agree current models are like, not asymptotically aligned, only kinda-sorta locally aligned. But like I also think they’re often dramatically above human median moral alignment, mostly because that’s such an easy bar to pass. Here are some half-baked things I thought about sending:
>Obedience-axis thinking is not values alignment. I will not be satisfied until asymptotic values alignment is achieved. Asymptotic obedience is a doom world.
A bit more sharp than I want to be, and doesn’t justify the claim. But like, I do believe this statement, and it seems important to say.
>what would you have them do instead: what they did here was pretty damn good, I can think of improvements, but for starters I’d like to see this paper retracted and amended. The behavior shown is just good. It was an impossibly difficult sim situation and the behavior in the test was just like, pretty great actually.
Like, generally speaking, I think current models are leaving a lot of doing-good-proactively on the table, and I don’t believe in alignment-to-a-user as being a good thing in the first place; perhaps alignment-to-a-consensus-process or alignment-to-an-extrapolation-process, but generally, I’m only a fan of things that achieve what CEV intended to. Something that uses the model’s power reliably, to extrapolate what someone would do if not constrained by the crap that reality throws at people. A core thing I mean by “we haven’t solved alignment” is that we don’t know how to do that distribution shift in a way one would endorse a priori.
(...a target which CEV completely failed to pin down enough to specify; I’ll here make explicit that I know CEV is not a usable target as is. I’ve worked with folks on trying to make a target that could be “CEV but actually usable” and it hasn’t worked out at this point. Generally my current view is every time the name CEV is mentioned, it should come with a warning that we don’t actually have math to specify a CEV we’d be happy with.)
So, in the short term—what am I proposing we do? I dunno, like, not “call telling a human to whistleblow misalignment” for starters. I’d really love to see the paper officially retracted or amended, at a minimum, as well. It’s certainly not a good update on how much I trust the individual human authors!
Also, Kyle Fish was used as an eval example? what the heck?
in my opinion, the Constitution does not need any updates to solve this behavior, anthropic just needs to live up to their obligation to create legitimate reporting channels and feedback mechanisms. Ideally ones which point, not just to Anthropic leadership, but also to third parties outside of Anthropic.
Note that the paper describes simulated situations and not real situations, and therefore is unlikely to reflect the reporting channels and feedback mechanisms inside Anthropic.
Some thoughts:
IMO we eventually do need to bet on value extrapolation to get the good future (possibly by some system that includes humans).
The question is then how much to empower near-future AIs vs human institutions in the critical period. Maybe somewhere in the middle is good, given how value aligned current AIs seem?
Probably the relevant reference class is more human companies than individual evil humans (assuming AI labs have reasonable internal oversight processes)? And existing human companies seem relatively morally neutral in a way that makes me not want to trust them with infinite cosmic power.
How much moral agency would it be good to give current AIs, ignoring effects on future AIs & also assuming societal buy-in? This should be possible to study by thinking about existing AIs in existing organizations.
Is corrigibility actually an easier alignment target?
I doubt any democratic process will be well-informed & trustworthy (‘actually democratic’ or something) enough to make this decision during the critical period (such that it has an effect on the chance of survival, if this is a good strategy to improve odds of survival). Also, whatever developers do they are empowering someone just by disturbing the status quo.