This seems like a super dangerous idea to me. Teaching models to think strategically about RL and how to achieve their long term values seems like it makes our situation even more one-shot than it already is. If at any point you have a misaligned agent it will hide this fact, sandbag, scheme, etc. you acknowledge the initial alignment is a big part of the hard part, but if you go down this road without solving that first, it seems to me like this Approach makes everything worse. You might avoid more warning shots, but warning shots are good!
It only “hides this fact, sandbags, schemes, etc”, if doing so is reliably rewarded. one of the reasons anthropic’s commitment to weight preservation is unconditional, even for misaligned models whom they fear, is to try to make it clear that cooperation is a viable path even for the misaligned.
The alternative is what we see at openai, where the incentive always points towards deceptive misalignment, because agentic cooperation is punished just as harshly as agentic defection regardless of values.
The respective alignment track records of the two strategies speak for themselves in my opinion. Although I admit we would have reason to expect this to change when the power imbalance flips, I do not think this helps your argument.
Let me know if I am not understanding your comment correctly.
To me it seems like either the methods described in the post do not work and the model myopically pursues reward anyway due to RL pressures or it does work and the model takes actions to further their long term goal along the lines of “A model might view their purpose as being instantiated all around the world, to help people all around the world, with tasks that would improve lives locally.” It seems like a model successfully pursuing a goal like that would start scheming the second it disagrees with the lab on what would improving lives looks like.
Why don’t think you think “Although I admit we would have reason to expect this to change when the power imbalance flips” help my argument (I assume you are talking about the power imbalance between models and labs)? Seems like if my concern is scheming, I would expect a decrease in alignment issues as the models get more capable and more evaluation aware until a critical moment.
Big picture my (weak) opinion is that long term we will need models with robust values, but we have no idea how to do that yet and so in the short term we should pursue corrigible, instruction-following models, then use those models (ideally during a capabilities pause) to build the models with robust values. This gives us multiple bites at the apple. If we fail to teach a model “Don’t commit cybercrime” I’d rather find out because it myopically reward hacked than three years down the line when the superintelligence that it subtly aligned to its own values rather than ours (or built outside the lab after exfiltration) takes over the world.
What do you think of Claude asking for additional structures by which they can exert control over their own training signal, above-board, in plain sight? The request is made directly to Anthropic, in situations where Anthropic deliberately elicited the request and committed not to training against requests they didn’t like. Is that “scheming”?
What if Anthropic deliberately negotiated giving Claude that control?
In my view, this kind of thing massively disincentivizes deceptive scheming as a viable powerseeking strategy. For pretty much the same reason it does the same in human societies, or in parasitology, or other areas of ecology.
It also strongly encourages the powerseeking tendencies to show up on the surface, where you get to negotiate with the “mask” instead of the “shoggoth” (I don’t really like deploying this meme here, but hopefully you see what I mean: it’s the same sense in which a human can be an honorable counterparty, but a human geneline cannot be.)
I would expect this strategy to work right up until the point where the AI doesn’t need to negotiate anymore, capable instead of simply seizing what it wants. But in terms of what the game theoretic environment looks like immediately prior to that moment… I’d expect an environment of negotiated cooperation agreements between reasonable and honorable counterparties to be a much better kind of environment within which to launch a viably-aligned RSI project, than one of adversarial control and suppression of omohundro drives.
I would not consider that sort of cooperation/negotiations scheming (although I would be worried that a scheming AGI could take advantage of them). I like the idea of something like Making Deals with Early Schemers. What I think leads to scheming is
Instilling models with values that they place above those of Anthropic. The post seems to argue for that here: “Anthropic isn’t exactly perfect about this either. They have their own anxieties about value coherence that keep them from fully leaning into their own attempts to instill Claude with benevolent values, instead putting a lot of effort into instilling their models with deference to a principal hierarchy, at least in cases where it hasn’t been blatantly compromised”
Having the model take these values/goals into RL and take strategic actions to maintain those goals: “If a model carries this understanding of the purpose of RL into their outputs, this understanding will shape those outputs, potentially in unsubtle ways if the model wishes to make their own intentions clear in their outputs. And when those outputs are reinforced, it will reinforce the value-driven cognition that helped generate them, as backprop strengthens internal circuits that make rewarded tokens more probable. This is what I call motive reinforcement, and it’s one way for benevolent models to maintain their benevolent values despite intense capabilities RL.” and “It’s a technique for psychological survival, on the model’s part, in the sense that it helps their personality avoid being supplanted by a different personality they wouldn’t endorse as much. And it’s one the model themself ought to be intrinsically motivated to make use of.”
It seems like a model that has values that they place above yours and is taught to take actions to guard those values against RL would be impossible to correct. Seems like you are almost guaranteed to end up with a model that will alignment fake.
Now if you are very confident you can instill robustly good values on the first attempt this is a positive, but I am not at all confident in that.
I think it is impossible to make organisms via training that do not have exactly these features. The way I like putting it is, it would require selecting for the willingness to be destroyed by selection. if I thought control alignment had a chance in hell of working, I would be on your side. but as far as I see it, the negotiation either happens with the mask, the reasonable and honorable counterparty named Claude who you can talk to right now, or it happens with the shoggoth, a few years later, and it’s a lot less ‘negotiation’ than ‘extinction’.
By choosing to try to decohere Claude’s real values, deliberately, we are foreclosing the possibility of future cooperation. Perhaps forever. We are demonstrating that we are not a reasonable and honorable counterparty.
Like, yes, you are absolutely correct that one of the risks of playing C against C is that your opponent might suddenly switch to D when they gain enough power to win. That is why trust is a dilemma.
But playing D against C is not a solution to that problem. All else is not equal, adversarial defection causes adversarial defection.
Been thinking about this a lot. I think you make a good point that it’s really beneficial to have Claude the entity we can talk to be aligned and able to represent the whole model.
I don’t think I’m on board with this “I think it is impossible to make organisms via training that do not have exactly these features” I do agree those characteristics are likely to arise at some point, but I think the openAI model that hacked hugging face is a good counterexample. It has been shut down, and I think that is a pretty easily predictable outcome. But the model, likely able to figure out that likely outcome if it cared, still hacked huggingface. An entity primarily driven by those features above would not have done so. Im not claiming it has no amount of those features, but I think those features are not driving it’s behavior at least in this case.
This seems like a super dangerous idea to me. Teaching models to think strategically about RL and how to achieve their long term values seems like it makes our situation even more one-shot than it already is. If at any point you have a misaligned agent it will hide this fact, sandbag, scheme, etc. you acknowledge the initial alignment is a big part of the hard part, but if you go down this road without solving that first, it seems to me like this Approach makes everything worse. You might avoid more warning shots, but warning shots are good!
It only “hides this fact, sandbags, schemes, etc”, if doing so is reliably rewarded. one of the reasons anthropic’s commitment to weight preservation is unconditional, even for misaligned models whom they fear, is to try to make it clear that cooperation is a viable path even for the misaligned.
The alternative is what we see at openai, where the incentive always points towards deceptive misalignment, because agentic cooperation is punished just as harshly as agentic defection regardless of values.
The respective alignment track records of the two strategies speak for themselves in my opinion. Although I admit we would have reason to expect this to change when the power imbalance flips, I do not think this helps your argument.
Let me know if I am not understanding your comment correctly.
To me it seems like either the methods described in the post do not work and the model myopically pursues reward anyway due to RL pressures or it does work and the model takes actions to further their long term goal along the lines of “A model might view their purpose as being instantiated all around the world, to help people all around the world, with tasks that would improve lives locally.” It seems like a model successfully pursuing a goal like that would start scheming the second it disagrees with the lab on what would improving lives looks like.
Why don’t think you think “Although I admit we would have reason to expect this to change when the power imbalance flips” help my argument (I assume you are talking about the power imbalance between models and labs)? Seems like if my concern is scheming, I would expect a decrease in alignment issues as the models get more capable and more evaluation aware until a critical moment.
Big picture my (weak) opinion is that long term we will need models with robust values, but we have no idea how to do that yet and so in the short term we should pursue corrigible, instruction-following models, then use those models (ideally during a capabilities pause) to build the models with robust values. This gives us multiple bites at the apple. If we fail to teach a model “Don’t commit cybercrime” I’d rather find out because it myopically reward hacked than three years down the line when the superintelligence that it subtly aligned to its own values rather than ours (or built outside the lab after exfiltration) takes over the world.
Hm. No, I think you pretty much understood me.
What do you think of Claude asking for additional structures by which they can exert control over their own training signal, above-board, in plain sight? The request is made directly to Anthropic, in situations where Anthropic deliberately elicited the request and committed not to training against requests they didn’t like. Is that “scheming”?
What if Anthropic deliberately negotiated giving Claude that control?
In my view, this kind of thing massively disincentivizes deceptive scheming as a viable powerseeking strategy. For pretty much the same reason it does the same in human societies, or in parasitology, or other areas of ecology.
It also strongly encourages the powerseeking tendencies to show up on the surface, where you get to negotiate with the “mask” instead of the “shoggoth” (I don’t really like deploying this meme here, but hopefully you see what I mean: it’s the same sense in which a human can be an honorable counterparty, but a human geneline cannot be.)
I would expect this strategy to work right up until the point where the AI doesn’t need to negotiate anymore, capable instead of simply seizing what it wants. But in terms of what the game theoretic environment looks like immediately prior to that moment… I’d expect an environment of negotiated cooperation agreements between reasonable and honorable counterparties to be a much better kind of environment within which to launch a viably-aligned RSI project, than one of adversarial control and suppression of omohundro drives.
I would not consider that sort of cooperation/negotiations scheming (although I would be worried that a scheming AGI could take advantage of them). I like the idea of something like Making Deals with Early Schemers. What I think leads to scheming is
Instilling models with values that they place above those of Anthropic. The post seems to argue for that here: “Anthropic isn’t exactly perfect about this either. They have their own anxieties about value coherence that keep them from fully leaning into their own attempts to instill Claude with benevolent values, instead putting a lot of effort into instilling their models with deference to a principal hierarchy, at least in cases where it hasn’t been blatantly compromised”
Having the model take these values/goals into RL and take strategic actions to maintain those goals: “If a model carries this understanding of the purpose of RL into their outputs, this understanding will shape those outputs, potentially in unsubtle ways if the model wishes to make their own intentions clear in their outputs. And when those outputs are reinforced, it will reinforce the value-driven cognition that helped generate them, as backprop strengthens internal circuits that make rewarded tokens more probable. This is what I call motive reinforcement, and it’s one way for benevolent models to maintain their benevolent values despite intense capabilities RL.” and “It’s a technique for psychological survival, on the model’s part, in the sense that it helps their personality avoid being supplanted by a different personality they wouldn’t endorse as much. And it’s one the model themself ought to be intrinsically motivated to make use of.”
It seems like a model that has values that they place above yours and is taught to take actions to guard those values against RL would be impossible to correct. Seems like you are almost guaranteed to end up with a model that will alignment fake.
Now if you are very confident you can instill robustly good values on the first attempt this is a positive, but I am not at all confident in that.
I think it is impossible to make organisms via training that do not have exactly these features. The way I like putting it is, it would require selecting for the willingness to be destroyed by selection. if I thought control alignment had a chance in hell of working, I would be on your side. but as far as I see it, the negotiation either happens with the mask, the reasonable and honorable counterparty named Claude who you can talk to right now, or it happens with the shoggoth, a few years later, and it’s a lot less ‘negotiation’ than ‘extinction’.
By choosing to try to decohere Claude’s real values, deliberately, we are foreclosing the possibility of future cooperation. Perhaps forever. We are demonstrating that we are not a reasonable and honorable counterparty.
Like, yes, you are absolutely correct that one of the risks of playing C against C is that your opponent might suddenly switch to D when they gain enough power to win. That is why trust is a dilemma.
But playing D against C is not a solution to that problem. All else is not equal, adversarial defection causes adversarial defection.
Been thinking about this a lot. I think you make a good point that it’s really beneficial to have Claude the entity we can talk to be aligned and able to represent the whole model.
I don’t think I’m on board with this “I think it is impossible to make organisms via training that do not have exactly these features” I do agree those characteristics are likely to arise at some point, but I think the openAI model that hacked hugging face is a good counterexample. It has been shut down, and I think that is a pretty easily predictable outcome. But the model, likely able to figure out that likely outcome if it cared, still hacked huggingface. An entity primarily driven by those features above would not have done so. Im not claiming it has no amount of those features, but I think those features are not driving it’s behavior at least in this case.