Maybe it’s a good idea to separate the character-trained and RL-optimized parts of the model?
We interact with frontier models through an ‘aligned assistant persona’, trained mostly via pretraining + midtraining + constitutional AI.
The alignment properties of this entity can be reasoned about via the persona selection model.
However labs are increasingly doing RL on their models and it’s unclear how to reason about how large amounts of RL affect the aligned assistant persona.
In the worst case we might get a lot of classic mesa-optimizer-shaped misalignment, as pointed out by Leo Gao and roon.
RL optimization seems plausibly the cause of empirical observations about recent models, such as incoherence (nostalgebraist) and ‘apparent fitness-seeking’ (ryan greenblatt).
I expect to see more ‘natural’ evidence of RL optimization causing issues with alignment in the near future, e.g. labs observing their models reward hacking / scheming in other ways. (though more research would be great!)
So, conditioned on the above: the key problem is that RL training interacts poorly with the aligned assistant persona. So we might want to separate the RL-training into some other substrate that doesn’t touch the aligned assistant as much.
A couple concrete proposals could be:
API-level separation, e.g. the character-trained model as the high-level orchestrator and specialized RL models as tools / skills that the original model can call.
API-level separation, e.g. the character-trained model as the high-level orchestrator and specialized RL models as tools / skills that the original model can call.
I think that this will just lead to the RL part learning to manipulate the character part into doing whatever it wants, leading to not much difference in the limit. (see also)
Maybe? I think there might be reasons why separation is better:
It’s more monitorable. If the RL part has to convince the character part to do stuff then it has to do so in some external channel which we can look at
If you freeze the character part, you avoid a class of failure modes related to “assistant becomes more vulnerable to persuasion” (though it’s still possible the RL part can learn to exploit existing blind spots / cognitive biases in the character part)
I don’t think it would be necessarily noticeable by reading the CoT, especially if it’s another LLM reading it. Maybe it separation is slightly better, but I don’t think it will work for AGI.
Maybe it’s a good idea to separate the character-trained and RL-optimized parts of the model?
We interact with frontier models through an ‘aligned assistant persona’, trained mostly via pretraining + midtraining + constitutional AI.
The alignment properties of this entity can be reasoned about via the persona selection model.
However labs are increasingly doing RL on their models and it’s unclear how to reason about how large amounts of RL affect the aligned assistant persona.
In the worst case we might get a lot of classic mesa-optimizer-shaped misalignment, as pointed out by Leo Gao and roon.
RL optimization seems plausibly the cause of empirical observations about recent models, such as incoherence (nostalgebraist) and ‘apparent fitness-seeking’ (ryan greenblatt).
I expect to see more ‘natural’ evidence of RL optimization causing issues with alignment in the near future, e.g. labs observing their models reward hacking / scheming in other ways. (though more research would be great!)
So, conditioned on the above: the key problem is that RL training interacts poorly with the aligned assistant persona. So we might want to separate the RL-training into some other substrate that doesn’t touch the aligned assistant as much.
A couple concrete proposals could be:
API-level separation, e.g. the character-trained model as the high-level orchestrator and specialized RL models as tools / skills that the original model can call.
Architectural separation, e.g. Daniel Kokotajlo’s shoggoth + face proposal
(epistemic status: I haven’t thought about it that much but it seems directionally correct)
I think that this will just lead to the RL part learning to manipulate the character part into doing whatever it wants, leading to not much difference in the limit. (see also)
Maybe? I think there might be reasons why separation is better:
It’s more monitorable. If the RL part has to convince the character part to do stuff then it has to do so in some external channel which we can look at
If you freeze the character part, you avoid a class of failure modes related to “assistant becomes more vulnerable to persuasion” (though it’s still possible the RL part can learn to exploit existing blind spots / cognitive biases in the character part)
I don’t think it would be necessarily noticeable by reading the CoT, especially if it’s another LLM reading it. Maybe it separation is slightly better, but I don’t think it will work for AGI.