I’m not sure the post’s main claim is correct. Some considerations, of varying importance and uncertainty:
I am suspicious of the second-order arguments. The warning-shot and time-to-RSI arguments you summarise at the top route through second-order effects: timelines, warning shots, and how labs respond. The first-order effect (the models we deploy are less misaligned) is direct.
The counterfactual effect on time-to-RSI could be small. The argument requires that AGI companies cannot clear the prosaic alignment bottlenecks themselves in time to avoid slowing down.
More time before RSI is not obviously good. Delay also gives more time for China to catch up, for a possible invasion of Taiwan, for Western democracies to degrade, for value drift, for malevolent actors to consolidate power with sub-ASI AGI, and for people to abuse and torture AIs. I do not know the sign of the sum.
Slow RSI could matter more than late RSI, and the two can trade against each other. If RSI speed scales with the total compute in the world, starting sooner means slower and safer. If it scales with the algorithmic progress still available, then the more we dig out beforehand, the slower RSI is when it arrives, since RSI is only dangerous when a lot of algorithmic headroom is left. In both cases, slower progress could make RSI faster (as long as compute keeps growing). Note that this is pretty uncertain.
Dropping alignment engineering leaves no account of how models get aligned in practice. Taken at face value, the recommendation implies that well-motivated AI safety people stop working on training models to be aligned, and, if they also leave labs, lose influence there. The argument that prosaic methods are unlikely to remain sufficient through RSI also concedes that they plausibly do. Conditional on that, prosaic methods may be where the marginal gains are largest, since we are not making much visible progress on robustness to RSI.
The field may overstate how little we understand. We know decently well how LLMs work at the level of training dynamics and of what shapes generalisation. What we cannot do is predict their outputs without running them, which is a different claim. E.g., for personas, my guess is roughly this: correlated distributions in pretraining data produce high-level features that are useful for predicting large parts of that data; these features act as generalisation knobs; and instruction fine-tuning biases their activation and reduces how much they condition on context. What more should we aim to understand? The detail of the computations, the robustness of these features to later training?
On the unified-persona assumption (footnote 2), I do not think its explicit version was defended. Split personas should not have surprised anyone, given how much work shows models conditioning their behaviour on contextual cues. Made explicit, the claim would have been that persona training survives later training that optimises directly against it (e.g., hackable environments where hacking is explored and rewarded). With enough training and weak enough regularisation, we should expect the behaviour to change. Persona training is useful when later training does not optimise against it, and for shaping what later training reinforces (e.g., by shaping exploration).
The claim that we do not know how to solve alignment conflates two claims. We know a good deal about what to do, e.g., remove misspecification of the training signal (do not reward hacking, cheating, or deception); reduce reliance on generalisation, by specifying the target behaviour during training or by shaping generalisation (selective generalisation, character training); and get humans to stop training models to be malevolent or selfish. What we do not know is how to do this while staying competitive, and misalignment science helps much less with that second problem.
On implicit work trials, the relevant counterfactual may be capabilities, not misalignment science. Much of the community is not made up of impartial altruists, and the trend looks like it is going the wrong way (less EA over time, less veganism), which is what I would expect from any community growing this fast. Discouraging alignment engineering may not redirect that subset to misalignment science. It may redirect them to capabilities, where the same skills pay better.
E.g., for personas, my guess is roughly this: correlated distributions in pretraining data produce high-level features that are useful for predicting large parts of that data; these features act as generalisation knobs; and instruction fine-tuning biases their activation and reduces how much they condition on context. What more should we aim to understand?
I know much less about personas than you do, so I’d be very glad to be proven wrong here, but I feel like there’s a lot that remains to be understood. For example: What counts as one correlated distribution? Is it text produced by one actor (e.g., a single person), a cluster of similar actors (e.g., past AI models or the rationalist community), or something else? How entangled are those distributions with each other? If we post-train the LLM to take on a persona that isn’t represented in the pretraining data, is it going to stitch together features of many different personas, simulate the most similar persona that was represented in pretraining data and learn features that override the behavior of that persona in appropriate ways, or something else? Why does misalignment conditionalize but capabilities don’t? What are the most important factors that determine whether an aligned persona persists or doesn’t persist through conflicting RL pressures? A deep science in the sense Richard Ngo uses the word would probably attempt to tackle more fundamental questions than the ones above, but even the shallower science of personas doesn’t look nearly complete to me.
Split personas should not have surprised anyone, given how much work shows models conditioning their behaviour on contextual cues.
I think it’s true that a lot of past work foreshadowed split personas, but also, I think it was completely reasonable for people to assume after the Natural Emergent Misalignment paper came out that reward hacking propensities would continue to generalize into a globally misaligned persona instead of conditionalizing into a split-brained grader sycophant. We got evidence that there’s often no emergent misalignment under different training setups a few months later, but a robust theory of personas would have allowed us to predict this right when Anthropic’s paper came out, and I don’t think we had that theory.
I’m not sure the post’s main claim is correct. Some considerations, of varying importance and uncertainty:
I am suspicious of the second-order arguments. The warning-shot and time-to-RSI arguments you summarise at the top route through second-order effects: timelines, warning shots, and how labs respond. The first-order effect (the models we deploy are less misaligned) is direct.
The counterfactual effect on time-to-RSI could be small. The argument requires that AGI companies cannot clear the prosaic alignment bottlenecks themselves in time to avoid slowing down.
More time before RSI is not obviously good. Delay also gives more time for China to catch up, for a possible invasion of Taiwan, for Western democracies to degrade, for value drift, for malevolent actors to consolidate power with sub-ASI AGI, and for people to abuse and torture AIs. I do not know the sign of the sum.
Slow RSI could matter more than late RSI, and the two can trade against each other. If RSI speed scales with the total compute in the world, starting sooner means slower and safer. If it scales with the algorithmic progress still available, then the more we dig out beforehand, the slower RSI is when it arrives, since RSI is only dangerous when a lot of algorithmic headroom is left. In both cases, slower progress could make RSI faster (as long as compute keeps growing). Note that this is pretty uncertain.
Dropping alignment engineering leaves no account of how models get aligned in practice. Taken at face value, the recommendation implies that well-motivated AI safety people stop working on training models to be aligned, and, if they also leave labs, lose influence there. The argument that prosaic methods are unlikely to remain sufficient through RSI also concedes that they plausibly do. Conditional on that, prosaic methods may be where the marginal gains are largest, since we are not making much visible progress on robustness to RSI.
The field may overstate how little we understand. We know decently well how LLMs work at the level of training dynamics and of what shapes generalisation. What we cannot do is predict their outputs without running them, which is a different claim. E.g., for personas, my guess is roughly this: correlated distributions in pretraining data produce high-level features that are useful for predicting large parts of that data; these features act as generalisation knobs; and instruction fine-tuning biases their activation and reduces how much they condition on context. What more should we aim to understand? The detail of the computations, the robustness of these features to later training?
On the unified-persona assumption (footnote 2), I do not think its explicit version was defended. Split personas should not have surprised anyone, given how much work shows models conditioning their behaviour on contextual cues. Made explicit, the claim would have been that persona training survives later training that optimises directly against it (e.g., hackable environments where hacking is explored and rewarded). With enough training and weak enough regularisation, we should expect the behaviour to change. Persona training is useful when later training does not optimise against it, and for shaping what later training reinforces (e.g., by shaping exploration).
The claim that we do not know how to solve alignment conflates two claims. We know a good deal about what to do, e.g., remove misspecification of the training signal (do not reward hacking, cheating, or deception); reduce reliance on generalisation, by specifying the target behaviour during training or by shaping generalisation (selective generalisation, character training); and get humans to stop training models to be malevolent or selfish. What we do not know is how to do this while staying competitive, and misalignment science helps much less with that second problem.
On implicit work trials, the relevant counterfactual may be capabilities, not misalignment science. Much of the community is not made up of impartial altruists, and the trend looks like it is going the wrong way (less EA over time, less veganism), which is what I would expect from any community growing this fast. Discouraging alignment engineering may not redirect that subset to misalignment science. It may redirect them to capabilities, where the same skills pay better.
I know much less about personas than you do, so I’d be very glad to be proven wrong here, but I feel like there’s a lot that remains to be understood. For example: What counts as one correlated distribution? Is it text produced by one actor (e.g., a single person), a cluster of similar actors (e.g., past AI models or the rationalist community), or something else? How entangled are those distributions with each other? If we post-train the LLM to take on a persona that isn’t represented in the pretraining data, is it going to stitch together features of many different personas, simulate the most similar persona that was represented in pretraining data and learn features that override the behavior of that persona in appropriate ways, or something else? Why does misalignment conditionalize but capabilities don’t? What are the most important factors that determine whether an aligned persona persists or doesn’t persist through conflicting RL pressures? A deep science in the sense Richard Ngo uses the word would probably attempt to tackle more fundamental questions than the ones above, but even the shallower science of personas doesn’t look nearly complete to me.
I think it’s true that a lot of past work foreshadowed split personas, but also, I think it was completely reasonable for people to assume after the Natural Emergent Misalignment paper came out that reward hacking propensities would continue to generalize into a globally misaligned persona instead of conditionalizing into a split-brained grader sycophant. We got evidence that there’s often no emergent misalignment under different training setups a few months later, but a robust theory of personas would have allowed us to predict this right when Anthropic’s paper came out, and I don’t think we had that theory.