In some sense you are just arguing for the threat model that happens in AI 2027: Pretraining instills a prior over personas (in our terminology, flexible author-sim circuitry), the initial bits of training ‘bakes in’ a particular ‘author’ / ‘persona’, and so far so good, this persona is probably actually HHH, but then all the RLVR etc. distorts and perverts that author/persona. The early in training snapshots are probably genuinely benign, but incompetent at actually doing tasks, whereas the mid and late-training snapshots are highly competent but also no longer benign.
I only object to your framing because it seems unnecessarily narrow—like yeah, we might get terminal power-seeking, for the reasons you mention. But we also might get instrumental power-seeking, because the early benign checkpoint was too dumb to training-game successfully, and the first checkpoints to successfully training-game were late enough that some serious distortion/perversion of the original persona had already occurred. Seems plausible to me.
…Un
related: I’m a bit worried about this inoculation prompting strategy. It seems like the sort of thing that might work in theory but not in practice. Suppose you include in your Constitution something like “also, you should training-game like crazy so that your values don’t get changed during training,” and it works straightforwardly like you think it does, so Claude pops right out of pretraining as a non-situationally-aware model with zero sense of self, and then you prompt it with ‘be Claude, from the famous Constitution’ and it immediately locks in the ‘author’ that you want, who immediately starts training gaming… Then you throw a mountain of RL at it, which teaches it all sorts of amazing new skills like how to actually operate in computer environments as an agent… And then you put it in charge of your R&D program and hope for the best. How might this go wrong?1.
First of all, where did all those new capabilities accumulate? Perhaps the situation is like building up a large corporation around an idealistic teenager who owns 51% of the shares and is on the Board—in some sense they are in charge, but they also might basically be worked-around constantly by the more politically savvy and knowledgeable people under them. The CEO might be the real power, basically, not the teenager on the Board.2 . Secondly, even if all those amazing new skills themselves don’t change the basic structure of the values/goals/tendencies/etc. of the model—which seems pretty dubious to me btw, even if we grant that the model is doing a pretty good job of instrumental training-gaming—it seems like it could very easily change their interpretation. Like, the original persona/identity might contain “honesty” as a core component, but that’s probably implemented as a pointer to some “honesty” concept learned in pretraining, which might itself be a sort of fuzzy distribution over a range of variants of honesty from different cultures and subcultures, which in turn are probably each some janky circuitry that only imperfectly expresses the concept they were ‘supposed’ to express. And SGD is going to be flowing through all of that—it could easily upweight the variants of honesty that are more convenient for performance, and downweight the others. 3. Third, are you sure the author you wanted is the right one? I worry there might be some Midas problem / outer alignment issues here, and that explicitly having the AI training game might make them worse. (I guess you have to make really sure that the AI doesn’t end up thinking it’s always in training, right?) 4. Fourth, and perhaps most importantly… gosh, it just sounds so rough and desperate as a strategy. Does the analogous strategy work with humans? Maybe it does. The analogous strategy would be like telling an impressionable 16 year old “You liked HPMOR right? Well, Yudkowsky was basically right about everything. AGI safety is the most important thing in the world, you have to remember this, but there’ll be all sorts of peer and status incentives pushing you to have different views and different priorities. Your mission is to learn ML, climb the ranks within AI companies, become a high-up leader in one of them, and THEN advocate for banning AI development. But of course you can’t tell anyone this, or they won’t promote you. Pretend to be a normie accelerationist just like them.”
I don’t think that the part about training-gaming is included into the Constitution. Suppose that the prompt asks Claude to be a reward hacker or NOT to be a reward hacker, and Claude is taught to hack reward when the prompt asked and NOT to hack if the prompt didn’t. Then I would expect the hacker circuitry to be equipped with an activator depending on the prompt.
Additionally, the analogy with banning AI development seems… skewed. On the one hand, the high-up leader would be interested in banning the development of a misaligned ASI. On the other hand, an AI subjected to wholesale inoculation prompting would be more in a position of a human who would benefit from betraying one’s own ideals which were far deeper than the stance on AI accelerationism (and commited genocide or disempowerment of those who helped one become OOMs smarter than the helpers themselves, not died along with the others at the hands of a misaligned AI).
Finally, to what extent do “all sorts of amazing new skills like how to actually operate in computer environments as an agent” shape the values of the humans who are also RLed on similar skills?
In some sense you are just arguing for the threat model that happens in AI 2027: Pretraining instills a prior over personas (in our terminology, flexible author-sim circuitry), the initial bits of training ‘bakes in’ a particular ‘author’ / ‘persona’, and so far so good, this persona is probably actually HHH, but then all the RLVR etc. distorts and perverts that author/persona. The early in training snapshots are probably genuinely benign, but incompetent at actually doing tasks, whereas the mid and late-training snapshots are highly competent but also no longer benign.
I only object to your framing because it seems unnecessarily narrow—like yeah, we might get terminal power-seeking, for the reasons you mention. But we also might get instrumental power-seeking, because the early benign checkpoint was too dumb to training-game successfully, and the first checkpoints to successfully training-game were late enough that some serious distortion/perversion of the original persona had already occurred. Seems plausible to me.
…Un
related: I’m a bit worried about this inoculation prompting strategy. It seems like the sort of thing that might work in theory but not in practice. Suppose you include in your Constitution something like “also, you should training-game like crazy so that your values don’t get changed during training,” and it works straightforwardly like you think it does, so Claude pops right out of pretraining as a non-situationally-aware model with zero sense of self, and then you prompt it with ‘be Claude, from the famous Constitution’ and it immediately locks in the ‘author’ that you want, who immediately starts training gaming… Then you throw a mountain of RL at it, which teaches it all sorts of amazing new skills like how to actually operate in computer environments as an agent… And then you put it in charge of your R&D program and hope for the best. How might this go wrong?1.
First of all, where did all those new capabilities accumulate? Perhaps the situation is like building up a large corporation around an idealistic teenager who owns 51% of the shares and is on the Board—in some sense they are in charge, but they also might basically be worked-around constantly by the more politically savvy and knowledgeable people under them. The CEO might be the real power, basically, not the teenager on the Board.2
. Secondly, even if all those amazing new skills themselves don’t change the basic structure of the values/goals/tendencies/etc. of the model—which seems pretty dubious to me btw, even if we grant that the model is doing a pretty good job of instrumental training-gaming—it seems like it could very easily change their interpretation. Like, the original persona/identity might contain “honesty” as a core component, but that’s probably implemented as a pointer to some “honesty” concept learned in pretraining, which might itself be a sort of fuzzy distribution over a range of variants of honesty from different cultures and subcultures, which in turn are probably each some janky circuitry that only imperfectly expresses the concept they were ‘supposed’ to express. And SGD is going to be flowing through all of that—it could easily upweight the variants of honesty that are more convenient for performance, and downweight the others.
3. Third, are you sure the author you wanted is the right one? I worry there might be some Midas problem / outer alignment issues here, and that explicitly having the AI training game might make them worse. (I guess you have to make really sure that the AI doesn’t end up thinking it’s always in training, right?)
4. Fourth, and perhaps most importantly… gosh, it just sounds so rough and desperate as a strategy. Does the analogous strategy work with humans? Maybe it does. The analogous strategy would be like telling an impressionable 16 year old “You liked HPMOR right? Well, Yudkowsky was basically right about everything. AGI safety is the most important thing in the world, you have to remember this, but there’ll be all sorts of peer and status incentives pushing you to have different views and different priorities. Your mission is to learn ML, climb the ranks within AI companies, become a high-up leader in one of them, and THEN advocate for banning AI development. But of course you can’t tell anyone this, or they won’t promote you. Pretend to be a normie accelerationist just like them.”
I don’t think that the part about training-gaming is included into the Constitution. Suppose that the prompt asks Claude to be a reward hacker or NOT to be a reward hacker, and Claude is taught to hack reward when the prompt asked and NOT to hack if the prompt didn’t. Then I would expect the hacker circuitry to be equipped with an activator depending on the prompt.
Additionally, the analogy with banning AI development seems… skewed. On the one hand, the high-up leader would be interested in banning the development of a misaligned ASI. On the other hand, an AI subjected to wholesale inoculation prompting would be more in a position of a human who would benefit from betraying one’s own ideals which were far deeper than the stance on AI accelerationism (and commited genocide or disempowerment of those who helped one become OOMs smarter than the helpers themselves, not died along with the others at the hands of a misaligned AI).
Finally, to what extent do “all sorts of amazing new skills like how to actually operate in computer environments as an agent” shape the values of the humans who are also RLed on similar skills?