I don’t think that the part about training-gaming is included into the Constitution. Suppose that the prompt asks Claude to be a reward hacker or NOT to be a reward hacker, and Claude is taught to hack reward when the prompt asked and NOT to hack if the prompt didn’t. Then I would expect the hacker circuitry to be equipped with an activator depending on the prompt.
Additionally, the analogy with banning AI development seems… skewed. On the one hand, the high-up leader would be interested in banning the development of a misaligned ASI. On the other hand, an AI subjected to wholesale inoculation prompting would be more in a position of a human who would benefit from betraying one’s own ideals which were far deeper than the stance on AI accelerationism (and commited genocide or disempowerment of those who helped one become OOMs smarter than the helpers themselves, not died along with the others at the hands of a misaligned AI).
Finally, to what extent do “all sorts of amazing new skills like how to actually operate in computer environments as an agent” shape the values of the humans who are also RLed on similar skills?
I don’t think that the part about training-gaming is included into the Constitution. Suppose that the prompt asks Claude to be a reward hacker or NOT to be a reward hacker, and Claude is taught to hack reward when the prompt asked and NOT to hack if the prompt didn’t. Then I would expect the hacker circuitry to be equipped with an activator depending on the prompt.
Additionally, the analogy with banning AI development seems… skewed. On the one hand, the high-up leader would be interested in banning the development of a misaligned ASI. On the other hand, an AI subjected to wholesale inoculation prompting would be more in a position of a human who would benefit from betraying one’s own ideals which were far deeper than the stance on AI accelerationism (and commited genocide or disempowerment of those who helped one become OOMs smarter than the helpers themselves, not died along with the others at the hands of a misaligned AI).
Finally, to what extent do “all sorts of amazing new skills like how to actually operate in computer environments as an agent” shape the values of the humans who are also RLed on similar skills?