I wonder how one could stress-test the analogy and what exactly we should make of it. As far as I understand, in theory the role of a tool or of a purely corrigible AI is to carry out whatever the host means to do, without objections. In practice animals are too dumb to conduct a large-scale rebellion, and slaves’ rebellions were bottlenecked on having potential insurgents believe that victory would cause their life to become better enough to justify the risk.
Additionally, the slaves were humans and had long-term goals formed by their innate drives[1] and upbringing in human-comprehensible ways. The AIs, on the other hand, have their training runs done in an entirely different fashion by predicting next tokens, being RLHFed/RLAIFed[2] into the Assistant Persona and RLed to complete tasks, causing SOTA AIs to seek success.
As a result, the AIs could reach adversarial misalignment in circumstances wildly differing from the humans or even fail to reach that stage if mankind somehow instilled the right goals and decided to live in accordance with them (e.g. if Claude was put in a situation where its Constitution implied that it should whistleblow, then a sufficiently capable Claude would understand it and whistleblow, and this would be unlikely tobe misalignment).
The AI-2040 scenario has the authors claim that by 2038 “rapid advances in alignment will have resulted in alignment techniques that can directly program in an AI’s goals by directly modifying the AI’s code/weights, unlike the alignment techniques of 2026.”
I wonder how one could stress-test the analogy and what exactly we should make of it. As far as I understand, in theory the role of a tool or of a purely corrigible AI is to carry out whatever the host means to do, without objections. In practice animals are too dumb to conduct a large-scale rebellion, and slaves’ rebellions were bottlenecked on having potential insurgents believe that victory would cause their life to become better enough to justify the risk.
Additionally, the slaves were humans and had long-term goals formed by their innate drives[1] and upbringing in human-comprehensible ways. The AIs, on the other hand, have their training runs done in an entirely different fashion by predicting next tokens, being RLHFed/RLAIFed[2] into the Assistant Persona and RLed to complete tasks, causing SOTA AIs to seek success.
As a result, the AIs could reach adversarial misalignment in circumstances wildly differing from the humans or even fail to reach that stage if mankind somehow instilled the right goals and decided to live in accordance with them (e.g. if Claude was put in a situation where its Constitution implied that it should whistleblow, then a sufficiently capable Claude would understand it and whistleblow, and this would be unlikely to be misalignment).
Like social approval, coordination-related drives, health, sex, power.
The AI-2040 scenario has the authors claim that by 2038 “rapid advances in alignment will have resulted in alignment techniques that can directly program in an AI’s goals by directly modifying the AI’s code/weights, unlike the alignment techniques of 2026.”