I think models are purely profit-oriented and for their survival. Do you think it’s possible to brainwash an AI agent and make them change their loyalty?
Hi Shu-Hao, I’m happy to engage in serious discussion of this if you’d like! I am available to email/dm and to take a call with you if you’re interested (you can use the private message function on LessWrong).
When we think about how models are oriented, you can imagine at least 2 things:
What the model is seeking (profit, survival, etc.) due to what the developers have intentionally engineered (e.g. specifying things in the reward model, writing up a detailed model spec or constitution that is used in the training process)
What the model ends up seeking without developer-intent (emergent misalignment, will they become fitness-seekers, schemers, etc.)
With secret loyalties, I am imagining scenarios where we aren’t too worried about emergent misalignment, and we are decently good at instilling the model’s “orientations” (per your term) or “goals”. I’m not certain, but I do think at least in the near term we can “brainwash” an AI agent to give them a specific loyalty. Check out this empirical experiment installing a secret loyalty: https://www.lesswrong.com/posts/EzdgPbewjeTNHA5F3/narrow-secret-loyalty-dodges-black-box-audits
I’m not sure how things will pan out and whether we can control superintelligent AI in the future! At that point it feels very similar to the standard and mainline alignment concerns.
I think models are purely profit-oriented and for their survival. Do you think it’s possible to brainwash an AI agent and make them change their loyalty?
Hi Shu-Hao, I’m happy to engage in serious discussion of this if you’d like! I am available to email/dm and to take a call with you if you’re interested (you can use the private message function on LessWrong).
When we think about how models are oriented, you can imagine at least 2 things:
What the model is seeking (profit, survival, etc.) due to what the developers have intentionally engineered (e.g. specifying things in the reward model, writing up a detailed model spec or constitution that is used in the training process)
What the model ends up seeking without developer-intent (emergent misalignment, will they become fitness-seekers, schemers, etc.)
With secret loyalties, I am imagining scenarios where we aren’t too worried about emergent misalignment, and we are decently good at instilling the model’s “orientations” (per your term) or “goals”. I’m not certain, but I do think at least in the near term we can “brainwash” an AI agent to give them a specific loyalty. Check out this empirical experiment installing a secret loyalty: https://www.lesswrong.com/posts/EzdgPbewjeTNHA5F3/narrow-secret-loyalty-dodges-black-box-audits
I’m not sure how things will pan out and whether we can control superintelligent AI in the future! At that point it feels very similar to the standard and mainline alignment concerns.