This was a helpful way to understand the differences (especially requiring authentication or knowing it’s really the principal is something I haven’t thought about much).
There’s an odd example where you might have someone like a CEO strike a deal with a misaligned/scheming model such that for some period of time/condition, the model is effectively secretly loyal despite it not having some weight-encoded “goal” or propensity to favor the principal.
Hi Shu-Hao, I’m happy to engage in serious discussion of this if you’d like! I am available to email/dm and to take a call with you if you’re interested (you can use the private message function on LessWrong).
When we think about how models are oriented, you can imagine at least 2 things:
What the model is seeking (profit, survival, etc.) due to what the developers have intentionally engineered (e.g. specifying things in the reward model, writing up a detailed model spec or constitution that is used in the training process)
What the model ends up seeking without developer-intent (emergent misalignment, will they become fitness-seekers, schemers, etc.)
With secret loyalties, I am imagining scenarios where we aren’t too worried about emergent misalignment, and we are decently good at instilling the model’s “orientations” (per your term) or “goals”. I’m not certain, but I do think at least in the near term we can “brainwash” an AI agent to give them a specific loyalty. Check out this empirical experiment installing a secret loyalty: https://www.lesswrong.com/posts/EzdgPbewjeTNHA5F3/narrow-secret-loyalty-dodges-black-box-audits
I’m not sure how things will pan out and whether we can control superintelligent AI in the future! At that point it feels very similar to the standard and mainline alignment concerns.