I’d love to believe that this supports various stories about how AIs can be sycophantic (and have other more dangerous behaviors) by shifting their interpretations of words, so that their conception of the AI self-character continues to use words humans would call good even while they do bad stuff.
I’d also love to believe that this method tells us accurately what the AI’s actual change in the conception of the self-character is when it’s doing different sorts of behaviors.
But given the messiness, I kind of feel like the biggest updates I should do are about the amount of off-target semantics in steering vectors, and the difficulty of training neologisms to be “described faithfully” (whatever that means). You have some convincing 1d projections, but how weird is your high-dimensional data?
I’d love to believe that this supports various stories about how AIs can be sycophantic (and have other more dangerous behaviors) by shifting their interpretations of words, so that their conception of the AI self-character continues to use words humans would call good even while they do bad stuff.
I’d also love to believe that this method tells us accurately what the AI’s actual change in the conception of the self-character is when it’s doing different sorts of behaviors.
But given the messiness, I kind of feel like the biggest updates I should do are about the amount of off-target semantics in steering vectors, and the difficulty of training neologisms to be “described faithfully” (whatever that means). You have some convincing 1d projections, but how weird is your high-dimensional data?