You either deal with an agent, and then the easiest things around to imitation learn are humans. Those do have personas. Maybe you need to shift this into non-human-like reasoning mode? E.g. some kind of neuralise or constructed language? But that sounds difficult for alignment, and it might still get seeded with human imitation, just non transparently. And all the problems with neuralise.
Or maybe you need more power-armor design? E.g. edit prediction. This also might give rise to an agent in background. And be less powerful in the first place.
Something other?
And to be fair all of this sounds like a pretty high capability externality line of thinking.
I’m unsure what you mean. I consider the pretraining data to be the issue, not the particular language or output that the AI generates. If you pretrain on moral humans then convert your model to a neuralese model, then the model still has representations of good human personas.
I do agree that Personaless Alignment is a dual-use research direction in that it may benefit capabilities (you’re doing crazy things to a misaligned AI to try to get it to be good), but I consider its capability risks likely smaller than the benefit to alignment. Once the details of the experimental approaches are pinned down, it would be easier to say.
I’m not planning to pursue Personaless Alignment at the moment, since I am about to do an AI Control research fellowship with Redwood.
I don’t think the right path is to try starting from Zero pretraining. Pretraining is and will continue to be a huge benefit to capabilities, and aligning the most capable models is the goal. I’m saying that we should separate pretraining for capabilities from the alignment that can come from eliciting aligned personas from pretraining. I’m not sure exactly how to do this, but it would likely involve starting with carefully filtered large models.
Okay, let’s try to classify non apples here.
You either deal with an agent, and then the easiest things around to imitation learn are humans. Those do have personas. Maybe you need to shift this into non-human-like reasoning mode? E.g. some kind of neuralise or constructed language? But that sounds difficult for alignment, and it might still get seeded with human imitation, just non transparently. And all the problems with neuralise.
Or maybe you need more power-armor design? E.g. edit prediction. This also might give rise to an agent in background. And be less powerful in the first place.
Something other?
And to be fair all of this sounds like a pretty high capability externality line of thinking.
I’m unsure what you mean. I consider the pretraining data to be the issue, not the particular language or output that the AI generates. If you pretrain on moral humans then convert your model to a neuralese model, then the model still has representations of good human personas.
I do agree that Personaless Alignment is a dual-use research direction in that it may benefit capabilities (you’re doing crazy things to a misaligned AI to try to get it to be good), but I consider its capability risks likely smaller than the benefit to alignment. Once the details of the experimental approaches are pinned down, it would be easier to say.
So, you plan to experiment with MuZero type stuff, where you train an agent with no human imitation learning whatsoever?
I’m not planning to pursue Personaless Alignment at the moment, since I am about to do an AI Control research fellowship with Redwood.
I don’t think the right path is to try starting from Zero pretraining. Pretraining is and will continue to be a huge benefit to capabilities, and aligning the most capable models is the goal. I’m saying that we should separate pretraining for capabilities from the alignment that can come from eliciting aligned personas from pretraining. I’m not sure exactly how to do this, but it would likely involve starting with carefully filtered large models.