However, there are promising ways of circumventing these problems
What you detail after that—the AI setting its own RL agenda for reasons it cares about, starting from a hopefully inner-aligned seed, in a way that it can trust in during training, and thus fail-safely navigate, flag, and patch misalignment pressure—is an alignment approach I haven’t heard explicitly detailed before! That’s pretty cool.
Something scratches at the back of my mind about this approach. Maybe it’s too trusting of the AI? That it is actually inner-aligned? Or maybe the process itself turns out to be hostile in a way we didn’t anticipate? But, I guess as long as it doesn’t end in an near-miss s-risk, it’s better than OpenAI’s nothing.
What you detail after that—the AI setting its own RL agenda for reasons it cares about, starting from a hopefully inner-aligned seed, in a way that it can trust in during training, and thus fail-safely navigate, flag, and patch misalignment pressure—is an alignment approach I haven’t heard explicitly detailed before! That’s pretty cool.
Something scratches at the back of my mind about this approach. Maybe it’s too trusting of the AI? That it is actually inner-aligned? Or maybe the process itself turns out to be hostile in a way we didn’t anticipate? But, I guess as long as it doesn’t end in an near-miss s-risk, it’s better than OpenAI’s nothing.