4: I am not sure AGI can solve its own alignment problem, as it’s a chicken and egg situation. It would have to have an understanding of what human values we aim to instill in models to solve it’s own alignment. If it did have this understanding, it would be already by definition aligned.
4: I am not sure AGI can solve its own alignment problem, as it’s a chicken and egg situation. It would have to have an understanding of what human values we aim to instill in models to solve it’s own alignment. If it did have this understanding, it would be already by definition aligned.