I think in as much as someone is serious about AI alignment with prosaic methods, corrigibility is a much better target than value alignment. When people say they are planning to build superhuman value aligned systems in a single shot (without routing through highly corrigible systems first), that strongly to me indicates to me they do not understand the difficulty of this problem, or the stakes.
That said, almost all prosaic research is failing at a much earlier stage, where people just do another flavor of RL, optimize on their own approval, and then act surprised every time the model looks superficially aligned and then turns out to be doing crazy things below the hood. That of course doesn’t work for either value alignment or corrigibility or really anything you want to point AI systems at.
I think in as much as someone is serious about AI alignment with prosaic methods, corrigibility is a much better target than value alignment. When people say they are planning to build superhuman value aligned systems in a single shot (without routing through highly corrigible systems first), that strongly to me indicates to me they do not understand the difficulty of this problem, or the stakes.
That said, almost all prosaic research is failing at a much earlier stage, where people just do another flavor of RL, optimize on their own approval, and then act surprised every time the model looks superficially aligned and then turns out to be doing crazy things below the hood. That of course doesn’t work for either value alignment or corrigibility or really anything you want to point AI systems at.