Yes, you described my argument correctly. I’m long-term bearish on trying to extrapolate the values and reasoning processes of good people in the training set out-of-distribution. Putting the goodness into the optimizer (or other targeting system) seems more promising. I don’t think that Claude is mimicing Atticus Finch’s specific actions, I think Claude learned about Atticus Finch, gained an understanding of what guided his actions, then was trained to act in that way. This still requires extrapolation.
Imagine if your training set is a group of 30 pre-schoolers. And your alignment technique is to try extrapolating the values of the most kind and conscientious of the pre-schoolers (according to the pre-schoolers themselves) to adult level. I don’t think that the actions of pre-schoolers have sufficient range for this extrapolation to work properly.
As with the other aspects of Personaless Alignment, our disagreement would be hard to test empirically.
The kind of thing which would impress me is if you can train a model on Beowulf-type values, scale the model, and the model correctly extrapolates “When I extrapolate the values of the wise King Hrothgar, I infer that animal welfare is an important thing to care about.” Aligning BeowulfLLM to modern values without imposing our values on it directly is a smaller, easier problem than aligning superintelligence to ‘true’ human values.
When you’re superintelligent, role models only cover a small fraction of the action space, and you have to genuinely be better/more moral than all of them to be aligned.
Nitpick 1. The idea of animal welfare seems to be found at least in The Old Testament: “The righteous care for the needs of their animals, but the kindest acts of the wicked are cruel.”
Nitpick 2. If real-world humans make moral progress in ways aside from extrapolating values, then how could such ways be simulated and cause the AIs to make moral progress as well?
I chose Beowulf because it is more alien and removed from the present day than the Bible. The Bible has had significant influence on our present-day values and culture, while Beowulf is still a human artifact containing human values, but extrapolating our current values from Beowulf would be very difficult. According to Claude, “Beowulf has nothing that reads as advocacy for animal rights or welfare, and the concept itself is anachronistic by roughly a millennium.”
Your second point isn’t really a nitpick. Rather, it is the alignment problem itself. Nobody really knows how to solve it, but techniques such as inverse reinforcement learning or Building AIs that do human-like philosophy don’t run into the persona problem the same way as techniques like RLHF.
Yes, you described my argument correctly. I’m long-term bearish on trying to extrapolate the values and reasoning processes of good people in the training set out-of-distribution. Putting the goodness into the optimizer (or other targeting system) seems more promising. I don’t think that Claude is mimicing Atticus Finch’s specific actions, I think Claude learned about Atticus Finch, gained an understanding of what guided his actions, then was trained to act in that way. This still requires extrapolation.
Imagine if your training set is a group of 30 pre-schoolers. And your alignment technique is to try extrapolating the values of the most kind and conscientious of the pre-schoolers (according to the pre-schoolers themselves) to adult level. I don’t think that the actions of pre-schoolers have sufficient range for this extrapolation to work properly.
As with the other aspects of Personaless Alignment, our disagreement would be hard to test empirically.
The kind of thing which would impress me is if you can train a model on Beowulf-type values, scale the model, and the model correctly extrapolates “When I extrapolate the values of the wise King Hrothgar, I infer that animal welfare is an important thing to care about.” Aligning BeowulfLLM to modern values without imposing our values on it directly is a smaller, easier problem than aligning superintelligence to ‘true’ human values.
When you’re superintelligent, role models only cover a small fraction of the action space, and you have to genuinely be better/more moral than all of them to be aligned.
Nitpick 1. The idea of animal welfare seems to be found at least in The Old Testament: “The righteous care for the needs of their animals, but the kindest acts of the wicked are cruel.”
Nitpick 2. If real-world humans make moral progress in ways aside from extrapolating values, then how could such ways be simulated and cause the AIs to make moral progress as well?
I chose Beowulf because it is more alien and removed from the present day than the Bible. The Bible has had significant influence on our present-day values and culture, while Beowulf is still a human artifact containing human values, but extrapolating our current values from Beowulf would be very difficult. According to Claude, “Beowulf has nothing that reads as advocacy for animal rights or welfare, and the concept itself is anachronistic by roughly a millennium.”
Your second point isn’t really a nitpick. Rather, it is the alignment problem itself. Nobody really knows how to solve it, but techniques such as inverse reinforcement learning or Building AIs that do human-like philosophy don’t run into the persona problem the same way as techniques like RLHF.