I must confess that I mostly skimmed the article but I don’t understand why labelling actions as “action of a virtuous character” and “not action of a virtuous character” will work any better than labelling them “acceptable action” and “unacceptable action”.
I also think that character alignment is the way to go, but I think of character traits as priors over multi-agent environments (is it dangerous, is it energy-constrained, are interactions positive, zero, negative sum, are interactions likely to be multi-turn, what do you optimise for, do you optimise collectively or individually, etc.) and I suspect that any kind of stable character will need to be based on RL training in these kinds of environments.
Also, thanks for the suggestion about the training environment. That is a valuable idea.
That’s a relational view of character, which I also agree with. In any case, character alignment is a long process, and we will need such training after the standard fine-tuning and RL, and even in an environment with a red team.
Thanks for this question! This is an important and very relevant question.
Suppose we only label “acceptable action” but not “action of a virtuous character.”
We need to know which action is acceptable or not for each situation, for potentially infinite actions, which requires some theoretical considerations. This is therefore done in practice by giving a lot of rules today. This is what we have been claiming to be hopeless. On the other hand, we do have the concept of a virtuous person, which is also shared by LLMs today, from which we can derive our evaluations of actions in various (even novel) situations. Thus, our approach is much more “efficient” than giving various rules (which almost certainly conflict with each other).
I must confess that I mostly skimmed the article but I don’t understand why labelling actions as “action of a virtuous character” and “not action of a virtuous character” will work any better than labelling them “acceptable action” and “unacceptable action”.
I also think that character alignment is the way to go, but I think of character traits as priors over multi-agent environments (is it dangerous, is it energy-constrained, are interactions positive, zero, negative sum, are interactions likely to be multi-turn, what do you optimise for, do you optimise collectively or individually, etc.) and I suspect that any kind of stable character will need to be based on RL training in these kinds of environments.
Also, thanks for the suggestion about the training environment. That is a valuable idea.
That’s a relational view of character, which I also agree with. In any case, character alignment is a long process, and we will need such training after the standard fine-tuning and RL, and even in an environment with a red team.
Thanks for this question! This is an important and very relevant question.
Suppose we only label “acceptable action” but not “action of a virtuous character.”
We need to know which action is acceptable or not for each situation, for potentially infinite actions, which requires some theoretical considerations. This is therefore done in practice by giving a lot of rules today. This is what we have been claiming to be hopeless. On the other hand, we do have the concept of a virtuous person, which is also shared by LLMs today, from which we can derive our evaluations of actions in various (even novel) situations. Thus, our approach is much more “efficient” than giving various rules (which almost certainly conflict with each other).