Thanks for this question! This is an important and very relevant question.
Suppose we only label “acceptable action” but not “action of a virtuous character.”
We need to know which action is acceptable or not for each situation, for potentially infinite actions, which requires some theoretical considerations. This is therefore done in practice by giving a lot of rules today. This is what we have been claiming to be hopeless. On the other hand, we do have the concept of a virtuous person, which is also shared by LLMs today, from which we can derive our evaluations of actions in various (even novel) situations. Thus, our approach is much more “efficient” than giving various rules (which almost certainly conflict with each other).
Thanks for this question! This is an important and very relevant question.
Suppose we only label “acceptable action” but not “action of a virtuous character.”
We need to know which action is acceptable or not for each situation, for potentially infinite actions, which requires some theoretical considerations. This is therefore done in practice by giving a lot of rules today. This is what we have been claiming to be hopeless. On the other hand, we do have the concept of a virtuous person, which is also shared by LLMs today, from which we can derive our evaluations of actions in various (even novel) situations. Thus, our approach is much more “efficient” than giving various rules (which almost certainly conflict with each other).