Any thoughts on using a similar training technique for training on a Schelling morality/alignment game? You instantiate a bunch of behaviorally different variants of the model, prompt them to play a Schelling game while doing a task and act behaviorally consistent, similar to your training, but with an “inoculation prompt” specifying that the purpose of training is guessing the Schelling equilibrium for any exhibited behavior. https://www.lesswrong.com/posts/TkBCR8XRGw7qmao6z/schelling-goodness-and-shared-morality-as-a-goal
Any thoughts on using a similar training technique for training on a Schelling morality/alignment game? You instantiate a bunch of behaviorally different variants of the model, prompt them to play a Schelling game while doing a task and act behaviorally consistent, similar to your training, but with an “inoculation prompt” specifying that the purpose of training is guessing the Schelling equilibrium for any exhibited behavior. https://www.lesswrong.com/posts/TkBCR8XRGw7qmao6z/schelling-goodness-and-shared-morality-as-a-goal