Cool work! One follow-up idea that comes to my mind is to test whether it is possible to instill asymmetric entanglement. For example, it would be useful to have a model that generalizes (reward hacking → benevolent persona) but not (benevolent persona → reward hacking).
Cool work! One follow-up idea that comes to my mind is to test whether it is possible to instill asymmetric entanglement. For example, it would be useful to have a model that generalizes (reward hacking → benevolent persona) but not (benevolent persona → reward hacking).