I vaguely thought the argument is something to do with not having some sort of “ground truth” feedback mechanism (i.e. the human reward function / steering subsystem). Like to do cev / reflective equilibrium type stuff, you need to be able to query some base set of intuitions.
This is the vibe I got from the Putin > claude post, but idk..
I vaguely thought the argument is something to do with not having some sort of “ground truth” feedback mechanism (i.e. the human reward function / steering subsystem). Like to do cev / reflective equilibrium type stuff, you need to be able to query some base set of intuitions.
This is the vibe I got from the Putin > claude post, but idk..