There seems to be a correlation between how we approach the flaws of argmaxing in AI safety arguments and their failure modes in game design; something about exploiting-the-larger-thing crowding-out the subtler thing in few-variant games and scoring systems.
There seems to be a correlation between how we approach the flaws of argmaxing in AI safety arguments and their failure modes in game design; something about exploiting-the-larger-thing crowding-out the subtler thing in few-variant games and scoring systems.