Mateusz Bagiński comments on Reward is not the optimization target