reward: Being a great country with high quality of life for it’s ruling elites
feels like this is more of a “goal” than a reward, right? Maybe semantically equivalent, but probably not, and matters when distinguishing between alignment failures and specific things that happen during RL.
If we consider classical reward hacking cases (like say RL CoastRunners example), there seems to be distinctions. I can look at the video of boat spinning in circles and be like “wow, it gets a really high score, also that’s definitely a hack or not the intended way to get that high score” in a way I can’t in this Worlds Fare example. Sure we can look at press being impressed by those pavilions, but was that hack, and can I point to reward function that got pushed really high? Sort of, but less so.
It seems useful to keep “reward hacking” distinct for those special cases like the CoastRunners example, without expanding it to any kind of alignment failure. I’m arguing this a lot on vibes. Making a formalism of what is reward hacking vs other alignment failures seems maybe doable (and perhaps people have), but I did not find something I liked in a quick search, and I don’t attempt with formalisms here.
After quick background search though, I will link to this recent post “Confusion around the term reward hacking” (Azarbal, 2026). There might be field-wide muddling.
feels like this is more of a “goal” than a reward, right? Maybe semantically equivalent, but probably not, and matters when distinguishing between alignment failures and specific things that happen during RL.
If we consider classical reward hacking cases (like say RL CoastRunners example), there seems to be distinctions. I can look at the video of boat spinning in circles and be like “wow, it gets a really high score, also that’s definitely a hack or not the intended way to get that high score” in a way I can’t in this Worlds Fare example. Sure we can look at press being impressed by those pavilions, but was that hack, and can I point to reward function that got pushed really high? Sort of, but less so.
It seems useful to keep “reward hacking” distinct for those special cases like the CoastRunners example, without expanding it to any kind of alignment failure. I’m arguing this a lot on vibes. Making a formalism of what is reward hacking vs other alignment failures seems maybe doable (and perhaps people have), but I did not find something I liked in a quick search, and I don’t attempt with formalisms here.
After quick background search though, I will link to this recent post “Confusion around the term reward hacking” (Azarbal, 2026). There might be field-wide muddling.