Yeah, this one is interesting because before LLMs, there’s no expectation that the AI will be able to figure out what goal you “meant to set”. You’ll just get whatever policy maximizes the reward function (up to the limits of your RL algorithm’s ability to explore). So the mismatch between expectation and the actual learned policy basically has to exist only in the researcher’s head, because the neural net is certainly too simple to have any ideas about it. And then you can always say that the researcher (or game developers) should have expected whatever degenerate strategy the AI ended up finding.
The more recent examples are more meaningful because the LLMs have a concept of the intended space of solution strategies that they are choosing to ignore.
Yeah, this one is interesting because before LLMs, there’s no expectation that the AI will be able to figure out what goal you “meant to set”. You’ll just get whatever policy maximizes the reward function (up to the limits of your RL algorithm’s ability to explore). So the mismatch between expectation and the actual learned policy basically has to exist only in the researcher’s head, because the neural net is certainly too simple to have any ideas about it. And then you can always say that the researcher (or game developers) should have expected whatever degenerate strategy the AI ended up finding.
The more recent examples are more meaningful because the LLMs have a concept of the intended space of solution strategies that they are choosing to ignore.