If you use the Anti-Nirvana trick, your agent just goes “nothing matters at all, the foe will mispredict and I’ll get -infinity reward” and rolls over and cries since all policies are optimal. Don’t do that one, it’s a bad idea.
Sorry, I meant the combination of best-case reasoning (sup instead of inf) and the anti-Nirvana trick. In that case the agent goes “Murphy won’t mispredict, since then I’d get -infinity reward which can’t be the best that I do”.
For your concrete example, that’s why you have multiple hypotheses that are learnable.
Hmm, that makes sense, I think? Perhaps I just haven’t really internalized the learning aspect of all of this.
Sorry, I meant the combination of best-case reasoning (sup instead of inf) and the anti-Nirvana trick. In that case the agent goes “Murphy won’t mispredict, since then I’d get -infinity reward which can’t be the best that I do”.
Hmm, that makes sense, I think? Perhaps I just haven’t really internalized the learning aspect of all of this.