I don’t disagree with the proposal being useful and likely already in use. I disagree that it can replace RL or is even in the same ballpark of sample efficiency. You said before that now anyone can train; disagree, because RL can train on novel problems and find ways to solve them, OPSD can’t if I understand correctly.
I don’t disagree with the proposal being useful and likely already in use. I disagree that it can replace RL or is even in the same ballpark of sample efficiency. You said before that now anyone can train; disagree, because RL can train on novel problems and find ways to solve them, OPSD can’t if I understand correctly.