I’m not really focused on “fair”—just understanding what these models are capable of.
I assume at some point there will be a model that can trivially beat experienced humans at Pokemon PVP, or your invented video game, without needing guides or hundreds of hours. I’m mostly trying to get a sense of how close we are to that, today.
That said, I’d also be impressed by an LLM that is smart enough to look up a strategy guide unprompted, especially if it can apply that information well enough to place at the top of leader boards.
Solving a metagame from first principles, without guides or hundreds of hours sounds like a superhuman skill. Which I agree current LLMs lack, but I bring up fairness because if we’re listing “not superhuman” as a limitation of current LLMs then we might as well complain they can’t prove the Riemann hypothesis.
I have been very impressed with Fable, so I’m making a note to self to spend a bunch of tokens having it learn Pokemon and see if it can put out a half-decent performance, a few months from now when the tokens are cheaper.
I’m not really focused on “fair”—just understanding what these models are capable of.
I assume at some point there will be a model that can trivially beat experienced humans at Pokemon PVP, or your invented video game, without needing guides or hundreds of hours. I’m mostly trying to get a sense of how close we are to that, today.
That said, I’d also be impressed by an LLM that is smart enough to look up a strategy guide unprompted, especially if it can apply that information well enough to place at the top of leader boards.
Solving a metagame from first principles, without guides or hundreds of hours sounds like a superhuman skill. Which I agree current LLMs lack, but I bring up fairness because if we’re listing “not superhuman” as a limitation of current LLMs then we might as well complain they can’t prove the Riemann hypothesis.
I have been very impressed with Fable, so I’m making a note to self to spend a bunch of tokens having it learn Pokemon and see if it can put out a half-decent performance, a few months from now when the tokens are cheaper.