My understanding is that most human players still struggle to put together a good PVP team, and the ones that do it well only got there after reading strategy guides for many hours.
It kind of raises the question of what a fair comparison is. If I invented a completely new videogame, kept it out of the training corpus and let humans practice it for hundreds of hours, obviously it wouldn’t be fair to compare those human experts to an unpracticed LLM. Are Pokemon PVP guides sufficiently represented in the corpus that we should expect an LLM to know their contents in the way we expect it to know the DSM V?
To be clear, I still expect cutting edge LLMs to lose whatever I would consider a fair comparison. But also, I saw Fable beat Pokemon Red in not that much longer than the average player, so I wouldn’t be surprised if the loss looked more like a difference in degree than the kind of thing one lists as a limitation.
I’m not really focused on “fair”—just understanding what these models are capable of.
I assume at some point there will be a model that can trivially beat experienced humans at Pokemon PVP, or your invented video game, without needing guides or hundreds of hours. I’m mostly trying to get a sense of how close we are to that, today.
That said, I’d also be impressed by an LLM that is smart enough to look up a strategy guide unprompted, especially if it can apply that information well enough to place at the top of leader boards.
Solving a metagame from first principles, without guides or hundreds of hours sounds like a superhuman skill. Which I agree current LLMs lack, but I bring up fairness because if we’re listing “not superhuman” as a limitation of current LLMs then we might as well complain they can’t prove the Riemann hypothesis.
I have been very impressed with Fable, so I’m making a note to self to spend a bunch of tokens having it learn Pokemon and see if it can put out a half-decent performance, a few months from now when the tokens are cheaper.
My understanding is that most human players still struggle to put together a good PVP team, and the ones that do it well only got there after reading strategy guides for many hours.
It kind of raises the question of what a fair comparison is. If I invented a completely new videogame, kept it out of the training corpus and let humans practice it for hundreds of hours, obviously it wouldn’t be fair to compare those human experts to an unpracticed LLM. Are Pokemon PVP guides sufficiently represented in the corpus that we should expect an LLM to know their contents in the way we expect it to know the DSM V?
To be clear, I still expect cutting edge LLMs to lose whatever I would consider a fair comparison. But also, I saw Fable beat Pokemon Red in not that much longer than the average player, so I wouldn’t be surprised if the loss looked more like a difference in degree than the kind of thing one lists as a limitation.
I’m not really focused on “fair”—just understanding what these models are capable of.
I assume at some point there will be a model that can trivially beat experienced humans at Pokemon PVP, or your invented video game, without needing guides or hundreds of hours. I’m mostly trying to get a sense of how close we are to that, today.
That said, I’d also be impressed by an LLM that is smart enough to look up a strategy guide unprompted, especially if it can apply that information well enough to place at the top of leader boards.
Solving a metagame from first principles, without guides or hundreds of hours sounds like a superhuman skill. Which I agree current LLMs lack, but I bring up fairness because if we’re listing “not superhuman” as a limitation of current LLMs then we might as well complain they can’t prove the Riemann hypothesis.
I have been very impressed with Fable, so I’m making a note to self to spend a bunch of tokens having it learn Pokemon and see if it can put out a half-decent performance, a few months from now when the tokens are cheaper.