The LLM’s tactic was still a simple, brute force solution: develop a single powerful Pokemon and use it’s raw stats to overwhelm the opposition.
It would be cool to see GPT-5 play Pokémon less like 6-year-old me, and more like an accomplished player, building a varied team of creatures that can take on any battle in the game.
The Pokemon Red speedrunning community is made up of extremely dedicated hobbyists who spend hundreds or thousands of hours trying to beat the game as quickly as possible, and the uncontroversial, obvious best strategy they employ is to develop a single powerful Pokemon and use its raw stats to overwhelm to opposition. The Pokemon games are simply not very well balanced, a single highly developed Pokemon can easily crush every battle and it seems unfair to call it a limitation that the LLM did not put on a cool show.
Choosing brute force isn’t a limitation, agreed. But being incapable of other strategies is!
It’s less a judgment about the tactic and more that it eliminates a lot of decisions in a fight (swapping out, type-type matchups, etc.). It doesn’t test “intelligence” as deeply as building a balanced team.
My understanding is LLMs still struggle to put together a good PVP team, where those elements can’t be ignored.
My understanding is that most human players still struggle to put together a good PVP team, and the ones that do it well only got there after reading strategy guides for many hours.
It kind of raises the question of what a fair comparison is. If I invented a completely new videogame, kept it out of the training corpus and let humans practice it for hundreds of hours, obviously it wouldn’t be fair to compare those human experts to an unpracticed LLM. Are Pokemon PVP guides sufficiently represented in the corpus that we should expect an LLM to know their contents in the way we expect it to know the DSM V?
To be clear, I still expect cutting edge LLMs to lose whatever I would consider a fair comparison. But also, I saw Fable beat Pokemon Red in not that much longer than the average player, so I wouldn’t be surprised if the loss looked more like a difference in degree than the kind of thing one lists as a limitation.
I’m not really focused on “fair”—just understanding what these models are capable of.
I assume at some point there will be a model that can trivially beat experienced humans at Pokemon PVP, or your invented video game, without needing guides or hundreds of hours. I’m mostly trying to get a sense of how close we are to that, today.
That said, I’d also be impressed by an LLM that is smart enough to look up a strategy guide unprompted, especially if it can apply that information well enough to place at the top of leader boards.
Solving a metagame from first principles, without guides or hundreds of hours sounds like a superhuman skill. Which I agree current LLMs lack, but I bring up fairness because if we’re listing “not superhuman” as a limitation of current LLMs then we might as well complain they can’t prove the Riemann hypothesis.
I have been very impressed with Fable, so I’m making a note to self to spend a bunch of tokens having it learn Pokemon and see if it can put out a half-decent performance, a few months from now when the tokens are cheaper.
The Pokemon Red speedrunning community is made up of extremely dedicated hobbyists who spend hundreds or thousands of hours trying to beat the game as quickly as possible, and the uncontroversial, obvious best strategy they employ is to develop a single powerful Pokemon and use its raw stats to overwhelm to opposition. The Pokemon games are simply not very well balanced, a single highly developed Pokemon can easily crush every battle and it seems unfair to call it a limitation that the LLM did not put on a cool show.
Choosing brute force isn’t a limitation, agreed. But being incapable of other strategies is!
It’s less a judgment about the tactic and more that it eliminates a lot of decisions in a fight (swapping out, type-type matchups, etc.). It doesn’t test “intelligence” as deeply as building a balanced team.
My understanding is LLMs still struggle to put together a good PVP team, where those elements can’t be ignored.
My understanding is that most human players still struggle to put together a good PVP team, and the ones that do it well only got there after reading strategy guides for many hours.
It kind of raises the question of what a fair comparison is. If I invented a completely new videogame, kept it out of the training corpus and let humans practice it for hundreds of hours, obviously it wouldn’t be fair to compare those human experts to an unpracticed LLM. Are Pokemon PVP guides sufficiently represented in the corpus that we should expect an LLM to know their contents in the way we expect it to know the DSM V?
To be clear, I still expect cutting edge LLMs to lose whatever I would consider a fair comparison. But also, I saw Fable beat Pokemon Red in not that much longer than the average player, so I wouldn’t be surprised if the loss looked more like a difference in degree than the kind of thing one lists as a limitation.
I’m not really focused on “fair”—just understanding what these models are capable of.
I assume at some point there will be a model that can trivially beat experienced humans at Pokemon PVP, or your invented video game, without needing guides or hundreds of hours. I’m mostly trying to get a sense of how close we are to that, today.
That said, I’d also be impressed by an LLM that is smart enough to look up a strategy guide unprompted, especially if it can apply that information well enough to place at the top of leader boards.
Solving a metagame from first principles, without guides or hundreds of hours sounds like a superhuman skill. Which I agree current LLMs lack, but I bring up fairness because if we’re listing “not superhuman” as a limitation of current LLMs then we might as well complain they can’t prove the Riemann hypothesis.
I have been very impressed with Fable, so I’m making a note to self to spend a bunch of tokens having it learn Pokemon and see if it can put out a half-decent performance, a few months from now when the tokens are cheaper.