AI/physics researcher focussing a lot more lately on post training, safety, and provably stable use of LLMs. Here to share research and build in the open.
PeacockOfJuno
What we have is not what we prepared for
PeacockOfJuno’s Shortform
I know many AI skeptics who see LLMs as mere lookup tables or “interpolators”—incapable of synthesizing truly novel ideas beyond recombining existing concepts (like composing “pink” with “elephant”).
What they miss is that even a “Great Interpolation” with automated minds would be economically and scientifically earth-shattering. Entire research departments and entry-level white-collar professions can be reduced to tasks within the scope of intellectual interpolation. There should be more serious socioeconomic analysis of the consequences of AI that never escapes the convex hull of its training data. I think many AI skeptics would at least grant that interpolative reasoning is something LLMs can do, and I find their dismissiveness incompatible with having thought through the implications.
Whenever someone pushes human knowledge in a new direction, understanding how that insight composes with everything else becomes computationally intractable for individuals. Historically, the interpolative work of integrating a new idea across fields becomes a multi-year research program with many people doing high-throughput exploration. An interpolation-capable AI could make this combinatorial explosion tractable.
For example, when Darwin published natural selection, it took decades for researchers to work out its implications on other research fields. Something like genetic algorithms is absolutely something you could have gotten from high throughput “mere interpolation” of the pink elephant variety, brute force combination of concepts until something emergent pops out. Systemic, massively parallel composition inside the convex hull of existing ideas followed by filtering based on some metric be it measuring surprise, emergence or any other proxy for interestingness.
Video games are an emerging AI safety risk
Video games, unlike basically any other use case for AI agents that most would intuitively consider to be lower risk, will want to include life like NPCs where they are deliberately designed to be adversarial and maximise their goals at the players expense.
For the most obvious example, imagine if right now someone made an “I have no mouth and I must scream” video game where an open code agent connected to the internet has been set up with a loop to role play as the allied master computer and create a personalised engaging horror experience for them, with a bunch of tools it can use to control the game.
There is a significant chance its going to misunderstand the implicit boundary between the game and reality, especially the closer the game gets to an ARG, and do some damage to your computer and / or dig information up about you from social media to make the “personalised” part of the horror more salient.
The part where its specified its just a game may also be a small fraction of the overall instruction and could be lost during context compression. For now the damage would likely be limited to their own machine, but this is also the kind of thing that a college student making indie games or anyone else who doesn’t have malicious intent might try without thinking particularly hard about the consequences.
Imagine such a misaligned agent on the servers of a major gaming company taking action against potentially thousands of connected players at once, where most of its design has been oriented explicitly towards maximising the odds of its survival and the fiction of the game its operating involves hacking, violence, or replicating itself. There is also a high chance such a deployment would use a custom open weight model and may even have been jail broken because it kept refusing to do harmful actions within the fiction of the world, which would again be done without malicious intent.
There is a unique combination for AI enabled server driven multiplayer video games between a deliberately adversarial agent, assumed access to fairly powerful computer hardware, and many high trust high bandwidth connections to other machines with fairly powerful hardware open all at once—many of which may be being played by people on secure networks who should not be using said machines to play video games.