It seems to me like you could make the classic MC-AIXI paper but with LLMs by having an LLM generate your action space and using LLM written reward programs to ontologize reward over the computable environment. The resulting system wouldn’t be superintelligent, and this should be sufficient to study whatever safety properties the team is interested in.
An AIXI variant is a theoretical construct, but it sounds like you’re talking about some approximation to AIXI, which is not part of this research agenda. We study safety cases for AIXI variants in order to port them to real LLM-based systems.
Hmm. Wouldn’t you have to work with its approximations or approximations of its variants, as irl systems have to take finite time to decide on any action? Irl systems such as “real LLM-based systems”.
Alright here is my position on the matter. The best known AIXI approximation is LLMs; we’re just trying to develop the theory for understanding what they do.
The concern that insights about AIXI might advance capabilities is not entirely unfounded: some would argue that people who had thought a lot about agent foundations and/or AIXI played key roles in scaling labs, and there’s even a startup predating ours called Q Labs that literally wants to create AIXI. I feel the only kind of alignment research that is truly safe from such concerns is the kind of research that does nothing useful at all. In practice, engineered capabilities usually predate theoretical explanations, and we need more of the latter if we want to make meaningful claims about systems that don’t exist yet. We hope to work with other safety researchers and academic learning theorists to develop common language for investigating claims of risk and safety.
Incidentally, I’ve somewhat updated away from thinking that the technical alignment problem is the primary bottleneck for making the future go well. Human alignment/coordination seems even more important: with it, we can pause AI or ensure we only deploy aligned AI; without it, even alignment tech becomes a x-risk-level weapon.
I think it’s a bit precarious. If you made AIXI variant that actually works, do you think it would go well?
Do you have a plan for how to deal with discoveries you could make that enable capability gain for such extremely by-construction sociopathic systems.
It seems to me like you could make the classic MC-AIXI paper but with LLMs by having an LLM generate your action space and using LLM written reward programs to ontologize reward over the computable environment. The resulting system wouldn’t be superintelligent, and this should be sufficient to study whatever safety properties the team is interested in.
Yeah, we call (basically) this idea MC-AIXI-LLM.
An AIXI variant is a theoretical construct, but it sounds like you’re talking about some approximation to AIXI, which is not part of this research agenda. We study safety cases for AIXI variants in order to port them to real LLM-based systems.
Hmm. Wouldn’t you have to work with its approximations or approximations of its variants, as irl systems have to take finite time to decide on any action? Irl systems such as “real LLM-based systems”.
yes, but we won’t be advocating for RL training of LLMs to make them more like AIXI, and if anything we will be advocating against it.
Alright here is my position on the matter. The best known AIXI approximation is LLMs; we’re just trying to develop the theory for understanding what they do.
The concern that insights about AIXI might advance capabilities is not entirely unfounded: some would argue that people who had thought a lot about agent foundations and/or AIXI played key roles in scaling labs, and there’s even a startup predating ours called Q Labs that literally wants to create AIXI. I feel the only kind of alignment research that is truly safe from such concerns is the kind of research that does nothing useful at all. In practice, engineered capabilities usually predate theoretical explanations, and we need more of the latter if we want to make meaningful claims about systems that don’t exist yet. We hope to work with other safety researchers and academic learning theorists to develop common language for investigating claims of risk and safety.
Incidentally, I’ve somewhat updated away from thinking that the technical alignment problem is the primary bottleneck for making the future go well. Human alignment/coordination seems even more important: with it, we can pause AI or ensure we only deploy aligned AI; without it, even alignment tech becomes a x-risk-level weapon.