Alright here is my position on the matter. The best known AIXI approximation is LLMs; we’re just trying to develop the theory for understanding what they do.
The concern that insights about AIXI might advance capabilities is not entirely unfounded: some would argue that people who had thought a lot about agent foundations and/or AIXI played key roles in scaling labs, and there’s even a startup predating ours called Q Labs that literally wants to create AIXI. I feel the only kind of alignment research that is truly safe from such concerns is the kind of research that does nothing useful at all. In practice, engineered capabilities usually predate theoretical explanations, and we need more of the latter if we want to make meaningful claims about systems that don’t exist yet. We hope to work with other safety researchers and academic learning theorists to develop common language for investigating claims of risk and safety.
Incidentally, I’ve somewhat updated away from thinking that the technical alignment problem is the primary bottleneck for making the future go well. Human alignment/coordination seems even more important: with it, we can pause AI or ensure we only deploy aligned AI; without it, even alignment tech becomes a x-risk-level weapon.
Alright here is my position on the matter. The best known AIXI approximation is LLMs; we’re just trying to develop the theory for understanding what they do.
The concern that insights about AIXI might advance capabilities is not entirely unfounded: some would argue that people who had thought a lot about agent foundations and/or AIXI played key roles in scaling labs, and there’s even a startup predating ours called Q Labs that literally wants to create AIXI. I feel the only kind of alignment research that is truly safe from such concerns is the kind of research that does nothing useful at all. In practice, engineered capabilities usually predate theoretical explanations, and we need more of the latter if we want to make meaningful claims about systems that don’t exist yet. We hope to work with other safety researchers and academic learning theorists to develop common language for investigating claims of risk and safety.
Incidentally, I’ve somewhat updated away from thinking that the technical alignment problem is the primary bottleneck for making the future go well. Human alignment/coordination seems even more important: with it, we can pause AI or ensure we only deploy aligned AI; without it, even alignment tech becomes a x-risk-level weapon.