I appreciate what this post is trying to point out, although I am sympathetic to some of the other comments that it doesn’t make an airtight case for its specific analogy-thesis.
A while ago I gave a talk at a local university in Tokyo about agent foundations. (You can look at the speaker notes to get a more-or-less verbatim idea of what I said.) I tried to give a relatively fair tour of the field, but the final slide section was titled “Does any of this matter?”, and the final slide, well, let me just quote it:
So overall, my impression of the field of agent foundations is that they haven’t caught up to the deep learning revolution. We have a bunch of research programs, mostly rooted in the 2014-era paradigm of aligning idealized, AIXI-ish agents. They’re proceeding fairly independently, and as we saw from the work on corrigibility, even when progress gets made, the field doesn’t have great way of building on this progress.
But if we go back to my earliest slides, the whole point of building these mathematical theories is that they should be universally applicable! We should be able to ask questions like:
Is ChatGPT using functional decision theory or logical induction in its forward passes?
How does Gemini updates its beliefs? Is it an infra-Bayesian reasoner?
Is Claude finding natural abstractions in its weights, or in the Minecraft world?
What are the shards that emerged during LLaMA’s training process? Can we find them, using interpretability tools?
These seem like pretty interesting questions to me! We’ve had since 2014 to develop a bunch of mathematics that supposedly applies to any agent. Since 2022, we’ve had actual machine intelligences, and they’re at least capable of simulating agency. Why haven’t people been applying these tools to modern LLMs?
I’ve seen some papers vaguely gesturing in these directions, but they mostly amount to asking the LLMs questions, like “how would you decide in Newcomb’s problem”. There was a very recent paper published about whether LLMs introspect, that had a clever twist involving comparing two different models, but was still done using black-box testing. That’s something, but… you don’t need a fully mathematical theory of agency for that!
I think these would be very interesting research problems. Either to try and succeed, or try and find out why all that agent foundations math doesn’t actually help with real-world systems after all. Both would be pretty exciting!
So yeah, I strongly sympathize with the vibe that agent foundations seems unmoored from the program of actually aligning agents, building castles in the sky on top of abstract mathematical formulations but never quite getting around to showing that they are useful.
(And I love a good abstract mathematical castle! It would give me great personal joy to sit down and become one of the few people in the world who understands all the infra-Bayesianism math. I just don’t see how it’s going to help us align AIs.)
I should probably write up that slideshow into a top-level post at some point, but, it’s a bit intimidating to anticipate what the reactions might be from all the people I’m critiquing. So, I’ll hide it in this comment for now.
Maybe someone should attempt an agent-foundations analysis of one of those DeepSeek AIs for which everything (architecture, training protocol, etc) is fully in the public domain?
I appreciate what this post is trying to point out, although I am sympathetic to some of the other comments that it doesn’t make an airtight case for its specific analogy-thesis.
A while ago I gave a talk at a local university in Tokyo about agent foundations. (You can look at the speaker notes to get a more-or-less verbatim idea of what I said.) I tried to give a relatively fair tour of the field, but the final slide section was titled “Does any of this matter?”, and the final slide, well, let me just quote it:
So yeah, I strongly sympathize with the vibe that agent foundations seems unmoored from the program of actually aligning agents, building castles in the sky on top of abstract mathematical formulations but never quite getting around to showing that they are useful.
(And I love a good abstract mathematical castle! It would give me great personal joy to sit down and become one of the few people in the world who understands all the infra-Bayesianism math. I just don’t see how it’s going to help us align AIs.)
I should probably write up that slideshow into a top-level post at some point, but, it’s a bit intimidating to anticipate what the reactions might be from all the people I’m critiquing. So, I’ll hide it in this comment for now.
Maybe someone should attempt an agent-foundations analysis of one of those DeepSeek AIs for which everything (architecture, training protocol, etc) is fully in the public domain?