They are probably referring to these:

They are probably referring to these:

AF is kinda a quite broad term, historically has been a lot of decision theory which does tend to make some of the assumptions you are referring to, but thinking about how to model agents more generally is also a core project of agent foundations
I think that the real reason work in agent foundations isn’t that applicable to current models is mostly that it is just a pretty young small field and still has a long way to go. Progress is very much bottlenecked by smart people getting work done, and eventually it absolutely will be able to help us understand LLMs, along with many other kinds of agents.
Could elaborate on why you think that a strong prior against goal-directedness remains after post training?
Here’s what that same distribution I used above looks like if you plot the closed-form pushforward density analytically. In this picture it’s easier for the visual cortex to pick up on the patterns (although it would still be nontrivial for a human to figure out what should be colored red and what should be colored blue if you erased the colors).
Yup! This is a very weird space to call a ‘thingspace’ but most transformations of it (anything that’s at least approximately injective, for example) will preserve the same concept.
without clusters we don’t have concepts
Nope! Here’s an example:
These dots were sampled from a 2-component Gaussian mixture and then put through a smooth invertible warp. There aren’t any clusters, but the concept is still present and recoverable from the data (altho too hard for our visual cortex to recover in this particular case).
The shortest rule to describe this scatter is that each dot is an independent draw from a mixture of two modes. You have to specify where the two modes are, and you can guess decently well which mode each dot is from. Thru this you’ve rediscovered color without needing clustering.
It seems you are claiming that information theory implies there is an objective measure of clusterhood
Not exactly. Rather, I am saying that clusterhood is not the right way to think about these things at all if we want to be free of an arbitrary choice of basis. Rather, we can study the information theoretic properties of the complete partition lattice over an input space and its corresponding probability distribution.
it seems to assume that for the “clusters in thingspace” there is an objectively natural choice of thingspace
usually natural abstractions research uses information theory to avoid privileging a certain basis. So we actually don’t have the problem of choosing a thingspace and looking for clusters geometrically!
btw here is a cool example of some empirical evidence for convergent abstractions in the some of the contexts we care about: https://www.nature.com/articles/nn.4244
I’d like to write up a more in depth reply to this post eventually, but I have a few thoughts right off the bat.
When I first read your opening claim that CAH is weaker and more likely to be true, I found myself wondering what you thought the difference was. Then, when you explained how you see the difference, I wasn’t satisfied—you didn’t explain as clearly as I would like how abstractions which are features of the world at distance come apart from abstractions which are convergently used by a variety of similarly pressured learners.
One load bearing assumption imo is that the relevant pressures which different learners share are ~ universal—namely, pressures to make good predictions, and pressures to use resources (time and space) efficiently. If we start requiring that agents have to share more pressures than these in order to converge on equivalent abstractions we have lost the kind of strength that I care about.
I currently guess that “good” as humans know it is not a natural abstraction, or an abstraction which different kinds of agents will robustly converge on. Of course this is a very important question indeed and it would be wonderful if it is a natural abstraction when you are trying to predict your observations of humans! And it’s totally possible that is the case.
Oh and last thing—agents with different sensory equipment should still converge on a lot of the same abstractions! I think you were just meaning that they might have some idiosyncratic abstractions in the mix as well, which I grant.
I think reading this post has possibly been worth five hundred bucks to me. Thanks!
TLDR: Make working optional, but require a medium friction process to opt out of working which is designed to prevent people from becoming aimless/depressed.
We might need jobs which are designed not to be productive to society but to provide purpose / fun / fulfillment.
If anyone has something they’d really like to be doing which is not a job, they can apply for money to go do that and probably get funded.
But most people who don’t have the motivation to design their own active and challenging life which keeps them happy will do the default thing which is get a job that seems like something they’d enjoy, and they will be happier than if left to their own devices without direction.
It’s a neat analogy, but I wish you would have tried to give some arguments for the claims you made about qualia, even just explicit references to existing arguments.
I don’t know why you think that the set of qualia should be finite or that qualia depend only on themselves.
This distinction is a great crux. Red is a actually a very large amount of information! Why would you think it’s only 3 bits? The brain does not use a ‘color’ slot with a compressed symbolic representation of color. Our representation of red is quite rich, containing information about associations (warmth, emotions, etc) and would be represented by a quite specific point in a complex high dimensional space.
i don’t really feel either imperils the centre of the argument
I think the center of the argument is basically correct
as for the “multiple agents, all hobbled by an unchanging terminal goal”: well, they’ll be outcompeted by the first one that gives it up.
This is not the scenario I’m imagining. I’m imagining multiple agents, some with thin terminal goals and others concerned purely with Omohundro drives, operating in context where the rational thing to do is the same whether or not you have non-Omohundro terminal goals.
In this case agents with thin terminal goals are not hobbled and they will not be out competed.
I mostly agree with the Landian ‘hypertrophy’ thesis that under selection pressure, the agents will have convergent instrumental goals as their terminal goals.
I also think the orthogonality thesis is poorly named. In the words of David Chalmers:
“orthogonal” in typical english means something more like “uncorrelated” than “dissociable”. “orthogonality thesis” was always a bad name for a thesis about (mere) dissociability.
I do think, however, that the orthogonality thesis’s traditional defenders have not held the strong version you argue against. Yudkowsky, for example, has mostly argued that a paperclipper would be reflectively stable by default, not that it would be equally fit in a competitive selection process.
I also think it’s super important to note that there are many different ways selection processes could look. Some of these could reward agents with specific terminal goals but many of them might not be sensitive to most differences in terminal goals if all agents act approximately independently of their terminal goals during the relevant timescale.
On the other hand, maybe being “logically omniscient,” the guesser would inevitably conclude that a codebook is the only reasonable scheme, and that would outweigh the enormous coincidence of the 5-character string being a valid name.
I don’t think this would be true because the string chooser picks a string after they know which person they are trying to point to
So, given that they need to point to Wei Yi, they might choose to use the name instead of the Schelling codebook if the Schelling codebook has a lower success probability than the name (even if the Schelling point is quite unique!). This is not a repeated game.
It’s hard to say what a UDT agent would do, though. It could go either way.
If they can find a Schelling codebook they might be able to get the perfect 5 ascii chars for each person. But I’m not sure one exists. (perhaps birth time? encoded how?)
A fun question is to think about what is the shortest string you can think of that would identify you. For me it’s the seven characters ‘satchlj’.
Xi and Ye have optimal names for this.
Upper bound for Americans might be USXXX-XX-XXXX
I think they are (at least OAI). Notably, GPT 5.5 Pro is not on the benchmark.