Research directions in condensation: varieties of objectivity
This is the first part of a survey of various ways that I’d like to see work on the theory of condensation develop. Condensation is a mathematical theory dealing with the organization of descriptions of the world into conceptual parts; some of the existing work on it is presented in the paper (Eisenstat 2025). This sequence will draw from that paper the definitions of random variable models and latent variable models, the notation for indexing subfamilies of variables in such models, and the objectivity theorem, Theorem 6.8. In the condensation paper, Section 3, Ideas, gives an overview of all of this. For more context on condensation, readers can refer to LessWrong posts including (Demski 2025, 2026; Gillen and Chiang 2026; Kirchner 2026). The first two posts of this sequence will introduce some central directions of current work—mainly, the concepts of almost perfect condensation and Kolmogorov (or algorithmic-information) condensation—which will be used in other sections, but the different parts are mostly independent beyond that.
I’d encourage those making a serious effort on any of these problem to contact me for further thoughts and coordination.
Thanks to Kaarel Hänni and James Cook for some of the ideas behind this research program, and to James Cook and Jeremy Gillen for comments on this document.
1. Varieties of objectivity
For the theory of condensation to be useful, we’d like hypotheses that we can assume about latent variable models that are general enough so that such latent variable models exist in cases of interesting random variable models, but narrow enough to imply objectivity-like conclusions—statements of approximate uniqueness along the lines of Theorem 6.8 of (Eisenstat 2025). We’ll refer to hypotheses of this kind as condensation properties. To know that such properties apply to interesting data, we’d like to come up with interesting random variable models satisfying them. (We can also look for interesting string models, which will be the analogue of random variable models when we introduce Kolmogorov condensation in the next post.) Work in these directions will also affect the problems discussed in later sections, which almost all either need to assume the existence of a latent variable model (or a latent string model in the Kolmogorov case) with good condensation properties, or else ask for such latent variable models in some particular context.
1.1. Almost perfect condensation
One strong condensation property is almost perfect condensation. I’ve defined this somewhat differently in different places, and there isn’t yet a good reference for this. We’ll look at a definition here, and we’ll discuss informally the consequences that follow from and motivate the definition, without giving a theorem statement or proof sketch.
Definition 1. Suppose that
Here, we should have an objectivity theorem as follows. If
While we can prove theorems like these, I’d like to better understand their range of applicability. I’d like to have better developed examples to which these theorems would apply, and to draw from them in order to better formulate similar theorem statements that more fully bring out what these methods can tell us.
We can look at a simple such example of almost perfect condensation. In the following, the latents will be the biases of dice, and the given variables will be observations that give us some information about these dice. That is, we will have a set of latents
We need further latents to account for the idiosyncratic information in
It would be helpful to have examples of almost perfect condensation that are a bit less artificial, and better demonstrate its relevance to constructing interesting latents. This could be in either the Shannon entropy or the Kolmogorov complexity setting. It would also be good if there were better ways to think about the rather long definition given, in which many somewhat arbitrary choices were made. For example, in Definition 1, many things that we want to be small are individually bounded by the same
1.2. Beyond almost perfect condensation
Almost perfect condensation has been constructed so as to derive reasonably strong and simple bounds from the objectivity theorem, and more generally aid in the interpretation of its conclusions. However, there may be many other cases where the objectivity theorem can tell us that some latent model is approximately unique enough to be interesting. We will look at some examples and discuss why almost perfect condensation will be too strict to allow the sort of analysis that we’d want, and argue that there’s reasonable hope for other ideas here. But first, we’ll review the objectivity theorem so that we can see how it applies here. We’ll work in the Shannon setting for simplicity, but these problems also exist in the Kolmogorov setting.
The general idea of the objectivity theorem is as follows. Let’s say
Now, suppose that
we can see that we can achieve any desired set
which will have size at most
By the definition of
From this, we can motivate condition (2) in the definition of almost perfect condensation. Each set
We can make two key observations here. First, we didn’t actually need to assume that for every sufficiently large
I expect that there is at least some, and possibly a lot, of progress to be made in these directions, which would extend the space of distributions that we can say something interesting about. There are plausibly many unanticipated phenomena here, but I’ll give some indication of a few directions to try.
1.2.1. Almost perfect condensation with a relation
In the definition of almost perfect condensation, we imposed the condition that
We can produce examples by thinking about situations where
This suggests an idea like the following. Instead of asking about those
Definition 2. Let
We say that
This definition has many of the same good properties as almost perfect condensation, for the same reasons. It is a little more opaque though; it would be easier to understand if we had appropriate examples. The geometric ideas mentioned above seem somewhat promising as a source, though there are some limitative reasons to think that many natural such random variable models, such as joint distributions coming from the colours at different points of a visual scene, will not admit this kind of almost perfect condensation. We will discuss this in the next subsection. However, there may still be some useful geometric idea. We can also put other kinds of structure on the set of observed variables. For example, maybe we have many experimental subjects, and we measure many properties of each subject, giving us a set of observations forming a 2-dimensional table. Note here that the different experimental subjects are being treated as giving us different random variables in our joint distribution, rather than being treated as different i.i.d. draws from the same distribution. This could therefore be a better fit for Kolmogorov condensation, which we will discuss in Section 2.
This also doesn’t say anything about where
As before, we may do better by adapting
1.2.2. Medium-scale effects
The hypothesis of well-separatedness is overly restrictive around what we might think of as medium scale effects, in a sense that we’ll look at next. The parameter—either
In the case of a visual scene, we might imagine trying to construct a latent variable model using different latent variables for different properties of the objects visible in the scene. For example, we can have a latent with a large contribution set specifying the approximate position and orientation of a tree, with smaller latents filling in details about its shape, e.g. its branches, which are more local and thus have smaller contribution sets. If
While this visual scene idea is described informally, we can come up with particular distributions in which these sort of problems exist, which can be good test cases for refining these ideas. Here I have in mind distributions like those coming from statistical physics, like a random walk or a Ising model near its critical temperature. Whether or not the more specific suggestions above are pointing in a useful direction, it would be good to have any kind of analysis of these distributions. Near the critical temperature, a typical Ising model configuration has a scale-free structure, where it is made up of large domains separated by walls, but then these domains have smaller islands with the same shapes. In order to understand these, physicists introduce coarse-grained variables in renormalization schemes, which are rather like the latent variables that we care about here. One version of this therefore is that we want to know what the objectivity theorem can tell us about the relationship between two different renormalization schemes.
The random walk example is simpler to describe. Given some number
Now, for an objectivity theorem, we don’t expect any two such trees to correspond exactly; they have different cut points. However, given any variable at layer
This corresponds to what the objectivity theorem should tell us in the case of almost perfect condensation with a relation. While together with well-separatedness properties we can get a bijection on latent variables, the same ideas give us these few-to-one relations without the well-separatedness assumption. Is this the best we can do? Is this the right way to understand this example? What other class of latent variable models for a random walk does this analysis generalize to? Does it generalize further to the critical Ising model, or to other such examples that we can come up with? It would be good to understand these examples better.
1.3. Almost perfect condensation and scoring functions
In Eisenstat (2025) and in Demski (2025), condensation is expressed in terms of some functions, the simple score and the conditioned score, which relate it to compression. Almost perfect condensation and related ideas are helpful for understanding why the objectivity theorem gives good bounds in certain situations, but the hypotheses of almost perfect condensation don’t have a clear relation to compression. It would be interesting if these two perspectives could be unified.
References
Demski, Abram. 2025. “Condensation. LessWrong.”
Demski, Abram. 2026. “Condensation & Relevance. LessWrong.”
Eisenstat, Sam. 2025. “Condensation: A Theory of Concepts.” ODYSSEY 2025 Conference.
Your idea of using statistical physics distributions as a testbed for condensation is really interesting. I’ve been exploring some ideas along these lines, but not (yet) using the condensation formalism. Specifically, I’ve been studying a hierarchical latent-variable model based on critical high-dimensional percolation: https://arxiv.org/abs/2606.20347 Each cluster is modeled as a random tree and is embedded as a branching random walk. A binary tree of latents splits each cluster into subclusters. Given a fixed cluster/tree, you can generate many latent hierarchies consistent with it (in principle—I haven’t actually implemented this yet). The converse is also true. If I understand correctly the random-walk setting you described, I believe this model essentially generalizes that from a latent hierarchy on a chain to one on a random tree. I’d be quite interested in understanding better what condensation can say about these kinds of settings.