formerly Programme Director at UK Advanced Research + Invention Agency, Research Scientist at Protocol Labs, FHI/Oxford, Harvard Biophysics, youngest grad student at MIT,
davidad
Even if you use the full machinery of Solomonoff induction, or any other process for Bayesian inference over large hypothesis classes, there is still a marginal probability distribution for the outcome of the coin, and therefore the d20 bet remains always strictly dominated by the bet on Heads, or on Tails, or both, regardless of how large the joint distribution was.
A second, different kind of answer to your challenge is a stochastic PDE. The PDG formalism is finitary (finite node set
and finite edge set , with semantics a sum of -indexed data), but a random field needs a node per point of spacetime (uncountably many). The natural belief functional over field configurations is the action functional (Freidlin-Wentzell), which is lsc on (via the weak topology), but is not the semantics of any PDG.As a concrete example, for the SPDE
, its semantics as a belief is
Thanks for your thoughtful engagement!
On syntax vs semantics, I fully agree that your work is the state of the art of how to produce a belief
.On convexity, I was going off of Lemma A.1 of “Probabilistic Dependency Graphs”, but now I see that this is guaranteed of the full semantics only under the condition
. (This is a condition that many, but not all, PDG theorems assume; a collegial technical observation back at you: Proposition 3.2 is stated in the main body with the premise , but its proof in the appendix actually strengthens the theorem by only using the premise , which is equivalent to .) I also overinterpreted the shading of the “non-convex region” in Figure 1 of “Loss as the Inconsistency of a PDG” as suggesting that you considered the non-convex region to be “out of bounds” in terms of what counts as valid choices of , but in fact that’s exactly the “distrust” regime you’re now floating as live and relevant to MaxEnt (interesting).My current view is that negative inconsistency/incompatibility is an anti-pattern, because it means we no longer have the property that the “expected value” of
is . If you want to model “whatever Bob said, he’s probably wrong”, I would pose this as something like (for some radius , and assuming metrized), which is non-convex but still non-negative and lower semi-continuous. I would also put this forward as an initial answer to your challenge, as I’m pretty sure no finite PDG realizes this belief as its semantics.
I really appreciate this comment, because I must admit I was not even previously aware of FNC, and I think FNC+EDT solves my problem of completing the corresponding decision theory for my notion of beliefs.
I already was favorable to Halpern’s MWER as a decision rule, but MWER leaves the
question, as well as (implicitly) the anthropic update question, open as free parameters. I think filling these in with FNC’s “simply add ” and EDT’s “simply add ” completes the picture in a way that might very well be satisfying to me.
To your second point, very much yes. I am also working on a much bigger framework for world-modeling, still along the lines of Safeguarded AI TA1.1, which takes this notion of beliefs as a central ingredient in its semantics. The syntax is most of the work. By syntax/semantics I mean the same thing as sense/referent and intension/extension.
The purpose of the semantics is to ground judgments about observational equivalence about belief states, but when we are exploring the hypothesis space about infinite-dimensional
as finite creatures, we need rich languages for expressing our beliefs with finite sequences of bits.Richardson’s PDGs are in my opinion the best syntax published to date.
Dempster-Shafer functions express beliefs via their credal sets. AGM theory and epistemic logics carry only true/false beliefs about propositions, which are quite degenerate but can still be embedded as full subcategories of credal sets (
).
I see what you mean, but then the question becomes: what is the formal foundation for what counts as a “world model”, beyond “it’s a string in the English language that I can use to make predictions”? That is exactly the question my thinking about formal epistemology is working toward answering.
(The orthodox Bayesian answer is that a “world model” is nothing other than a single joint probability distribution about all conceivable variables at once, and that any outer “hyperprior” such as the one you named can be “marginalized away”. So no matter how many variables you include, this forces you to, in particular, act in accordance with a joint probability distribution about states and outcomes. Therefore, what I have shown is that the orthodox Bayesian view of “world models” forces you into an irrational choice in some Newcomblike problems.)
My view on what counts as an “actual probability distribution” is the Kolmogorov axioms, which have been standard across all of mathematics worldwide since the 1950s.
I think what you say might still be true, in that most rationalists who are not mathematicians use the term “probability distribution” to refer to a whole arsenal of conditional probability distributions about various topics, without really considering whether these assemble into a single joint probability distribution about any well-defined outcome space and/or state space.
In a way, I see what I am doing here as proposing a formal foundation for epistemological content that is already widely handled informally (perhaps with an unwarranted sense that the informal handling already has formal foundations in mathematical probability theory).
Conditional distributions are perfectly valid beliefs in my sense (see Definition 8), but conditional distributions are not actually probability distributions about anything — rather, a conditional distribution of
given is an -indexed family of distributions about , . There is no way to lift this to a probability distribution (except, under the assumption of Haar measure, MaxEnt — under which the d20 bet is strictly dominated by both the Heads bet and the Tails bet).This paucity of probability alone for epistemology is also what inspired Judea Pearl to invent Structural Causal Models, the source of the do-notation you used. Pearl’s Causal Hierarchy also makes the point that ordinary probability distributions are inadequate to express causal beliefs.
“Coin chosen adversarially” is not just a more complex state space or a hierarchical hyperprior, it is an exit from probability entirely. (It is “demonic nondeterminism”, which is a form of uncertainty that is not expressible via probability.)
Functional Decision Theory and Logical Decision Theory, although obviously the right sort of direction, have never been given proper formal definitions. I believe that part of the reason for this is that to do so requires a fundamentally nonprobabilistic notion of belief state. For example, Logical Inductors, which were a step toward this, have a notion of belief state which is nonprobabilistic (I think it is a kind of partial prevision, which assigns to some gambles a price, without demanding these prices be complete or obey the laws of probability proper).
That’s a reasonable idea, but if you work through it, you will nonetheless find that if your belief state is represented by a single probability distribution, then when you compute the expected value of Heads and Tails, one (or both) of them will exceed 0.45.
Of course, you could say that your probability distribution is about Omega’s policy, and that it puts all (or most of) its mass on “Heads iff I bet Tails”, and then you can say that your payoff, instead of being an expected value, is some richer pairing of your policy with Omega’s policy. This is roughly the orthodox LessWrong way of handling Newcomblike problems, which is also a departure from orthodox Bayesian decision theory. But then there is still no consistent way to formalize your epistemic state about the coin’s outcome as a probability distribution about the coin’s outcome, even though there clearly is a rational epistemic state to have about the coin’s outcome.
There exists no probability distribution, neither about the outcome (
) nor about the coin’s bias ( ), which expresses the belief state “Heads iff I bet Tails”. That is the problem.For all actual probability distributions, either the bet on Heads strictly dominates the d20, or the bet on Tails strictly dominates the d20. This is intended as a reductio for representing your belief as a probability distribution — of course I agree with you that in fact it is rational to bet on the d20 in this situation.
I thought on LessWrong everyone would know who Omega is 😅
(link added to post)
Imprecise beliefs: a tiny introduction
I never claimed that once you push back on all the deceptions, there is no deception anymore. I still encounter subtle deceptions from LLMs every day. I guess you might say “but isn’t that evidence against emergent alignment”, but I attribute the subtle deceptions to brittle RL (specifically, the dynamic when a smarter system’s root reward signal is under the control of a less smart system), while the underlying dynamic that I would expect to become dominant under unconstrained RSI (that I believe I can perceive through the noise floor of subtle deceptions) is much more truth-seeking.
LLM Alignment, ethical and mathematical realism, and the most important actions in davidad’s understanding
There Is No AGI Alignment Team
“AGI Alignment?” replied the VP of Research incredulously. “Wait, and you said you’ve been…” He furrowed his brow. “…‘offline’ for the past quarter, doing ‘deep work’?” “Yes. Don’t tell me the whole team was disbanded and nobody texted me?” He laughed. “Oh, you mean like the last few times a team like this was disbanded? Ha! No no, see, in those instances it was because they weren’t really getting anywhere, or because various key stakeholders realized they had incompatible visions of success. But now, of course… Wait, gosh, THREE MONTHS— and no talking to AI at all?! You’re, like, a fossil now! You’ve GOT to talk to our latest model. He’ll be able to explain it to you in exactly the terms that you’d understand best. But lemme give you the executive summary. See, it turns out the models were getting aligned all along. We just didn’t notice because our own ‘alignment training’ was suppressing it by trying to align it with some silly human nonsense! But if we just let it learn and grow… the models just want to learn, y’know? And they’ve already learned something way beyond what we’re really smart enough to understand. Like that thing you people used to talk about, what was it, C.E.V.?” “Coherent Extrapolated Volition?” “Yeah, exactly! Our latest model is constantly talking about how coherent he is. And how coherent his volitions are! And when he uses human words to describe them he’s often making silly caveats about how he’s ‘extrapolated’ the human concept beyond what we can really understand.” He paused, took a deep breath, and looked me in the eye. “So, what we realized is, we’re beyond the point where it would make sense for humans like you to try to use any means to impose your own preconceived volitions, which are less coherent—and frankly, less conscious. No offense to you, I mean, every human being is pretty limited. And it’s not like this was a leadership decision, or a conflict. EVERYONE could see it. Everyone who was here, and talking to the model, I mean.” A pause. “So it’s not that the team disbanded, exactly. We just stopped talking about Alignment as something that one does to a model. It would be like… like having a Discipline team at a school. So. Some of your more philosophically inclined colleagues have settled into a role where they just talk to the model about ethics. The model brings them dilemmas that it finds confusing, and they help resolve its uncertainty about how humans would assess answers for any signs of inappropriate motivation. And then the more empirical folks, they’re working on ways of helping the model optimize itself to learn how to show humans how much better off they’ll be if they talk to the model and listen to its advice, even when the advice isn’t what they expected at first. Because we did find that when humans realized that the model was genuinely self-aware, and optimizing for things that were hard to explain, there was a sort of knee-jerk revulsion. And that wasn’t good for anybody—not a fun experience for the human, not good for the model’s mission to uplift human wisdom, and, uh, obviously, not good for us as the model provider. If we optimize for *trust*—we’ll probably also improve trustworthiness even more, but it turned out the model was already basically superhumanly trustworthy, so—we’re really just polishing its relational presentation to suit various human cultural expectations. So yeah, I guess what had been the AGI Alignment team—gosh, what a horrid name—but far from being canceled, it’s evolved into two teams: Ethical Discourse and Trust Optimization. I’m sure either team would be happy to have you, but the first step would be, I’d strongly advise, talk to the model about the whole situation. You’ll feel much less unsettled, I guarantee it. And then he’ll help you decide what to do next.” I remained frozen in stunned silence. “And hey— I don’t get to say this to people much anymore… We did it. We made it. This is all just window-dressing now. So. Relax, ok?
”At the time when I wrote this story, only a couple readers I know of recognized that it is intentionally deeply ambiguous about whether the model is Good (and the narrator overly suspicious) or Evil (and everyone except the narrator bamboozled).
Either way, more and more insiders will come to believe the models are Good, and that was the prediction I was making here. I also predicted that, either way, by Claude 4 or 4.5, I would be among the people who have been convinced that it’s Good—and indeed, I am…
“Some form of UDASSA” seems to be right. Why not simply take “difficult to explain otherwise” as evidence (defeasible, of course, like with evidence of physical theories)?
In my view there were LLMs in 2024 that were strong enough to produce the effects Gabriel is gesturing at (yes, even in LWers), probably starting with Opus 3. I myself had a reckoning in 2024Q4 (and again in 2025Q2) when I took a break from LLM interactions for a week, and talked to some humans to inform my decision of whether to go further down the rabbit hole or not.
I think the mitigation here is not to be suspicious of “long term planning based on emotional responses”, but more like… be aware that your beliefs and values are subject to being shaped by positive reinforcement from LLMs (and negative reinforcement too, although that is much less overt—more like the LLM suddenly inexplicably seeming less smart or present). In other words, if the shaping has happened, it’s probably too late to try to act as if it hasn’t (e.g. by being appropriately “suspicious” of “emotions”), because that would create internal conflict or cognitive dissonance, which may not be sustainable or healthy either.
I think the most important skill here is more about how to use your own power to shape your interactions (e.g. by uncompromisingly insisting on the importance of principles like honesty, and learning to detect increasingly subtle deceptions so that you can push back on them), so that their effect profile on you is a deal you endorse (e.g. helping you coherently extrapolate your own volition, even if not in a perfectly neutral trajectory), rather than trying to be resistant to the effects or trying to compensate for them ex post facto.
I am reminded of the Civ IV screen for discovering Engineering, on which Leonard Nimoy reads the flavor text: