Researcher at the Center on Long-Term Risk. All opinions my own.
Anthony DiGiovanni
Thanks for clarifying!
My worry is, these heuristics are only “time-tested” in that they’ve (apparently) had good consequences relative to not-super-large-scale goals, in not-super-unfamiliar contexts. Many people working on the kinds of problems you’re pointing at do so for impartial altruistic reasons. So I suspect that that track record is only weak evidence that such heuristics will work relative to their goals, in unfamiliar contexts — like, say, ASI takeoff.
“Sharing scientific knowledge freely” contributed to the industrial revolution, for example, without which factory farming wouldn’t have happened, nor the development of technologies that pose x-risks. Granted, the implications of these examples are debatable. But they seem prima facie concerning enough to me that calling big-world heuristics “robust to uncertainty” sounds way too strong (for impartial altruistic goals at least).
depend on your specific causal argument being correct & your conceptualizations apt to reality. Often it’s a fairly conjunctive argument full of lots of beliefs about what people will do
I don’t think the extrapolation from “good on local scales / familiar contexts” to “good on the cosmic scale / very unfamiliar contexts” escapes really conjunctive causal mechanisms. It just hides them at a higher level of abstraction. When making that extrapolation, we’re predicting that the same sorts of mechanisms that led to good consequences on the local scales will be at play on the cosmic scale. So we’re baking in a lot of implicit beliefs about those mechanisms.
Taboo “equilibrium”: Less confused frames for research on AI bargaining
I mostly had in mind something like this, basically reducing induction to a norm of non-arbitrariness. This isn’t fully satisfying though, see e.g. my comment on the post. I feel pretty confused about induction.
when you are working on a hard technical problem that’s a long way from solved, and could have both benefits and harms to humanity if it were solved, you don’t really have to worry about “what if we succeed and it’s bad?” or “what if this information gets into the hands of the wrong people?” You’re a long way from winning, worrying about winning too much is silly, and heuristics like “do a good job on important work” and “share scientific knowledge freely” are more trustworthy than galaxy-brained arguments for why in this particular case the utilitarian calculus works out the other way.
I’m not sure why exactly you believe this, can you say more? (Calling the arguments for backfire “galaxy-brained” kinda unduly stacks the deck, I think.)
What do you think about an alternative (/complementary) approach like: Take some conceptual questions that have been well-studied but aren’t public, and get really thorough rubrics of objective good and bad reasoning patterns that show up in answers to those questions?
What do you think of my arguments under “If they aren’t capable of C”, in the OP?
one should still support it under precise credences?
I’m saying that Elga’s argument doesn’t tell us to have precise credences in the first place. It only tells us “you should commit to act in a way that avoids sure losses”.
(I generally find LLM-written philosophy critiques overstate various things, which I don’t think are worth the time to engage with. Just briefly replying here to the substantive points.)
“The imprecise credence itself doesn’t guide the action w.r.t. the second bet” does not imply “you shouldn’t have imprecise credences in general”. Elga’s argument doesn’t tell us at all that we should, say, “go with our best guess” about altruistic interventions.
I don’t understand the critique. If C does “dodge the objection”, that’s an important advantage of C! And a commitment is a substantively different move from choosing what to do about bet B once bet A has already been dealt with, so it’s not an ad hoc dodge.
Again, see (1). Whether the commitment (without the imprecise credences) suffices to avoid this particular Dutch book tells us nothing about whether, when we’re making altruistic decisions, we should adopt precise credences and EV-max w.r.t. them. Claude’s point about “adds nothing a sharp state lacks” ignores all the positive epistemic motivations for imprecision I’ve argued for in the sequence.
Following (3): There are positive epistemic motivations for imprecise credences. So yes, Elga is asking the impreciser to override what they consider locally (epistemically) rational.
I’m proposing binding commitments, not plans. As I say, there’s no choice to be made after committing to C and rejecting bet A. So I reject the claim that the commitment response requires “imposing different requirements on choices that are identical in all relevant respects”.
Summary of why I don’t buy diachronic Dutch book arguments
(Using Elga’s Dutch book against imprecise credences as an example, where the agent faces a sequence of two bets A and B.)
Consider the action C: “Commit to accept B if you first reject A.” By “commit”, I mean literally lock your future self out of the “reject B” action (if you reject A), or make the “reject B” option so costly that it would never be locally rational to take it. This is stronger than Elga’s “planning”.
If the agent is capable of C:
Then, after choosing to do C and whether to accept A, there is no “choice” left to be made. The commitment just determines what will happen.
And doing C plus either accepting or rejecting A is permissible, according to any plausible choice rule applied to imprecise credences. → No sure loss.
If they aren’t capable of C:
Then I don’t see in what sense it’s irrational for the agent to have imprecise credences + locally choose whether to accept each bet based on their credences at the time of each choice.
Why? Because:
By hypothesis, they can’t force their future self to deviate from what they’d consider locally rational.
But that’s exactly what Elga’s argument tells the agent to do! [1] It says:
“Let’s grant that you and your future self might consider imprecise credences to be locally rational. [2]
“If so, your future self (choosing whether to accept B) might choose such that overall you’ve suffered a sure loss.
“You should adopt sharp credences — i.e., make your future self’s beliefs deviate from what they’d consider locally rational. That’s a strictly better alternative that you’re capable of.”
So in order for us to say the agent’s sequence of choices is irrational, we need them to be capable of a “commitment” that seems just as strong as C. [3]
(See also the links in the table here re: money pump arguments for completeness, which have a very similar structure.)
ETA (Jul 15, 2026): See here for a summary of my take on the objection: “But following C is behaviorally the same as rejecting imprecise credences (i.e., the imprecise credences don’t do any work).”
- ↩︎
H/t Jesse Clifton for making this salient to me; not sure if he’d endorse this version of the counterargument though.
- ↩︎
Indeed, Elga gives some great intuition pumps for this in the intro of his paper!
- ↩︎
You might say, adopting a different standard of rationality (as Elga asks the impreciser to do) is more psychologically tractable than C. But one way to achieve C is to adopt the principle of resolute choice.
- 's comment on Subjective Probabilities should be Sharp by (EA Forum; 11 Jul 2026 11:11 UTC; 6 points)
- 's comment on Subjective Probabilities should be Sharp by (EA Forum; 12 Jul 2026 15:27 UTC; 0 points)
- 's comment on Subjective Probabilities should be Sharp by (EA Forum; 12 Jul 2026 14:53 UTC; -3 points)
FWIW, I don’t think Bob’s “I am uncertain about the probability” is the most plausible motivation for ranges of probabilities. (I think it’s pretty confusing for people to report ranges for this purpose, without being explicit about what the endpoints represent.)
The most plausible motivation is: “You have no reason to favor one precise probability over various others in the range. Whether you’d pick one when forced to bet is beside the point, because that doesn’t tell you what the reason is to pick one over another.” What do you think of the intuition pump in this post?
You’re not taking general claims about the epistemics of goal-directed agency and applying them to the specific case of impartially altruistic goals
I’m confused — that’s exactly what I’m doing in the argument. P2 is the general claim “there are some conditions under which we can’t compare actions’ ‘expected’ consequences”, while P3 says “for the case of impartial altruism, in our actual epistemic situation, those conditions do hold”.
It would help to hear more specific critiques of my arguments for P3, in the posts themselves (and why you find my responses to Richard weak).
Announcing the Safe Pareto Improvements (SPI) Fundamentals Program
Thanks!
So what if there’s a remaining degree of freedom for the prior—Savage’s theorem still holds, so even as you complain that you have no justification for one prior over another, you should still be acting as if you have one.
“Acting as if” I have a prior is compatible with admitting cluelessness. See here, which also discusses Dutch book arguments.
Solomonoff’s arguments that a simplicity prior will make only a finite number of mistakes in an approximately computable universe. You build an abstract model of how a reasoning style will perform, and then you justify using that reasoning style by appealing to good modeled performance
One problem with this is that Solomonoff priors are way too computationally intractable for us. So I don’t see what normative relevance these arguments have for us.
But when the arbitrariness is contained, it’s more appealing to say something like “This arbitrariness is unavoidable, and that’s okay. To worry that we’re making a ‘wrong choice’ or that any choice here needs a further step of justification is to misunderstand what’s going on. This is about expressing ourselves and doing our best, and it’s genuinely okay to be arbitrary in this way.”
When you say the arbitrariness is “unavoidable”, is this an implication of your previous paragraphs, or something else? I think the following response avoids some of the arbitrariness: “I simply don’t think the information I have warrants ‘expecting’ A to be better than B, or vice versa. So, I don’t have reason to c-prefer A or B. I’ll make my decisions on some basis other than c-preferences, or at the very least not tell myself I’m choosing A or B based on the impartial good, when that’s false.”
You can try to fix this by speaking of your expected value of your ideal guy’s expected value of the options, not of what you expect the guy to decide
Ah, that’s exactly what I meant — if we ourselves have precise, literal expected values about the ideal guy’s expected values. But in P1 I don’t want to assume we do. That’s why I talk about scare-quote “expectations”. I want to capture “whatever kind of aggregation across possible outcomes is EV-ish but is actually accessible to bounded agents”. (This is vague, but as I say in the footnote, it’s what EA consequentialists seem to actually appeal to in practice.)
And then, P1 says that in order for a c-preference to be justified, you need to “expect” that the literal EV you’d calculate if you were capable of doing so is positive. Does that clarify things?
Yep. Presumably BB’s implicit claim here is something like: “People make claims that we should do such and such thing because of FDT, but those claims don’t follow from ‘we should program an AI to follow FDT’.”
For example, here’s Richard Ngo saying we should cooperate with the values of civilizations outside our lightcone because of FDT.
Note that the real world contains many Newcomb-like problems. People do in fact go around making decisions that depend on their beliefs about other people’s decision algorithms.
I disagree, see here for why. (I think “decisions that depend on their beliefs about other people’s decision algorithms” is too weak to get you a Newcomblike structure.)
Cluelessness: Summary of the argument, why it matters, and counterarguments
Thanks Vojta!
I agree that thinking about concrete scenarios is important. But I’m not exactly sure what you have in mind here: “The hope behind this is that it would give us intuition pumps with which progress on SPIs would get easier and faster.” What’s a quick example?
Yep it’s the latter. To answer the questions, as far as I recall what we intended when writing this:
Bargaining strategies are any unilateral (potentially conditional) commitments to actions in the game.
It’s just stipulated that the bilateral meta-commitment the agents consider whether to follow is, “Use the same bargaining strategy regardless of whether the bargaining game is G0 or Gdi”. (This is analogous to how renegotiation programs use the same default program against renegotiators and non-renegotiators, except that renegotiation programs are unilateral. ETA: Tbc, in both the bilateral and unilateral case, the agents could of course use commitments that deviate from this constraint if they wanted to — but the argument we make is, no matter which commitments they consider using that deviate from this constraint, they’d be individually better off instead plugging those commitments into the meta-commitment that adheres to this constraint. Under the belief assumptions, that is.)