The final document is now referenced from my linkpost at https://www.lesswrong.com/posts/JAxgk7sBYYeK4dyw4/when-to-take-one-box-when-to-take-two-a-concrete-analysis-of. The draft for which this is the SHA-256 hash is the file newcomb-draft-2026-08-06.pdf in the repository at https://gitlab.com/radfordneal/newcomb.
There haven’t been any claims of AI generating anything similar that I’m aware of, so publishing the hash turned out to be unnecessary. But perhaps a necessary precaution in current circumstances.
Hi. Thanks for your detailed reply.
Regarding smoking lesion scenarios, my thoughts (see the Discussion section) were that both when the lesion affects desire and when it affect the decision rule used, FDT would recommend smoking, since imagining that your decision function outputs “smoke” or “don’t smoke” affects your utility only through its affect on whether you obtain the pleasure of smoking—it’s assumed that you have no concern for other people’s pleasure.
So on this view, it shouldn’t matter whether the lesion affects the decision rule used in the sense of inclining people to use FDT, or inclining them to use CDT, or inclining them to use some other decision rule that results in them smoking, or some combination of these. I assume that all you know is the correlatio between decision (however made) and cancer, but I don’t think it would matter if you had more detailed information.
If you have altruistic motivations, it’s a different problem, which also seems interesting, and in which the proportion of smokers who use FDT starts to seem relevant, so the action FDT recommends would depend on beliefs about that, and also (as you note) on any information about cancer in FDT agents in particular.
As you say, the proportion of past Participants (on whose actions the Predictor based their prediction) who used FDT could be relevant for the impersonal Newcomb problem too. This seems a bit problematical, since presumably one wants to count as “FDT agents” people who have never heard of FDT, but are consciously or unconsciously behaving as an FDT agent.
The arguments seem similar (though more complicated) in a non-fantastical standard Newcomb problem with prediction using categories. It seems so extravagent to me, though, to be imagining that so many people decided one way or another as you decide—if you decide to one-box, you’re imagining that you get $1M, but also that all those other people got $1M, which could affect the general economy, causing inflation and other bad things, so maybe two-box instead? It seems too much...