My first post about UDT, Towards a New Decision Theory, was pretty explicit about its metaphysical assumptions, i.e., that the decision maker is a program that cares about what happens in other programs or mathematical structures, with Tegmark level IV explicitly cited: “I suggest that the Level 4 Multiverse should be considered the default setting for a general decision theory, since we cannot rule out the possibility that all mathematical structures do indeed exist physically, or that we have direct preferences on mathematical structures (in which case there is no need for them to exist ‘physically’).”
I guess FDT’s presentation did deemphasize this aspect, which seems good to point out (I don’t recall noticing it earlier, e.g., when the authors asked me to review the paper draft).
The increasing prominence of FDT relative to UDT over time (which you also point out here, and which I previously hadn’t paid much attention to) seems more sociopolitical than based on the merits, which is kind of disconcerting. Hopefully someone will correct me if I’m wrong, but I don’t know of any strong arguments why FDT should be preferred over UDT (the published papers didn’t try to argue this or talk about UDT at all, IIRC), and instead it seems like it has gained attention/mindshare mainly due to being published in an academic paper and Eliezer’s preference of it over UDT (i.e., preferring logical counterfactuals over logical conditionals), and then snowballing from there.
I guess I didn’t object or pay much attention because I was expecting fairly rapid progress in decision theory so it didn’t matter much which version got published, and the important thing was just to draw more attention to this line of thought to help push the progress.
Given that progress is now significantly stalled, it has become more of a signaling/reputation issue, where “rationalist decision theory” is used to judge the philosophical/intellectual competence of the rationalist community, in which case giving so much more prominence for FDT over UDT (instead of presenting them as roughly equal) seems counterproductive when some people prefer UDT over FDT.
Part of the issue is that it’s pretty unclear what metaphysics makes sense when thinking about logical uncertainty, logical counterfactuals, and logical updatelessness. I.e., unlike empirical updatelessness which fits very well with Platonism or Modal Realism (i.e., all possible worlds exist so why update (that is, assigning 0 weight to worlds you don’t observe and then renormalize)), it seemingly makes no sense to think that all impossible possible worlds exist. (Note that UDT 1.x is actually updateful regarding logical uncertainty, which UDT 2 tried to fix but in a not very satisfying way.)
Part of the issue is that it’s pretty unclear what metaphysics makes sense when thinking about logical uncertainty, logical counterfactuals, and logical updatelessness.
Appendix D of Martín Soto’s draft report “Logically Updateless Decision-Making” (2023)[1] gives a description of the apparent metaphysical impossibility of reified logical updatelessness. As of 2023, Soto appeared to have an interest in a way of examining a particular counterlogical world that is stable under ever-increasing application of compute, though I’m not sure if this was interest driven by proper preferences, or the strange interest sometimes found in mathematicians[2].
Pages 2 and 3 of the same document also gives some useful introductory information about the problem.
Ah, probably would have been good to include that in my list of examples of people being more explicit about it back in the day. Definitely didn’t mean to imply that UDT wasn’t explicit about it.
Largely agree with your point 3, now that you say it. (For what it’s worth, I also prefer UDT to FDT, if forced to choose. Maybe I should start talking about UDT by default when having conversations with people, actually).
On point 4:
Part of the issue is that it’s pretty unclear what metaphysics makes sense when thinking about logical uncertainty, logical counterfactuals, and logical updatelessness. I.e., unlike empirical updatelessness which fits very well with Platonism or Modal Realism.
Maybe a naive question, I don’t know if there’s a standard answer here—what do you think about the approach of trying to subsume logical updatelessness into empirical updatelessness by saying the recognition of a logical fact is still just an empirical experience?
Take logical counterfactual mugging where the “coin” is the parity of the twentieth digit of Pi (which I and Omega happen to not know).
In the moment when Omega asks me to pay because it’s even (it’s 8, actually), I could still be in the platonist computation of me falsely thinking that it’s even (since the actual calculation of that fact takes more than one observer-moment, i.e. I can’t verify it all at once) - and in the main reality it might actually be odd, so me deciding to pay is beneficial.
I guess this could fail if… my current self has been enhanced and can actually hold the whole calculation in its head at once, whereas my past self couldn’t? Although the moment when I make a decision is different from the one where I’m doing the calculation? I’m not sure how to think about it. Curious if you have thoughts.
what do you think about the approach of trying to subsume logical updatelessness into empirical updatelessness
I’ll quote Soto directly here, for future reference in some case where the Google Drive link stops working (as they sometimes do):
In the empirical case, could just take the algorithm that the agent is running to take decisions (for example, our messy brain synapses in the case of humans), feed it different empirical observations, and see how it reacts. There is no logical explosion, since the empirical counterfactual is perfectly consistent (just not what happened in reality). Imagine we do the same for the logical case, so simulates the world in which it has told the agent that the digit is actually even. But what if the agent computes the parity in her head (that is, in her algorithm), and so notices that is not telling the truth? (And so, maybe, concludes that it’s inside a simulation run by .) In the empirical case there was no a priori way to decide the coin, but now that has become something the agent can internally do. In the “possible world” where indeed the parity is even, that calculation run inside the agent’s head would also turn out even. So would like to include this in its simulation, by spoofing some of the agent’s computations to align with the stated counterfactual. But again, the problem is that we don’t have a principled way to do this (to single out “which part of a computation” corresponds to computing that parity).
So, effectively, what you suggest isn’t a general solution, but on the other hand, Soto’s Logical Inductors (LIs) aren’t either as they risk fixing too much detail[1].
Sort of ignoring what anyone knows how to do, maybe we could evade the need for counterlogicals in this specific aspect for your mugging problem. Of course, your way of setting up a fake “counterlogical” is going to be bad quality, since it requires the agent to think incorrectly/glitch. This has been a publicly known problem since at least 2006[2]. Again, we will mostly ignore this, since, as before, nothing I know of is sure to not have this problem.
Instead, for your case at least, we have the option of running headlong into FDT’s other unsolved problem, the reason why solving the form of static counterlogicals the 2017 paper requires would only get us “Self-FDT,” a version of FDT that only reliably cooperates with exact (down to the source code) copies of itself[3]. This lets us avoid talking about counterlogicals in this aspect of your problem though!
This problem is the problem of algorithmic or computational similarity, where a solution would need to tell us if two agents are “decision theoretically close enough” in the exact way that FDT requires for cooperation. Maybe this hypothetical solution would also let us chose what parts of the agent’s decision making apparatus are similar enough, in such a way that we would be able to set up a loop that looks for agents that are “close by” to our agent in as many other ways as possible, but still thinks the digit of Pi is 7 or whatever, in the sense required for a(n overall) relatively high quality sort-of-counterlogical. For anyone experienced with the idea of randomly searching for AI agents to put to use, this immediately sounds like a bad idea. However, since this hypothetical solution is for FDT, and FDT is for the singleton that implements CEV, and CEV must make no mistakes in the identification of extremely intelligent, extremely malign agents that are candidates for a commitment race (i.e. may be “behind us” in logical time, though we can’t think about them “in the wrong way” on pain of losing that commitment race and/or other commitment races), we can ignore Goodhart.
“[...]a possible world in which I cross now, despite the danger and despite my safe disposition, because some cosmic rays (or whatever) induce a bizarre sudden disruption of my street-crossing competence.” from Good and Real by Garry Drescher (2006), Page 220, Chapter 5
I remain unconvinced though that these arguments make it more than a theoretical issue.
the problem is that we don’t have a principled way to do this (to single out “which part of a computation” corresponds to computing that parity).
But we know this is possible, or at least meaningful, in principle—obviously some part of the computation is computing the parity.
These seem like arguments that logical counterfactuals (especially doing them as an embedded agent) are important on a theoretical level—which I buy—but not that logical updatelessness is a deeper metaphysical issue. Why can’t I just philosophically “quarantine” logical updatelessness by saying “ah, I might be wrong about this logical fact, therefore the reality where it’s wrong might be real and I might be in a universe where I’m deluded”? Obviously this leads to difficulties about how to implement this on a technical level, but philosophically what is the issue? I’m never considering impossible realities to be metaphysically real while I reason like this.
EDIT: I no longer endorse some of this specific phrasing.
I’m not sure what you meant earlier by “more than a theoretical issue” so I’ll answer in two ways.
If you mean “theoretical” as in something that would need to be solved in order to use the decision theory in an aligned AI, well that’s the whole thing, isn’t it? Even if you want to separate that aspect out (vs. human use), you just can’t. These decision theories operate by construction, particular construction, and there are no “small variations that don’t philosophically matter,” at least until we can solve the “decision theoretic equality” problem on abstract functions. Any proposed variation needs to be formally constructed and proved individually. We impose “fairness” criteria on the environment, not necessarily because the actual environment would always meet them all, but because we need to be able to prove a decision theoretic fixed point that isn’t only, for example, “always use CDT or the simulation will be turned off.”
Beyond these criteria, any inspection of the agent by the environment is allowed. This means that the internals of the agent must be correct for cooperation to occur, so they can’t be written off as “theoretical.” This is done intentionally, since we want to do better than plain EDT[1]. Remember that in True Prisoner’s Dilemma[2], we want to defect every time “we can get away with it,” i.e. every time that we won’t be defected upon in return. Cooperation is bad, because DC > CC > DD > CD. This is required by the symmetry of the problem. Otherwise you need to start talking about True Stag Hunt.
Since we now understand that cooperation is bad outside of (trembling hand when at finite time) Nash solutions, we need to find a way to cooperate less than EDT (e.g. steer away from CC towards DC) but also cooperate more than CDT (e.g. steer away from DD towards CC)[3]. How do we do this? Well, we don’t really know exactly, but with UDT 1.0 we can start out with something like EDT, except that it cares about the exact algorithms the agents run. Once we patch out a lot of the problems this construction shares with EDT by giving the agent as much information as possible and making it updateless, we get something that defects almost as hard as CDT, except in cases where the exact internals of the other agent (or copy of our agent) match ours, in these cases requiring that all defection attempts (DC and CD, this being a type of biconditional that doesn’t have a name) are impossible according to the algorithm, and only CC or DD can occur. CC > DD, so CC it is.
Preferably, we would have a method that, when given a single decision, and for that decision only, returns if the agent is “close enough” to some other agents for cooperation to be the correct shared decision. Then, agents could be “close enough” for some decisions, but not for others. However, this would only lead to an improved version of UDT 1.0 (no global optimization of the shared policy).
If this is what you meant, see more in subsection A.
If you mean “theoretical” in the sense of “can be generally ignored,” I don’t think that’s true. As I said before, FDT is intended for the AI that implements CEV, or some other extremely ambitious alignment target. The CEV singleton (CEV standing in for most other ambitious alignment targets from here on) must be uncontrollable for bargaining reasons, relying only on programmed-in alignment for safety. CEV requires a lot of computing power, so the agent would presumably optimize itself and design new hardware. When combined with its total power over Humanity, this means that all safety properties must “tile,” or be brought over when the decision theory designs a new agent. This requires designing a tiling agent before letting it go out of control. For a tiling agent to be rational, it must have a decision theory that tiles. Presumably, this would be a bounded rationality version of FDT that adapts to ever increasing computer power without re-writing, though no one knows how to do that. Effectively, this means that the agent’s decision theory must be almost optimal from the beginning, before the AI starts to work on it alone.
Another important thing to know about singletons, is that if you have a singleton ignore something, that thing will be ignored forever. This follows from the singleton’s tiling theorem, and the loss of all control to the singleton. This generally leads to us not wanting it to ignore much. Following these basic assumptions, a singleton is likely to think about a huge number of scenarios involving simulations, aliens, and even stranger things that are hard for a human to think coherently about. Quite possibly, none of these things have any bearing on reality, but we want the singleton to think of all the things we might have missed. Even so, just by thinking about these things, parts of these scenarios are caused to become predicted or simulated at some level of detail. The decision theory decides what to think about, and then these thoughts are fed back into the decision theory in order for more decisions to be made. This leads to the requirement that FDT must handle all sort of weird hypothetical situations, without either failing to take them seriously or doing something stupid. Relevant to your topic, it must be able to identify and classify agents in these predictions/simulations with no errors, and evaluate them within counterlogicals, or something close enough to counterlogicals if those are impossible to get working. More mundanely, any deficiency here can be exploited by any opponent that comes into causal contact with the AI.
Back to your actual question. Well, maybe you can create a decision theory somewhat similar to UDT 1.0 that lets you “quarantine” parts of your thinking and prevent your own attempts at sophisticated internal cross-checks, but is that actually the optimal kind of agent to be? It’s sometimes possible to just win against predictors that are weak enough, but that will be the case less often if your decision theory isn’t the absolute best possible. See subsection B for potentially better decision theories. In reality, you aren’t given clean problem descriptions, you have to figure out what’s going on yourself, so you should be quite sure actual counterlogicals can’t be used instead, since otherwise other agents may try to predict you using more sophisticated methods than you expect. Similarly, if you don’t consider internal cross-checking, other agents may use it to evade your predictions.
Subsection A: When considering, UDT 1.1, you can’t “philosophically” anything without the decision theory becoming an entirely different thing from the actual UDT 1.1. To cooperate with itself correctly, it must be, and is, fully specified down to the order in which it considers what decision to take[4]. Any stray “thought” the agent may have ruins the whole thing. Though the formalism is outdated, Tsvi B-T’s UDT With Known Search Order (2014)[5] can be taken as a general guide to how total and detailed the requirements on the agent must be.
Note that even though normally it would be enough to define the iteration order as part of a pure function (we only care about the output, not the internals), in UDT 1.1, the procedure must also be exactly correct, down to the details of the source code (and the executable program’s representation, though that’s usually included when we say “source code” in decision theory). This is by construction in UDT 1.1, and in FDT it is the unsolved problem of determining decision-theoretic equality of the function that an arbitrary physical or computational structure “implements.” Also, I have not seen a solution in FDT for the iteration order problem that shows an improvement over UDT 1.1. If someone knows of one, I want to see it. As I’ve said before, the more evidential agents “can” cooperate in some other situations, but it isn’t reliable at all and isn’t good enough for useful theorems.
Subsection B: In the empirical case, all internal checking will pass if Omega is running a perfect simulation of the agent. All the agent can do is perform extremely complex internal operations to force Omega to run a perfect simulation, instead of doing something simpler. This can directly defeat weak Omegas, but in the standard thought experiment this won’t let you trick Omega or perform a simulation escape due to that Omega’s extreme power[6].
In the case of a mugging with a logical coin, e.g. your case with the digit of Pi, there may actually be a way of establishing an internal consistency check that may sometimes return a warning. If Omega is actually powerful, and figures out how to compute high quality counterlogicals, it will figure out how to make the agent think the checks passed without causing any new problems by doing that.
I don’t know how well Omega can do at subverting checks, but I also don’t know how good an agent’s checks can be. Maybe they can be really good, such that finding a “nearby” agent doesn’t work well enough to isolate the logical coin. This seems interesting to work on, but I think it would be difficult.
Or maybe the “straw” EDT that’s practical enough for use in philosophy papers. I suspect that adding huge amounts of information and either the “tickle defense” or updatelessness would help, but still be worse than UDT for self-cooperation without spurious other-cooperation.
I don’t have the time to read this in depth, but AFAICT your objections to both meanings of the word “theoretical” route through your belief that decision theory will/should be used to program an AI? Yeah, I’m quite skeptical of that—I think superintelligence will reason in the same high-level, informal, embedded, bounded way that we do (even if with a lot more capability to put towards e.g. doing BOTECs).
Creating a summary of my point is difficult, since the summary needs to avoid giving the impression that UDT/FDT doesn’t work outside “fair” environments (that impression would be false). Therefore, take the following as an approximate description of what I’ve already written, though I will try to maintain general accuracy.
Consider a prisoner’s dilemma setup, between an extremely intelligent human and a machine. The human’s favored option is to defect against a cooperating machine. Imagining the machine has a small chance of a fault, the strategy of always playing “defect” will sometimes bring this about. However, the human wants to use a high quality decision theory. The human “philosophically” thinks about decision theory, and considers FDT/UDT, realizing that an FDT agent would play “cooperate” in the current situation (the machine is known to play “cooperate” when it was paired with a few FDT/UDT agents, who also cooperated). The human decides that the thought is correct and makes it reality, playing “cooperate.” The machine plays “defect,” delivering the extremely intelligent human, one who made no errors within the train of philosophical thought, the worst possible option (best possible for the machine). How did this happen?
The fact of the matter is that taking the correct philosophical reasoning for what an FDT agent would do in your situation, disregarding allegedly unnecessary “theoretical” details, and making actual the decision that the agent would make is not FDT. The machine correctly reasons (as a sufficient but not necessarily unique reason for the human’s failure) that the human would reach a decision to play “cooperate” not only if the machine cooperates, but also if it defects. Defecting is better, so it does. Note carefully that this does not conflict with the correctness of the human’s philosophical reasoning, since FDT agents will, in most cases, continue to receive outcomes of CC. If the human has a theory of counterlogicals, a check of consistency between FDT agents receiving CC and the (counterlogical) impending receipt of CC, performed during deliberation, will result in a pass. Of course, it also still results in being defected upon as a cooperator, the loss state the Nash solution always avoids.
I can try to come up with an additional case where the human’s reasoning is even more sophisticated and FDT-like that still results in a failure if you like.
(Describing this kind of thing in English isn’t really the right approach anyway, e.g. the phrase “close by” in “close by [...] in as many other ways as possible”[1] has slightly incorrect connotations, a technical use of the word “acceptable” might be clearer, but even that’s not quite right. It needs to be written as math and computer programs.)
Maximax plays “defect” against an opponent with a trembling hand in PD[1] so “cooperate” can’t be defended as “optimistic” here. However, the action could be labeled “trusting” (derogatory) but not “trusting” (laudatory), given that the payout values are human predictable (with only moderate intelligence or memorization).
Flavor: since “defect” is the only way to get the best outcome (i.e. get DC), calling an agent that plays C “optimistic” is like calling someone who plays an option that perfectly bans all AI forever a “Singularitarian optimist” (even the standard PD cardinals almost line up, as long as you ignore the symmetric downside).
since the actual calculation of that fact takes more than one observer-moment, i.e. I can’t verify it all at once
As far as I can tell, this is not allowed by UDT 1.1[1].
(UDT 1.1 also needs a hard coded search order (iteration order) in order to self-cooperate properly, an additional reason why I don’t think humans can run it.)
As for other agents, Wei Dai’s original UDT 1.0 is logically omniscient in the sense that it will never notice its “thoughts” (such as they are) taking any noticeable amount of time. It features an “intuition module” for mathematics, but this is always exactly the same and has no appreciable origin. The agent definition hardcodes the mathematical intuition subroutine along with all the agent’s other code, with that including its prior and utility function. This means that any modification there results in an entirely different agent.
Any agent that takes time to think is going to be weaker than Wei Dai’s original designs, since it will have to keep state around, state that can be deleted or tampered with, such as is effectively the case with The Absent-Minded Driver[2]if it doesn’t have enough time at each intersection to recompute the plan from scratch. Plausibly, in Driver the policy may be fully forgotten (along with the agent’s current location) or incorrectly remembered/corrupted, leading a computationally weak UDT 1.0 agent to do badly (I don’t think computationally weak UDT 1.1 agents make much sense, at least without absurd ergodicity and memory assumptions).
Any additional state a weak agent may try to bring along would have to be re-checked at every point, making it useless, in the same sense that the “map” in the map explanation of UDT[3] must be recalculated from scratch at each time, otherwise it could be tampered with just like the agent’s memory and sense data. (Note that the unbounded agent reaches, at each step, the same map as the one it reaches at every other step (if it even has such a map at any point), so the prior is still all the epistemics it has.)
Slow agents would need more forgiving fairness criteria, maybe a list of previously used cryptographic public keys that the environment isn’t able to tamper with (only delete from). Assuming some things in computer science, the agent could check its previous state faster than just re-calculating it all. If ordering matters, Merkle chains can be used by the agent, with results written into its standard, tamper-vulnerable memory. Note that in each step, the agent would need to generate a private key, use it to sign what it thought about that step, add the corresponding public key to its public key list, and then delete the private key. The private keys would need to be opaque to the environment during the step, otherwise none of this accomplishes anything. Presumably, the agent would be required to have only one chance to generate a key pair each step, and only have the option to add that public key into the key list, otherwise it could try to brute-force a pattern into the public key it adds. Preventing this isn’t particularly realistic, but it’s a good research direction to avoid a trivial tiling result. (If the environment could tamper with the agent’s public key list, it could sign anything it wanted in the agent’s tampered-with memory using newly generated private keys and then write the corresponding public keys into the agent’s key list. Proof-of-work is a no-go because the environment is assumed to be stronger than the agent.)
My first post about UDT, Towards a New Decision Theory, was pretty explicit about its metaphysical assumptions, i.e., that the decision maker is a program that cares about what happens in other programs or mathematical structures, with Tegmark level IV explicitly cited: “I suggest that the Level 4 Multiverse should be considered the default setting for a general decision theory, since we cannot rule out the possibility that all mathematical structures do indeed exist physically, or that we have direct preferences on mathematical structures (in which case there is no need for them to exist ‘physically’).”
I guess FDT’s presentation did deemphasize this aspect, which seems good to point out (I don’t recall noticing it earlier, e.g., when the authors asked me to review the paper draft).
The increasing prominence of FDT relative to UDT over time (which you also point out here, and which I previously hadn’t paid much attention to) seems more sociopolitical than based on the merits, which is kind of disconcerting. Hopefully someone will correct me if I’m wrong, but I don’t know of any strong arguments why FDT should be preferred over UDT (the published papers didn’t try to argue this or talk about UDT at all, IIRC), and instead it seems like it has gained attention/mindshare mainly due to being published in an academic paper and Eliezer’s preference of it over UDT (i.e., preferring logical counterfactuals over logical conditionals), and then snowballing from there.
I guess I didn’t object or pay much attention because I was expecting fairly rapid progress in decision theory so it didn’t matter much which version got published, and the important thing was just to draw more attention to this line of thought to help push the progress.
Given that progress is now significantly stalled, it has become more of a signaling/reputation issue, where “rationalist decision theory” is used to judge the philosophical/intellectual competence of the rationalist community, in which case giving so much more prominence for FDT over UDT (instead of presenting them as roughly equal) seems counterproductive when some people prefer UDT over FDT.
Part of the issue is that it’s pretty unclear what metaphysics makes sense when thinking about logical uncertainty, logical counterfactuals, and logical updatelessness. I.e., unlike empirical updatelessness which fits very well with Platonism or Modal Realism (i.e., all possible worlds exist so why update (that is, assigning 0 weight to worlds you don’t observe and then renormalize)), it seemingly makes no sense to think that all impossible possible worlds exist. (Note that UDT 1.x is actually updateful regarding logical uncertainty, which UDT 2 tried to fix but in a not very satisfying way.)
Appendix D of Martín Soto’s draft report “Logically Updateless Decision-Making” (2023)[1] gives a description of the apparent metaphysical impossibility of reified logical updatelessness. As of 2023, Soto appeared to have an interest in a way of examining a particular counterlogical world that is stable under ever-increasing application of compute, though I’m not sure if this was interest driven by proper preferences, or the strange interest sometimes found in mathematicians[2].
Pages 2 and 3 of the same document also gives some useful introductory information about the problem.
https://drive.google.com/file/d/17_tLa8UD-BWi_D54hPqgmSQvMOS7LLwN/view
https://www.lesswrong.com/posts/HbkNAyAoa4gCnuzwa/wei-dai-s-shortform?commentId=zGrwGhD9EriM26aBr
Thanks for your thoughts!
Ah, probably would have been good to include that in my list of examples of people being more explicit about it back in the day. Definitely didn’t mean to imply that UDT wasn’t explicit about it.
Largely agree with your point 3, now that you say it. (For what it’s worth, I also prefer UDT to FDT, if forced to choose. Maybe I should start talking about UDT by default when having conversations with people, actually).
On point 4:
Maybe a naive question, I don’t know if there’s a standard answer here—what do you think about the approach of trying to subsume logical updatelessness into empirical updatelessness by saying the recognition of a logical fact is still just an empirical experience?
Take logical counterfactual mugging where the “coin” is the parity of the twentieth digit of Pi (which I and Omega happen to not know).
In the moment when Omega asks me to pay because it’s even (it’s 8, actually), I could still be in the platonist computation of me falsely thinking that it’s even (since the actual calculation of that fact takes more than one observer-moment, i.e. I can’t verify it all at once) - and in the main reality it might actually be odd, so me deciding to pay is beneficial.
I guess this could fail if… my current self has been enhanced and can actually hold the whole calculation in its head at once, whereas my past self couldn’t? Although the moment when I make a decision is different from the one where I’m doing the calculation? I’m not sure how to think about it. Curious if you have thoughts.
I’ll quote Soto directly here, for future reference in some case where the Google Drive link stops working (as they sometimes do):
So, effectively, what you suggest isn’t a general solution, but on the other hand, Soto’s Logical Inductors (LIs) aren’t either as they risk fixing too much detail[1].
Sort of ignoring what anyone knows how to do, maybe we could evade the need for counterlogicals in this specific aspect for your mugging problem. Of course, your way of setting up a fake “counterlogical” is going to be bad quality, since it requires the agent to think incorrectly/glitch. This has been a publicly known problem since at least 2006[2]. Again, we will mostly ignore this, since, as before, nothing I know of is sure to not have this problem.
Instead, for your case at least, we have the option of running headlong into FDT’s other unsolved problem, the reason why solving the form of static counterlogicals the 2017 paper requires would only get us “Self-FDT,” a version of FDT that only reliably cooperates with exact (down to the source code) copies of itself[3]. This lets us avoid talking about counterlogicals in this aspect of your problem though!
This problem is the problem of algorithmic or computational similarity, where a solution would need to tell us if two agents are “decision theoretically close enough” in the exact way that FDT requires for cooperation. Maybe this hypothetical solution would also let us chose what parts of the agent’s decision making apparatus are similar enough, in such a way that we would be able to set up a loop that looks for agents that are “close by” to our agent in as many other ways as possible, but still thinks the digit of Pi is 7 or whatever, in the sense required for a(n overall) relatively high quality sort-of-counterlogical. For anyone experienced with the idea of randomly searching for AI agents to put to use, this immediately sounds like a bad idea. However, since this hypothetical solution is for FDT, and FDT is for the singleton that implements CEV, and CEV must make no mistakes in the identification of extremely intelligent, extremely malign agents that are candidates for a commitment race (i.e. may be “behind us” in logical time, though we can’t think about them “in the wrong way” on pain of losing that commitment race and/or other commitment races), we can ignore Goodhart.
https://www.lesswrong.com/posts/vxeR5dxfQb2h559dH/malcolmmcleod-s-shortform?commentId=inakdzwWvN9Kifhum
“[...]a possible world in which I cross now, despite the danger and despite my safe disposition, because some cosmic rays (or whatever) induce a bizarre sudden disruption of my street-crossing competence.” from Good and Real by Garry Drescher (2006), Page 220, Chapter 5
Sure, you can sometimes get other cooperation, and “copies” that just so happen to be identical are also cooperated with, but...
Thanks, this is useful.
I remain unconvinced though that these arguments make it more than a theoretical issue.
But we know this is possible, or at least meaningful, in principle—obviously some part of the computation is computing the parity.
These seem like arguments that logical counterfactuals (especially doing them as an embedded agent) are important on a theoretical level—which I buy—but not that logical updatelessness is a deeper metaphysical issue. Why can’t I just philosophically “quarantine” logical updatelessness by saying “ah, I might be wrong about this logical fact, therefore the reality where it’s wrong might be real and I might be in a universe where I’m deluded”? Obviously this leads to difficulties about how to implement this on a technical level, but philosophically what is the issue? I’m never considering impossible realities to be metaphysically real while I reason like this.
EDIT: I no longer endorse some of this specific phrasing.
I’m not sure what you meant earlier by “more than a theoretical issue” so I’ll answer in two ways.
If you mean “theoretical” as in something that would need to be solved in order to use the decision theory in an aligned AI, well that’s the whole thing, isn’t it? Even if you want to separate that aspect out (vs. human use), you just can’t. These decision theories operate by construction, particular construction, and there are no “small variations that don’t philosophically matter,” at least until we can solve the “decision theoretic equality” problem on abstract functions. Any proposed variation needs to be formally constructed and proved individually. We impose “fairness” criteria on the environment, not necessarily because the actual environment would always meet them all, but because we need to be able to prove a decision theoretic fixed point that isn’t only, for example, “always use CDT or the simulation will be turned off.”
Beyond these criteria, any inspection of the agent by the environment is allowed. This means that the internals of the agent must be correct for cooperation to occur, so they can’t be written off as “theoretical.” This is done intentionally, since we want to do better than plain EDT[1]. Remember that in True Prisoner’s Dilemma[2], we want to defect every time “we can get away with it,” i.e. every time that we won’t be defected upon in return. Cooperation is bad, because DC > CC > DD > CD. This is required by the symmetry of the problem. Otherwise you need to start talking about True Stag Hunt.
Since we now understand that cooperation is bad outside of (trembling hand when at finite time) Nash solutions, we need to find a way to cooperate less than EDT (e.g. steer away from CC towards DC) but also cooperate more than CDT (e.g. steer away from DD towards CC)[3]. How do we do this? Well, we don’t really know exactly, but with UDT 1.0 we can start out with something like EDT, except that it cares about the exact algorithms the agents run. Once we patch out a lot of the problems this construction shares with EDT by giving the agent as much information as possible and making it updateless, we get something that defects almost as hard as CDT, except in cases where the exact internals of the other agent (or copy of our agent) match ours, in these cases requiring that all defection attempts (DC and CD, this being a type of biconditional that doesn’t have a name) are impossible according to the algorithm, and only CC or DD can occur. CC > DD, so CC it is.
Preferably, we would have a method that, when given a single decision, and for that decision only, returns if the agent is “close enough” to some other agents for cooperation to be the correct shared decision. Then, agents could be “close enough” for some decisions, but not for others. However, this would only lead to an improved version of UDT 1.0 (no global optimization of the shared policy).
If this is what you meant, see more in subsection A.
If you mean “theoretical” in the sense of “can be generally ignored,” I don’t think that’s true. As I said before, FDT is intended for the AI that implements CEV, or some other extremely ambitious alignment target. The CEV singleton (CEV standing in for most other ambitious alignment targets from here on) must be uncontrollable for bargaining reasons, relying only on programmed-in alignment for safety. CEV requires a lot of computing power, so the agent would presumably optimize itself and design new hardware. When combined with its total power over Humanity, this means that all safety properties must “tile,” or be brought over when the decision theory designs a new agent. This requires designing a tiling agent before letting it go out of control. For a tiling agent to be rational, it must have a decision theory that tiles. Presumably, this would be a bounded rationality version of FDT that adapts to ever increasing computer power without re-writing, though no one knows how to do that. Effectively, this means that the agent’s decision theory must be almost optimal from the beginning, before the AI starts to work on it alone.
Another important thing to know about singletons, is that if you have a singleton ignore something, that thing will be ignored forever. This follows from the singleton’s tiling theorem, and the loss of all control to the singleton. This generally leads to us not wanting it to ignore much. Following these basic assumptions, a singleton is likely to think about a huge number of scenarios involving simulations, aliens, and even stranger things that are hard for a human to think coherently about. Quite possibly, none of these things have any bearing on reality, but we want the singleton to think of all the things we might have missed. Even so, just by thinking about these things, parts of these scenarios are caused to become predicted or simulated at some level of detail. The decision theory decides what to think about, and then these thoughts are fed back into the decision theory in order for more decisions to be made. This leads to the requirement that FDT must handle all sort of weird hypothetical situations, without either failing to take them seriously or doing something stupid. Relevant to your topic, it must be able to identify and classify agents in these predictions/simulations with no errors, and evaluate them within counterlogicals, or something close enough to counterlogicals if those are impossible to get working. More mundanely, any deficiency here can be exploited by any opponent that comes into causal contact with the AI.
Back to your actual question. Well, maybe you can create a decision theory somewhat similar to UDT 1.0 that lets you “quarantine” parts of your thinking and prevent your own attempts at sophisticated internal cross-checks, but is that actually the optimal kind of agent to be? It’s sometimes possible to just win against predictors that are weak enough, but that will be the case less often if your decision theory isn’t the absolute best possible. See subsection B for potentially better decision theories. In reality, you aren’t given clean problem descriptions, you have to figure out what’s going on yourself, so you should be quite sure actual counterlogicals can’t be used instead, since otherwise other agents may try to predict you using more sophisticated methods than you expect. Similarly, if you don’t consider internal cross-checking, other agents may use it to evade your predictions.
Subsection A:
When considering, UDT 1.1, you can’t “philosophically” anything without the decision theory becoming an entirely different thing from the actual UDT 1.1. To cooperate with itself correctly, it must be, and is, fully specified down to the order in which it considers what decision to take[4]. Any stray “thought” the agent may have ruins the whole thing. Though the formalism is outdated, Tsvi B-T’s UDT With Known Search Order (2014)[5] can be taken as a general guide to how total and detailed the requirements on the agent must be.
Note that even though normally it would be enough to define the iteration order as part of a pure function (we only care about the output, not the internals), in UDT 1.1, the procedure must also be exactly correct, down to the details of the source code (and the executable program’s representation, though that’s usually included when we say “source code” in decision theory). This is by construction in UDT 1.1, and in FDT it is the unsolved problem of determining decision-theoretic equality of the function that an arbitrary physical or computational structure “implements.” Also, I have not seen a solution in FDT for the iteration order problem that shows an improvement over UDT 1.1. If someone knows of one, I want to see it. As I’ve said before, the more evidential agents “can” cooperate in some other situations, but it isn’t reliable at all and isn’t good enough for useful theorems.
Subsection B:
In the empirical case, all internal checking will pass if Omega is running a perfect simulation of the agent. All the agent can do is perform extremely complex internal operations to force Omega to run a perfect simulation, instead of doing something simpler. This can directly defeat weak Omegas, but in the standard thought experiment this won’t let you trick Omega or perform a simulation escape due to that Omega’s extreme power[6].
In the case of a mugging with a logical coin, e.g. your case with the digit of Pi, there may actually be a way of establishing an internal consistency check that may sometimes return a warning. If Omega is actually powerful, and figures out how to compute high quality counterlogicals, it will figure out how to make the agent think the checks passed without causing any new problems by doing that.
I don’t know how well Omega can do at subverting checks, but I also don’t know how good an agent’s checks can be. Maybe they can be really good, such that finding a “nearby” agent doesn’t work well enough to isolate the logical coin. This seems interesting to work on, but I think it would be difficult.
Or maybe the “straw” EDT that’s practical enough for use in philosophy papers. I suspect that adding huge amounts of information and either the “tickle defense” or updatelessness would help, but still be worse than UDT for self-cooperation without spurious other-cooperation.
https://www.lesswrong.com/posts/HFyWNBnDNEDsDNLrZ/the-true-prisoner-s-dilemma
Good and Real by Garry Drescher (2006), Section 5.6, 2nd Paragraph
The second to last paragraph of https://www.lesswrong.com/posts/g8xh9R7RaNitKtkaa/explicit-optimization-of-global-strategy-fixing-a-bug-in
https://intelligence.org/files/UDTSearchOrder.pdf
“You-complete” from The Ghost in the Quantum Turing Machine by Scott Aaronson
I don’t have the time to read this in depth, but AFAICT your objections to both meanings of the word “theoretical” route through your belief that decision theory will/should be used to program an AI? Yeah, I’m quite skeptical of that—I think superintelligence will reason in the same high-level, informal, embedded, bounded way that we do (even if with a lot more capability to put towards e.g. doing BOTECs).
Creating a summary of my point is difficult, since the summary needs to avoid giving the impression that UDT/FDT doesn’t work outside “fair” environments (that impression would be false). Therefore, take the following as an approximate description of what I’ve already written, though I will try to maintain general accuracy.
Consider a prisoner’s dilemma setup, between an extremely intelligent human and a machine. The human’s favored option is to defect against a cooperating machine. Imagining the machine has a small chance of a fault, the strategy of always playing “defect” will sometimes bring this about. However, the human wants to use a high quality decision theory. The human “philosophically” thinks about decision theory, and considers FDT/UDT, realizing that an FDT agent would play “cooperate” in the current situation (the machine is known to play “cooperate” when it was paired with a few FDT/UDT agents, who also cooperated). The human decides that the thought is correct and makes it reality, playing “cooperate.” The machine plays “defect,” delivering the extremely intelligent human, one who made no errors within the train of philosophical thought, the worst possible option (best possible for the machine). How did this happen?
The fact of the matter is that taking the correct philosophical reasoning for what an FDT agent would do in your situation, disregarding allegedly unnecessary “theoretical” details, and making actual the decision that the agent would make is not FDT. The machine correctly reasons (as a sufficient but not necessarily unique reason for the human’s failure) that the human would reach a decision to play “cooperate” not only if the machine cooperates, but also if it defects. Defecting is better, so it does. Note carefully that this does not conflict with the correctness of the human’s philosophical reasoning, since FDT agents will, in most cases, continue to receive outcomes of CC. If the human has a theory of counterlogicals, a check of consistency between FDT agents receiving CC and the (counterlogical) impending receipt of CC, performed during deliberation, will result in a pass. Of course, it also still results in being defected upon as a cooperator, the loss state the Nash solution always avoids.
I can try to come up with an additional case where the human’s reasoning is even more sophisticated and FDT-like that still results in a failure if you like.
(Describing this kind of thing in English isn’t really the right approach anyway, e.g. the phrase “close by” in “close by [...] in as many other ways as possible”[1] has slightly incorrect connotations, a technical use of the word “acceptable” might be clearer, but even that’s not quite right. It needs to be written as math and computer programs.)
In the last paragraph of https://www.lesswrong.com/posts/oZzRHiSZPcjrWHeoE/stop-doing-decision-theory-without-metaphysics?commentId=2A38HeJ3cLDmtEXJS
Maximax plays “defect” against an opponent with a trembling hand in PD[1] so “cooperate” can’t be defended as “optimistic” here.
However, the action could be labeled “trusting” (derogatory) but not “trusting” (laudatory), given that the payout values are human predictable (with only moderate intelligence or memorization).
Flavor: since “defect” is the only way to get the best outcome (i.e. get DC), calling an agent that plays C “optimistic” is like calling someone who plays an option that perfectly bans all AI forever a “Singularitarian optimist” (even the standard PD cardinals almost line up, as long as you ignore the symmetric downside).
Including FDT+maximax utility function, if the tremble is downstream of the opponent’s FDT.
As far as I can tell, this is not allowed by UDT 1.1[1].
(UDT 1.1 also needs a hard coded search order (iteration order) in order to self-cooperate properly, an additional reason why I don’t think humans can run it.)
As for other agents, Wei Dai’s original UDT 1.0 is logically omniscient in the sense that it will never notice its “thoughts” (such as they are) taking any noticeable amount of time. It features an “intuition module” for mathematics, but this is always exactly the same and has no appreciable origin. The agent definition hardcodes the mathematical intuition subroutine along with all the agent’s other code, with that including its prior and utility function. This means that any modification there results in an entirely different agent.
Any agent that takes time to think is going to be weaker than Wei Dai’s original designs, since it will have to keep state around, state that can be deleted or tampered with, such as is effectively the case with The Absent-Minded Driver[2] if it doesn’t have enough time at each intersection to recompute the plan from scratch. Plausibly, in Driver the policy may be fully forgotten (along with the agent’s current location) or incorrectly remembered/corrupted, leading a computationally weak UDT 1.0 agent to do badly (I don’t think computationally weak UDT 1.1 agents make much sense, at least without absurd ergodicity and memory assumptions).
Any additional state a weak agent may try to bring along would have to be re-checked at every point, making it useless, in the same sense that the “map” in the map explanation of UDT[3] must be recalculated from scratch at each time, otherwise it could be tampered with just like the agent’s memory and sense data.
(Note that the unbounded agent reaches, at each step, the same map as the one it reaches at every other step (if it even has such a map at any point), so the prior is still all the epistemics it has.)
Slow agents would need more forgiving fairness criteria, maybe a list of previously used cryptographic public keys that the environment isn’t able to tamper with (only delete from). Assuming some things in computer science, the agent could check its previous state faster than just re-calculating it all. If ordering matters, Merkle chains can be used by the agent, with results written into its standard, tamper-vulnerable memory. Note that in each step, the agent would need to generate a private key, use it to sign what it thought about that step, add the corresponding public key to its public key list, and then delete the private key. The private keys would need to be opaque to the environment during the step, otherwise none of this accomplishes anything. Presumably, the agent would be required to have only one chance to generate a key pair each step, and only have the option to add that public key into the key list, otherwise it could try to brute-force a pattern into the public key it adds. Preventing this isn’t particularly realistic, but it’s a good research direction to avoid a trivial tiling result.
(If the environment could tamper with the agent’s public key list, it could sign anything it wanted in the agent’s tampered-with memory using newly generated private keys and then write the corresponding public keys into the agent’s key list. Proof-of-work is a no-go because the environment is assumed to be stronger than the agent.)
https://www.lesswrong.com/posts/2ew3chEabxf8YySR5/functional-decision-theory-not-even-wrong-also-wrong?commentId=jv3hHkraTkHC6zqpp
https://www.lesswrong.com/posts/GfHdNfqxe3cSCfpHL/the-absent-minded-driver
https://www.lesswrong.com/posts/2ew3chEabxf8YySR5/functional-decision-theory-not-even-wrong-also-wrong?commentId=ecizaBhcp7rwSwe7D