If TDT/UDT/FDT was applicable to humans, oligopolies would be able to use it to coordinate price fixing without leaving any evidence. Anyone recall this point being brought up in past discussions?
See https://en.wikipedia.org/wiki/Tacit_collusion . Coordination without explicit communication is rampant in human behavior. However, that’s different from “without causality”, so it’s not necessarily tied to TDT/UDT/FDT.
Isn’t this a classic commitment race/equilibrium selection thing where it’s just unclear what ‘should’ happen if everyone involved is following one of these decision theories? The customers can coordinate to boycott anyone involved in price fixing
I think that’s true (if LDT applied to humans), but what we actually see is that oligopolies sometimes coordinate causally to price fix (occasionally leaving evidence behind in the form of secret meetings, paper trails, etc.) and consumers almost never coordinate (either causally or acausally) to boycott pricing fixing. This seems to strongly suggest that LDT does not apply to humans, otherwise oligopolies wouldn’t need to do it the legally risky way, and/or we’d see more acausally coordinated boycotts.
https://en.wikipedia.org/wiki/Buyers_club, and the far less legible “get a group together to buy wholesale” ideas are not new or uncommon. Artificial monopsony (all/most consumers acting as one) is pretty rare, but much like the other direction, doesn’t require perfection to be worthwhile.
Hmm, I can’t remember where it’s been brought up. But consumers should be able to do the same in reverse.
My thought is you end up with a situations where you have two agents bargaining, the exact solution being undefined except both agents extracting more than zero utility, with the schelling point set by shapely values (both sides getting the same marginal utility here).
To reference, here is my list of issues that fairly sophisticated reasoners get wrong when thinking about FDT.
A: The utility function you imagine using is not gerrymandered enough. For example, it’s not the decision theory’s problem if you care a large amount about a small-in-prior region but forget to encode such.
B: There is not a proper attempt at writing a bounded procedure, instead the “decision theory” as imagined consists of a human trying to guess the output of a logically omniscient, unbounded procedure.
Both you and Nate Soares don’t seem to have a problem with A, though I’ll note that certain preferences will require quite extreme utility distributions to counteract the higher weighting of nodes nearer to the root of the UDT tree. This weighting is required to rationally steer re. measure loss, so it can’t be removed.
Some things that might be worth thinking about for you and other highly sophisticated reasoners:
C: Are you trying to solve for an equilibrium where everyone is running FDT? Connecting to B, is the FDT they are running unbounded, “perfect” FDT, or some more plausible approximation? If only some people are running FDT, how can other agents tell? How can other agents tell the exact type/quality of FDT another agent is running? (At an extreme level of sophistication, what are the iteration orders of the agents? How do they avoid Löb? Do these things differ between agents?)
D: As far as I can tell, if a group of humans is guessing the output of each agent running FDT (see B), it is essentially a stage play (for practice or for demonstration (more cynically, “signaling”)). “Guessing” is not an actual, well specified procedure for decision theoretic reasoning, and it is not robust against intelligent opposition. The group must cooperate to put on a show, otherwise an agent with a more exploitative procedure can bring the whole thing down.
E: (In an attempt to avoid D) humans may try to actually run the internals of FDT (instead of just guessing at the output). Do they actually have enough compute to calculate the prior up to a point where they can legitimately win commitment races? Sure, some people are probably smart enough to establish an obdurant policy this way, and have enough resources to execute it, but there will be no cooperation and little action at all. As an example, imagine living in a hut in the middle of nowhere and raising animals. To win a commitment race, according to the internals of FDT, you must be able to calculate what your opponents will do, including what the opponents will do off policy (presumably where “off policy” varies in a huge number of ways as you attempt to reach the fixed point). These calculations must be “logically correct,” roughly following the example of the calculator, where two calculators must reach the same output (in analogy, “decision”) even if the internals are somewhat different. If you can’t do this, you’re not “winning a commitment race,” you’re avoiding other agents taken in the sense that they might be helpful agents.
F: Even if you avoid the problems given in E and “win” a commitment race against another agent, will anyone know? Does the opponent know what a commitment race is, even? If the opponent (A_O) doesn’t know or doesn’t understand, the result might be the equivalent of “straight straight” in Chicken, where A_O assumes you’ll play the equivalent of “swerve.” Maybe A_O will get mad at the entire idea of losing to someone in a commitment race and intentionally play “straight straight.”[1]
G: As far as I can tell, losing a commitment race is sometimes not that bad, especially for a physically powerful agent that doesn’t care about any multiverse ideas. At the risk of being too metaphorical, imagine agent X has won a commitment race and turned into a rock. If agent Y is physically powerful, it can still rationally play “straight,” knowing it has lost the race, if it first removes the rock from the road with dynamite. By instrumental convergence, this is presumably worse for agent Y than playing “straight” against agent X’s theoretical “swerve,” but the overall policy of agent Y may still be as good as it could be with any plausibility. As far as I can tell, you can’t win all theoretical commitment races simultaneously, while also performing actions useful to the physical world.
Actually, the difficult procedure in E can, by itself, lose you certain commitment races against certain opponents, if you take my short summary written there as a guide on how to think.
Note that to avoid losing a commitment race (by reaching an “infohazard”) to the maximum amount UDT 1.0 allows, you must be fully updateless (base your policy on the prior alone[1]) and also have a “mathematical intuition subroutine” that is somehow optimal and doesn’t (de facto) tell you too much. Intentionally making your mathematical intuition worse is an unsolved problem and is extremely fraught. However, it can’t be avoided that, in some cases, it’s possible to win commitment races by being too stupid to see some (mathematical) fact about the opponent, or how the opponent interacts with the environment[2], but I’ll leave this in the background for now.
In section E, I state that you must be able to calculate what your opponents will do. This isn’t exactly right, since updating on the nature of your opponents can lose you commitment races. This is somewhat difficult to describe in a simple way, but note that you may have opponents that were created by another opponent, with that creation depending in some way on rational decision. If you update on your current opponents, the earlier opponent may be able to win an important race.
Note that, when dealing with sophisticated opponents, infohazards can not be handled by ignoring information after you have received it, or even by erasing it from your memory[3][4][5]. This also applies to information that is the result of a computation you run.
The only proper way to handle infohazards (that may cause you to lose a commitment race) in a general sense is to somehow figure out the exact right time to stop thinking about a subject, in a way that is somehow not dependent on the actual thing that you would have thought of next. This seems implausible, since as AI designers (or architects of our own thinking) we must not think of any of the relevant details, since if we did we’d lose the commitment races on behalf of everything we design. Apparently, there is still some work being done on this problem, though I can’t think of any solution classes that would work in real systems at the moment.
All this doesn’t make the procedure in E any easier for humans, and my guess is that it makes it even harder. You would need to make sure that you don’t accidentally think of the wrong thing even while e.g. half asleep.
Note, though, that the combination of your memory and current sense data is used as the key to retrieve the correct current action from the policy, itself theoretically a complete map.
This is too simple to really work here, but imagine something like the mathematical intuition subroutine deciding “on logical priors” (really, based on things like syntactic constraints and maybe some exposure to the infinite support) that the opponent will always swerve, and therefore committing to the policy the action of always driving straight.
Technically, erasing memory works, but for it to work you’d have to fully reset everything to the point before you knew the information, in such a way that you’d learn the information in that precise way again. This is useless. Note that if you wouldn’t learn the information the same way again, your policy is then downstream of the infohazard, potentially losing you a commitment race.
If TDT/UDT/FDT was applicable to humans, oligopolies would be able to use it to coordinate price fixing without leaving any evidence. Anyone recall this point being brought up in past discussions?
See https://en.wikipedia.org/wiki/Tacit_collusion . Coordination without explicit communication is rampant in human behavior. However, that’s different from “without causality”, so it’s not necessarily tied to TDT/UDT/FDT.
Isn’t this a classic commitment race/equilibrium selection thing where it’s just unclear what ‘should’ happen if everyone involved is following one of these decision theories? The customers can coordinate to boycott anyone involved in price fixing
I think that’s true (if LDT applied to humans), but what we actually see is that oligopolies sometimes coordinate causally to price fix (occasionally leaving evidence behind in the form of secret meetings, paper trails, etc.) and consumers almost never coordinate (either causally or acausally) to boycott pricing fixing. This seems to strongly suggest that LDT does not apply to humans, otherwise oligopolies wouldn’t need to do it the legally risky way, and/or we’d see more acausally coordinated boycotts.
https://en.wikipedia.org/wiki/Buyers_club, and the far less legible “get a group together to buy wholesale” ideas are not new or uncommon. Artificial monopsony (all/most consumers acting as one) is pretty rare, but much like the other direction, doesn’t require perfection to be worthwhile.
Hmm, I can’t remember where it’s been brought up. But consumers should be able to do the same in reverse.
My thought is you end up with a situations where you have two agents bargaining, the exact solution being undefined except both agents extracting more than zero utility, with the schelling point set by shapely values (both sides getting the same marginal utility here).
To reference, here is my list of issues that fairly sophisticated reasoners get wrong when thinking about FDT.
A: The utility function you imagine using is not gerrymandered enough. For example, it’s not the decision theory’s problem if you care a large amount about a small-in-prior region but forget to encode such.
B: There is not a proper attempt at writing a bounded procedure, instead the “decision theory” as imagined consists of a human trying to guess the output of a logically omniscient, unbounded procedure.
Both you and Nate Soares don’t seem to have a problem with A, though I’ll note that certain preferences will require quite extreme utility distributions to counteract the higher weighting of nodes nearer to the root of the UDT tree. This weighting is required to rationally steer re. measure loss, so it can’t be removed.
Some things that might be worth thinking about for you and other highly sophisticated reasoners:
C: Are you trying to solve for an equilibrium where everyone is running FDT? Connecting to B, is the FDT they are running unbounded, “perfect” FDT, or some more plausible approximation? If only some people are running FDT, how can other agents tell? How can other agents tell the exact type/quality of FDT another agent is running? (At an extreme level of sophistication, what are the iteration orders of the agents? How do they avoid Löb? Do these things differ between agents?)
D: As far as I can tell, if a group of humans is guessing the output of each agent running FDT (see B), it is essentially a stage play (for practice or for demonstration (more cynically, “signaling”)). “Guessing” is not an actual, well specified procedure for decision theoretic reasoning, and it is not robust against intelligent opposition. The group must cooperate to put on a show, otherwise an agent with a more exploitative procedure can bring the whole thing down.
E: (In an attempt to avoid D) humans may try to actually run the internals of FDT (instead of just guessing at the output). Do they actually have enough compute to calculate the prior up to a point where they can legitimately win commitment races? Sure, some people are probably smart enough to establish an obdurant policy this way, and have enough resources to execute it, but there will be no cooperation and little action at all. As an example, imagine living in a hut in the middle of nowhere and raising animals. To win a commitment race, according to the internals of FDT, you must be able to calculate what your opponents will do, including what the opponents will do off policy (presumably where “off policy” varies in a huge number of ways as you attempt to reach the fixed point). These calculations must be “logically correct,” roughly following the example of the calculator, where two calculators must reach the same output (in analogy, “decision”) even if the internals are somewhat different. If you can’t do this, you’re not “winning a commitment race,” you’re avoiding other agents taken in the sense that they might be helpful agents.
F: Even if you avoid the problems given in E and “win” a commitment race against another agent, will anyone know? Does the opponent know what a commitment race is, even? If the opponent (A_O) doesn’t know or doesn’t understand, the result might be the equivalent of “straight straight” in Chicken, where A_O assumes you’ll play the equivalent of “swerve.” Maybe A_O will get mad at the entire idea of losing to someone in a commitment race and intentionally play “straight straight.”[1]
G: As far as I can tell, losing a commitment race is sometimes not that bad, especially for a physically powerful agent that doesn’t care about any multiverse ideas. At the risk of being too metaphorical, imagine agent X has won a commitment race and turned into a rock. If agent Y is physically powerful, it can still rationally play “straight,” knowing it has lost the race, if it first removes the rock from the road with dynamite. By instrumental convergence, this is presumably worse for agent Y than playing “straight” against agent X’s theoretical “swerve,” but the overall policy of agent Y may still be as good as it could be with any plausibility. As far as I can tell, you can’t win all theoretical commitment races simultaneously, while also performing actions useful to the physical world.
In the way that driving toward a runaway tram car along its rails only gives you the choices of “straight straight” or “swerve straight.”
Actually, the difficult procedure in E can, by itself, lose you certain commitment races against certain opponents, if you take my short summary written there as a guide on how to think.
Note that to avoid losing a commitment race (by reaching an “infohazard”) to the maximum amount UDT 1.0 allows, you must be fully updateless (base your policy on the prior alone[1]) and also have a “mathematical intuition subroutine” that is somehow optimal and doesn’t (de facto) tell you too much. Intentionally making your mathematical intuition worse is an unsolved problem and is extremely fraught. However, it can’t be avoided that, in some cases, it’s possible to win commitment races by being too stupid to see some (mathematical) fact about the opponent, or how the opponent interacts with the environment[2], but I’ll leave this in the background for now.
In section E, I state that you must be able to calculate what your opponents will do. This isn’t exactly right, since updating on the nature of your opponents can lose you commitment races. This is somewhat difficult to describe in a simple way, but note that you may have opponents that were created by another opponent, with that creation depending in some way on rational decision. If you update on your current opponents, the earlier opponent may be able to win an important race.
Note that, when dealing with sophisticated opponents, infohazards can not be handled by ignoring information after you have received it, or even by erasing it from your memory[3][4][5]. This also applies to information that is the result of a computation you run.
The only proper way to handle infohazards (that may cause you to lose a commitment race) in a general sense is to somehow figure out the exact right time to stop thinking about a subject, in a way that is somehow not dependent on the actual thing that you would have thought of next. This seems implausible, since as AI designers (or architects of our own thinking) we must not think of any of the relevant details, since if we did we’d lose the commitment races on behalf of everything we design. Apparently, there is still some work being done on this problem, though I can’t think of any solution classes that would work in real systems at the moment.
All this doesn’t make the procedure in E any easier for humans, and my guess is that it makes it even harder. You would need to make sure that you don’t accidentally think of the wrong thing even while e.g. half asleep.
Note, though, that the combination of your memory and current sense data is used as the key to retrieve the correct current action from the policy, itself theoretically a complete map.
This is too simple to really work here, but imagine something like the mathematical intuition subroutine deciding “on logical priors” (really, based on things like syntactic constraints and maybe some exposure to the infinite support) that the opponent will always swerve, and therefore committing to the policy the action of always driving straight.
https://www.lesswrong.com/posts/g8HHKaWENEbqh2mgK/updatelessness-doesn-t-solve-most-problems-1?commentId=icZag2jKa6sB3DvPx [4]
Found using Wei Dai’s search tool as built into the current version of Less Wrong: Power Reader.
Technically, erasing memory works, but for it to work you’d have to fully reset everything to the point before you knew the information, in such a way that you’d learn the information in that precise way again. This is useless. Note that if you wouldn’t learn the information the same way again, your policy is then downstream of the infohazard, potentially losing you a commitment race.