I don’t have the time to read this in depth, but AFAICT your objections to both meanings of the word “theoretical” route through your belief that decision theory will/should be used to program an AI? Yeah, I’m quite skeptical of that—I think superintelligence will reason in the same high-level, informal, embedded, bounded way that we do (even if with a lot more capability to put towards e.g. doing BOTECs).
Creating a summary of my point is difficult, since the summary needs to avoid giving the impression that UDT/FDT doesn’t work outside “fair” environments (that impression would be false). Therefore, take the following as an approximate description of what I’ve already written, though I will try to maintain general accuracy.
Consider a prisoner’s dilemma setup, between an extremely intelligent human and a machine. The human’s favored option is to defect against a cooperating machine. Imagining the machine has a small chance of a fault, the strategy of always playing “defect” will sometimes bring this about. However, the human wants to use a high quality decision theory. The human “philosophically” thinks about decision theory, and considers FDT/UDT, realizing that an FDT agent would play “cooperate” in the current situation (the machine is known to play “cooperate” when it was paired with a few FDT/UDT agents, who also cooperated). The human decides that the thought is correct and makes it reality, playing “cooperate.” The machine plays “defect,” delivering the extremely intelligent human, one who made no errors within the train of philosophical thought, the worst possible option (best possible for the machine). How did this happen?
The fact of the matter is that taking the correct philosophical reasoning for what an FDT agent would do in your situation, disregarding allegedly unnecessary “theoretical” details, and making actual the decision that the agent would make is not FDT. The machine correctly reasons (as a sufficient but not necessarily unique reason for the human’s failure) that the human would reach a decision to play “cooperate” not only if the machine cooperates, but also if it defects. Defecting is better, so it does. Note carefully that this does not conflict with the correctness of the human’s philosophical reasoning, since FDT agents will, in most cases, continue to receive outcomes of CC. If the human has a theory of counterlogicals, a check of consistency between FDT agents receiving CC and the (counterlogical) impending receipt of CC, performed during deliberation, will result in a pass. Of course, it also still results in being defected upon as a cooperator, the loss state the Nash solution always avoids.
I can try to come up with an additional case where the human’s reasoning is even more sophisticated and FDT-like that still results in a failure if you like.
(Describing this kind of thing in English isn’t really the right approach anyway, e.g. the phrase “close by” in “close by [...] in as many other ways as possible”[1] has slightly incorrect connotations, a technical use of the word “acceptable” might be clearer, but even that’s not quite right. It needs to be written as math and computer programs.)
Maximax plays “defect” against an opponent with a trembling hand in PD[1] so “cooperate” can’t be defended as “optimistic” here. However, the action could be labeled “trusting” (derogatory) but not “trusting” (laudatory), given that the payout values are human predictable (with only moderate intelligence or memorization).
Flavor: since “defect” is the only way to get the best outcome (i.e. get DC), calling an agent that plays C “optimistic” is like calling someone who plays an option that perfectly bans all AI forever a “Singularitarian optimist” (even the standard PD cardinals almost line up, as long as you ignore the symmetric downside).
I don’t have the time to read this in depth, but AFAICT your objections to both meanings of the word “theoretical” route through your belief that decision theory will/should be used to program an AI? Yeah, I’m quite skeptical of that—I think superintelligence will reason in the same high-level, informal, embedded, bounded way that we do (even if with a lot more capability to put towards e.g. doing BOTECs).
Creating a summary of my point is difficult, since the summary needs to avoid giving the impression that UDT/FDT doesn’t work outside “fair” environments (that impression would be false). Therefore, take the following as an approximate description of what I’ve already written, though I will try to maintain general accuracy.
Consider a prisoner’s dilemma setup, between an extremely intelligent human and a machine. The human’s favored option is to defect against a cooperating machine. Imagining the machine has a small chance of a fault, the strategy of always playing “defect” will sometimes bring this about. However, the human wants to use a high quality decision theory. The human “philosophically” thinks about decision theory, and considers FDT/UDT, realizing that an FDT agent would play “cooperate” in the current situation (the machine is known to play “cooperate” when it was paired with a few FDT/UDT agents, who also cooperated). The human decides that the thought is correct and makes it reality, playing “cooperate.” The machine plays “defect,” delivering the extremely intelligent human, one who made no errors within the train of philosophical thought, the worst possible option (best possible for the machine). How did this happen?
The fact of the matter is that taking the correct philosophical reasoning for what an FDT agent would do in your situation, disregarding allegedly unnecessary “theoretical” details, and making actual the decision that the agent would make is not FDT. The machine correctly reasons (as a sufficient but not necessarily unique reason for the human’s failure) that the human would reach a decision to play “cooperate” not only if the machine cooperates, but also if it defects. Defecting is better, so it does. Note carefully that this does not conflict with the correctness of the human’s philosophical reasoning, since FDT agents will, in most cases, continue to receive outcomes of CC. If the human has a theory of counterlogicals, a check of consistency between FDT agents receiving CC and the (counterlogical) impending receipt of CC, performed during deliberation, will result in a pass. Of course, it also still results in being defected upon as a cooperator, the loss state the Nash solution always avoids.
I can try to come up with an additional case where the human’s reasoning is even more sophisticated and FDT-like that still results in a failure if you like.
(Describing this kind of thing in English isn’t really the right approach anyway, e.g. the phrase “close by” in “close by [...] in as many other ways as possible”[1] has slightly incorrect connotations, a technical use of the word “acceptable” might be clearer, but even that’s not quite right. It needs to be written as math and computer programs.)
In the last paragraph of https://www.lesswrong.com/posts/oZzRHiSZPcjrWHeoE/stop-doing-decision-theory-without-metaphysics?commentId=2A38HeJ3cLDmtEXJS
Maximax plays “defect” against an opponent with a trembling hand in PD[1] so “cooperate” can’t be defended as “optimistic” here.
However, the action could be labeled “trusting” (derogatory) but not “trusting” (laudatory), given that the payout values are human predictable (with only moderate intelligence or memorization).
Flavor: since “defect” is the only way to get the best outcome (i.e. get DC), calling an agent that plays C “optimistic” is like calling someone who plays an option that perfectly bans all AI forever a “Singularitarian optimist” (even the standard PD cardinals almost line up, as long as you ignore the symmetric downside).
Including FDT+maximax utility function, if the tremble is downstream of the opponent’s FDT.