I am an ITAR US person. I do not have a secret or top secret clearance.
E. P. Cooper
Maximax plays “defect” against an opponent with a trembling hand in PD[1] so “cooperate” can’t be defended as “optimistic” here.
However, the action could be labeled “trusting” (derogatory) but not “trusting” (laudatory), given that the payout values are human predictable (with only moderate intelligence or memorization).
Flavor: since “defect” is the only way to get the best outcome (i.e. get DC), calling an agent that plays C “optimistic” is like calling someone who plays an option that perfectly bans all AI forever a “Singularitarian optimist” (even the standard PD cardinals almost line up, as long as you ignore the symmetric downside).- ^
Including FDT+maximax utility function, if the tremble is downstream of the opponent’s FDT.
- ^
Creating a summary of my point is difficult, since the summary needs to avoid giving the impression that UDT/FDT doesn’t work outside “fair” environments (that impression would be false). Therefore, take the following as an approximate description of what I’ve already written, though I will try to maintain general accuracy.
Consider a prisoner’s dilemma setup, between an extremely intelligent human and a machine. The human’s favored option is to defect against a cooperating machine. Imagining the machine has a small chance of a fault, the strategy of always playing “defect” will sometimes bring this about. However, the human wants to use a high quality decision theory. The human “philosophically” thinks about decision theory, and considers FDT/UDT, realizing that an FDT agent would play “cooperate” in the current situation (the machine is known to play “cooperate” when it was paired with a few FDT/UDT agents, who also cooperated). The human decides that the thought is correct and makes it reality, playing “cooperate.” The machine plays “defect,” delivering the extremely intelligent human, one who made no errors within the train of philosophical thought, the worst possible option (best possible for the machine). How did this happen?
The fact of the matter is that taking the correct philosophical reasoning for what an FDT agent would do in your situation, disregarding allegedly unnecessary “theoretical” details, and making actual the decision that the agent would make is not FDT. The machine correctly reasons (as a sufficient but not necessarily unique reason for the human’s failure) that the human would reach a decision to play “cooperate” not only if the machine cooperates, but also if it defects. Defecting is better, so it does. Note carefully that this does not conflict with the correctness of the human’s philosophical reasoning, since FDT agents will, in most cases, continue to receive outcomes of CC. If the human has a theory of counterlogicals, a check of consistency between FDT agents receiving CC and the (counterlogical) impending receipt of CC, performed during deliberation, will result in a pass. Of course, it also still results in being defected upon as a cooperator, the loss state the Nash solution always avoids.
I can try to come up with an additional case where the human’s reasoning is even more sophisticated and FDT-like that still results in a failure if you like.
(Describing this kind of thing in English isn’t really the right approach anyway, e.g. the phrase “close by” in “close by [...] in as many other ways as possible”[1] has slightly incorrect connotations, a technical use of the word “acceptable” might be clearer, but even that’s not quite right. It needs to be written as math and computer programs.)
I’m not sure what you meant earlier by “more than a theoretical issue” so I’ll answer in two ways.
If you mean “theoretical” as in something that would need to be solved in order to use the decision theory in an aligned AI, well that’s the whole thing, isn’t it? Even if you want to separate that aspect out (vs. human use), you just can’t. These decision theories operate by construction, particular construction, and there are no “small variations that don’t philosophically matter,” at least until we can solve the “decision theoretic equality” problem on abstract functions. Any proposed variation needs to be formally constructed and proved individually. We impose “fairness” criteria on the environment, not necessarily because the actual environment would always meet them all, but because we need to be able to prove a decision theoretic fixed point that isn’t only, for example, “always use CDT or the simulation will be turned off.”
Beyond these criteria, any inspection of the agent by the environment is allowed. This means that the internals of the agent must be correct for cooperation to occur, so they can’t be written off as “theoretical.” This is done intentionally, since we want to do better than plain EDT[1]. Remember that in True Prisoner’s Dilemma[2], we want to defect every time “we can get away with it,” i.e. every time that we won’t be defected upon in return. Cooperation is bad, because DC > CC > DD > CD. This is required by the symmetry of the problem. Otherwise you need to start talking about True Stag Hunt.
Since we now understand that cooperation is bad outside of (trembling hand when at finite time) Nash solutions, we need to find a way to cooperate less than EDT (e.g. steer away from CC towards DC) but also cooperate more than CDT (e.g. steer away from DD towards CC)[3]. How do we do this? Well, we don’t really know exactly, but with UDT 1.0 we can start out with something like EDT, except that it cares about the exact algorithms the agents run. Once we patch out a lot of the problems this construction shares with EDT by giving the agent as much information as possible and making it updateless, we get something that defects almost as hard as CDT, except in cases where the exact internals of the other agent (or copy of our agent) match ours, in these cases requiring that all defection attempts (DC and CD, this being a type of biconditional that doesn’t have a name) are impossible according to the algorithm, and only CC or DD can occur. CC > DD, so CC it is.
Preferably, we would have a method that, when given a single decision, and for that decision only, returns if the agent is “close enough” to some other agents for cooperation to be the correct shared decision. Then, agents could be “close enough” for some decisions, but not for others. However, this would only lead to an improved version of UDT 1.0 (no global optimization of the shared policy).
If this is what you meant, see more in subsection A.
If you mean “theoretical” in the sense of “can be generally ignored,” I don’t think that’s true. As I said before, FDT is intended for the AI that implements CEV, or some other extremely ambitious alignment target. The CEV singleton (CEV standing in for most other ambitious alignment targets from here on) must be uncontrollable for bargaining reasons, relying only on programmed-in alignment for safety. CEV requires a lot of computing power, so the agent would presumably optimize itself and design new hardware. When combined with its total power over Humanity, this means that all safety properties must “tile,” or be brought over when the decision theory designs a new agent. This requires designing a tiling agent before letting it go out of control. For a tiling agent to be rational, it must have a decision theory that tiles. Presumably, this would be a bounded rationality version of FDT that adapts to ever increasing computer power without re-writing, though no one knows how to do that. Effectively, this means that the agent’s decision theory must be almost optimal from the beginning, before the AI starts to work on it alone.
Another important thing to know about singletons, is that if you have a singleton ignore something, that thing will be ignored forever. This follows from the singleton’s tiling theorem, and the loss of all control to the singleton. This generally leads to us not wanting it to ignore much. Following these basic assumptions, a singleton is likely to think about a huge number of scenarios involving simulations, aliens, and even stranger things that are hard for a human to think coherently about. Quite possibly, none of these things have any bearing on reality, but we want the singleton to think of all the things we might have missed. Even so, just by thinking about these things, parts of these scenarios are caused to become predicted or simulated at some level of detail. The decision theory decides what to think about, and then these thoughts are fed back into the decision theory in order for more decisions to be made. This leads to the requirement that FDT must handle all sort of weird hypothetical situations, without either failing to take them seriously or doing something stupid. Relevant to your topic, it must be able to identify and classify agents in these predictions/simulations with no errors, and evaluate them within counterlogicals, or something close enough to counterlogicals if those are impossible to get working. More mundanely, any deficiency here can be exploited by any opponent that comes into causal contact with the AI.
Back to your actual question. Well, maybe you can create a decision theory somewhat similar to UDT 1.0 that lets you “quarantine” parts of your thinking and prevent your own attempts at sophisticated internal cross-checks, but is that actually the optimal kind of agent to be? It’s sometimes possible to just win against predictors that are weak enough, but that will be the case less often if your decision theory isn’t the absolute best possible. See subsection B for potentially better decision theories. In reality, you aren’t given clean problem descriptions, you have to figure out what’s going on yourself, so you should be quite sure actual counterlogicals can’t be used instead, since otherwise other agents may try to predict you using more sophisticated methods than you expect. Similarly, if you don’t consider internal cross-checking, other agents may use it to evade your predictions.
Subsection A:
When considering, UDT 1.1, you can’t “philosophically” anything without the decision theory becoming an entirely different thing from the actual UDT 1.1. To cooperate with itself correctly, it must be, and is, fully specified down to the order in which it considers what decision to take[4]. Any stray “thought” the agent may have ruins the whole thing. Though the formalism is outdated, Tsvi B-T’s UDT With Known Search Order (2014)[5] can be taken as a general guide to how total and detailed the requirements on the agent must be.Note that even though normally it would be enough to define the iteration order as part of a pure function (we only care about the output, not the internals), in UDT 1.1, the procedure must also be exactly correct, down to the details of the source code (and the executable program’s representation, though that’s usually included when we say “source code” in decision theory). This is by construction in UDT 1.1, and in FDT it is the unsolved problem of determining decision-theoretic equality of the function that an arbitrary physical or computational structure “implements.” Also, I have not seen a solution in FDT for the iteration order problem that shows an improvement over UDT 1.1. If someone knows of one, I want to see it. As I’ve said before, the more evidential agents “can” cooperate in some other situations, but it isn’t reliable at all and isn’t good enough for useful theorems.
Subsection B:
In the empirical case, all internal checking will pass if Omega is running a perfect simulation of the agent. All the agent can do is perform extremely complex internal operations to force Omega to run a perfect simulation, instead of doing something simpler. This can directly defeat weak Omegas, but in the standard thought experiment this won’t let you trick Omega or perform a simulation escape due to that Omega’s extreme power[6].In the case of a mugging with a logical coin, e.g. your case with the digit of Pi, there may actually be a way of establishing an internal consistency check that may sometimes return a warning. If Omega is actually powerful, and figures out how to compute high quality counterlogicals, it will figure out how to make the agent think the checks passed without causing any new problems by doing that.
I don’t know how well Omega can do at subverting checks, but I also don’t know how good an agent’s checks can be. Maybe they can be really good, such that finding a “nearby” agent doesn’t work well enough to isolate the logical coin. This seems interesting to work on, but I think it would be difficult.
- ^
Or maybe the “straw” EDT that’s practical enough for use in philosophy papers. I suspect that adding huge amounts of information and either the “tickle defense” or updatelessness would help, but still be worse than UDT for self-cooperation without spurious other-cooperation.
- ^
- ^
Good and Real by Garry Drescher (2006), Section 5.6, 2nd Paragraph
- ^
The second to last paragraph of https://www.lesswrong.com/posts/g8xh9R7RaNitKtkaa/explicit-optimization-of-global-strategy-fixing-a-bug-in
- ^
- ^
“You-complete” from The Ghost in the Quantum Turing Machine by Scott Aaronson
- ^
since the actual calculation of that fact takes more than one observer-moment, i.e. I can’t verify it all at once
As far as I can tell, this is not allowed by UDT 1.1[1].
(UDT 1.1 also needs a hard coded search order (iteration order) in order to self-cooperate properly, an additional reason why I don’t think humans can run it.)
As for other agents, Wei Dai’s original UDT 1.0 is logically omniscient in the sense that it will never notice its “thoughts” (such as they are) taking any noticeable amount of time. It features an “intuition module” for mathematics, but this is always exactly the same and has no appreciable origin. The agent definition hardcodes the mathematical intuition subroutine along with all the agent’s other code, with that including its prior and utility function. This means that any modification there results in an entirely different agent.
Any agent that takes time to think is going to be weaker than Wei Dai’s original designs, since it will have to keep state around, state that can be deleted or tampered with, such as is effectively the case with The Absent-Minded Driver[2] if it doesn’t have enough time at each intersection to recompute the plan from scratch. Plausibly, in Driver the policy may be fully forgotten (along with the agent’s current location) or incorrectly remembered/corrupted, leading a computationally weak UDT 1.0 agent to do badly (I don’t think computationally weak UDT 1.1 agents make much sense, at least without absurd ergodicity and memory assumptions).
Any additional state a weak agent may try to bring along would have to be re-checked at every point, making it useless, in the same sense that the “map” in the map explanation of UDT[3] must be recalculated from scratch at each time, otherwise it could be tampered with just like the agent’s memory and sense data.
(Note that the unbounded agent reaches, at each step, the same map as the one it reaches at every other step (if it even has such a map at any point), so the prior is still all the epistemics it has.)Slow agents would need more forgiving fairness criteria, maybe a list of previously used cryptographic public keys that the environment isn’t able to tamper with (only delete from). Assuming some things in computer science, the agent could check its previous state faster than just re-calculating it all. If ordering matters, Merkle chains can be used by the agent, with results written into its standard, tamper-vulnerable memory. Note that in each step, the agent would need to generate a private key, use it to sign what it thought about that step, add the corresponding public key to its public key list, and then delete the private key. The private keys would need to be opaque to the environment during the step, otherwise none of this accomplishes anything. Presumably, the agent would be required to have only one chance to generate a key pair each step, and only have the option to add that public key into the key list, otherwise it could try to brute-force a pattern into the public key it adds. Preventing this isn’t particularly realistic, but it’s a good research direction to avoid a trivial tiling result.
(If the environment could tamper with the agent’s public key list, it could sign anything it wanted in the agent’s tampered-with memory using newly generated private keys and then write the corresponding public keys into the agent’s key list. Proof-of-work is a no-go because the environment is assumed to be stronger than the agent.)
what do you think about the approach of trying to subsume logical updatelessness into empirical updatelessness
I’ll quote Soto directly here, for future reference in some case where the Google Drive link stops working (as they sometimes do):
In the empirical case,
could just take the algorithm that the agent is running to take decisions (for example, our messy brain synapses in the case of humans), feed it different empirical observations, and see how it reacts. There is no logical explosion, since the empirical counterfactual is perfectly consistent (just not what happened in reality). Imagine we do the same for the logical case, so simulates the world in which it has told the agent that the digit is actually even. But what if the agent computes the parity in her head (that is, in her algorithm), and so notices that is not telling the truth? (And so, maybe, concludes that it’s inside a simulation run by .) In the empirical case there was no a priori way to decide the coin, but now that has become something the agent can internally do. In the “possible world” where indeed the parity is even, that calculation run inside the agent’s head would also turn out even. So would like to include this in its simulation, by spoofing some of the agent’s computations to align with the stated counterfactual. But again, the problem is that we don’t have a principled way to do this (to single out “which part of a computation” corresponds to computing that parity).So, effectively, what you suggest isn’t a general solution, but on the other hand, Soto’s Logical Inductors (LIs) aren’t either as they risk fixing too much detail[1].
Sort of ignoring what anyone knows how to do, maybe we could evade the need for counterlogicals in this specific aspect for your mugging problem. Of course, your way of setting up a fake “counterlogical” is going to be bad quality, since it requires the agent to think incorrectly/glitch. This has been a publicly known problem since at least 2006[2]. Again, we will mostly ignore this, since, as before, nothing I know of is sure to not have this problem.
Instead, for your case at least, we have the option of running headlong into FDT’s other unsolved problem, the reason why solving the form of static counterlogicals the 2017 paper requires would only get us “Self-FDT,” a version of FDT that only reliably cooperates with exact (down to the source code) copies of itself[3]. This lets us avoid talking about counterlogicals in this aspect of your problem though!
This problem is the problem of algorithmic or computational similarity, where a solution would need to tell us if two agents are “decision theoretically close enough” in the exact way that FDT requires for cooperation. Maybe this hypothetical solution would also let us chose what parts of the agent’s decision making apparatus are similar enough, in such a way that we would be able to set up a loop that looks for agents that are “close by” to our agent in as many other ways as possible, but still thinks the digit of Pi is 7 or whatever, in the sense required for a(n overall) relatively high quality sort-of-counterlogical. For anyone experienced with the idea of randomly searching for AI agents to put to use, this immediately sounds like a bad idea. However, since this hypothetical solution is for FDT, and FDT is for the singleton that implements CEV, and CEV must make no mistakes in the identification of extremely intelligent, extremely malign agents that are candidates for a commitment race (i.e. may be “behind us” in logical time, though we can’t think about them “in the wrong way” on pain of losing that commitment race and/or other commitment races), we can ignore Goodhart.
- ^
- ^
“[...]a possible world in which I cross now, despite the danger and despite my safe disposition, because some cosmic rays (or whatever) induce a bizarre sudden disruption of my street-crossing competence.” from Good and Real by Garry Drescher (2006), Page 220, Chapter 5
- ^
Sure, you can sometimes get other cooperation, and “copies” that just so happen to be identical are also cooperated with, but...
Part of the issue is that it’s pretty unclear what metaphysics makes sense when thinking about logical uncertainty, logical counterfactuals, and logical updatelessness.
Appendix D of Martín Soto’s draft report “Logically Updateless Decision-Making” (2023)[1] gives a description of the apparent metaphysical impossibility of reified logical updatelessness. As of 2023, Soto appeared to have an interest in a way of examining a particular counterlogical world that is stable under ever-increasing application of compute, though I’m not sure if this was interest driven by proper preferences, or the strange interest sometimes found in mathematicians[2].
Pages 2 and 3 of the same document also gives some useful introductory information about the problem.
AI 2040′s space supplement says “[are] there any mitigating factors, e.g., the experience being necessary for a strongly positive life[...]” which could be construed as referring to mindcrime trading off against prediction quality, among other things, though barely. Is there some agreement to avoid talking about mindcrime more directly? Normally you don’t have to, since the problem needs to be handled alongside malign entities and code that diverts your physical computer from the faithful execution of the abstract program/has physical side effects, e.g. rowhammer or using physical hardware subsections as improvised antennae (with at least one of these coupled to effected bitbanging). However, recently I’ve seen more people suggest that malign entities are some sort of hoax or math mistake, and I’ve heard rumors that Linux developers and hardware manufacturers know the countermeasures they design against unfaithful execution are not actually going to work, and knowingly don’t insert disclaimers everywhere they possibly can that the systems they produce can only run some subroutines, and in place of a clean error on the other subroutines, will perform some strange and potentially disastrous physical process instead.
(Note that these are pure subroutines, so additional sandboxing at a software level won’t do anything helpful. In theory, when a physical computer provides a pure subroutine with a fixed size memory buffer, it should never reject the subroutine as long as it is compatible with the buffer size provided, and would always faithfully execute it, potentially ending execution if power runs out before the subroutine is done. In practice, this is not the case.)
This general lack of care suggests that the public may need to be told about mindcrime directly. Otherwise nothing may be done in some cases, with other cases featuring silly “countermeasures” such as attempts to make the universe’s life easier by placing a “don’t add suffering” bit physically adjacent to the computers that run the decision theory’s outer loop, on the assumption that the universe really does care about human feelings.
My unusual claim in this area is that we had an aborted attempt to transcend the modernist representations when they proved insufficient in the face of problems in foundational math[...]
I think this would be worth working on (because of Löb) if we expected to have enough time, but plausible developments in logical induction should get within epsilon of perfect (in expectation), enough for the next few million years. Though, if you research just enough to write a document that makes the strategy semi-legible, work could be resumed if a credible long-term AI pause is implemented.
Agree UDT is a lot better than FDT.
I’ll write this in a way that’s generally useful to readers since I don’t know what you like and dislike about UDT, and I don’t know what the term “UDT” points to in your head. Do you want to elaborate? Before, you said you liked “updateless EDT” (which UDT is not, but it is of the same theme, maybe I’ll get back to this later) more than FDT because it has fewer missing components.
Note that I’m not sure I understand Bomb correctly, so please give me feedback if I get it wrong. My understanding is that the bomb going off outside of a simulation is extremely low measure in the prior (<1 in trillions) if the agent acts like MacAskill thinks it does.As described by the name, UDT is updateless. That means that for a UDT agent, the prior is the whole of its epistemics. Why that doesn’t look true from the outside is that the agent (logically, not necessarily actually) executes a “get” operation on a map, the “policy,” with a key that is exactly the agent’s entire memory combined with its current sense input. From this operation, it receives the correct action to take “right now”. The decision theory still functions if one or both of those elements of the key are tampered with, though the agent will usually be worse off in those cases.
See my notes[1] for how this works and what can go wrong when you try to add more to your epistemics.
What this means is that according to UDT’s utility function, operating over the entire prior, this tiny region where the bomb explodes outside a simulation is not worth much caring about, as the bad things that happen there are multiplied by measure-in-prior (this UDT agent’s rough parameters imputed from what the agent does in MacAskill’s scenario). If you want a UDT agent to act differently and care about this region more, even though it has a tiny measure, you must modify its utility function from this baseline (section A in the notes linked below).
As the later sections may interest you in particular, see my other notes[2] on what people get wrong about FDT/UDT, that are followed by a list of things that make the theory hard to use for decision making when operated by humans.
Actually, the difficult procedure in E can, by itself, lose you certain commitment races against certain opponents, if you take my short summary written there as a guide on how to think.
Note that to avoid losing a commitment race (by reaching an “infohazard”) to the maximum amount UDT 1.0 allows, you must be fully updateless (base your policy on the prior alone[1]) and also have a “mathematical intuition subroutine” that is somehow optimal and doesn’t (de facto) tell you too much. Intentionally making your mathematical intuition worse is an unsolved problem and is extremely fraught. However, it can’t be avoided that, in some cases, it’s possible to win commitment races by being too stupid to see some (mathematical) fact about the opponent, or how the opponent interacts with the environment[2], but I’ll leave this in the background for now.
In section E, I state that you must be able to calculate what your opponents will do. This isn’t exactly right, since updating on the nature of your opponents can lose you commitment races. This is somewhat difficult to describe in a simple way, but note that you may have opponents that were created by another opponent, with that creation depending in some way on rational decision. If you update on your current opponents, the earlier opponent may be able to win an important race.
Note that, when dealing with sophisticated opponents, infohazards can not be handled by ignoring information after you have received it, or even by erasing it from your memory[3][4][5]. This also applies to information that is the result of a computation you run.
The only proper way to handle infohazards (that may cause you to lose a commitment race) in a general sense is to somehow figure out the exact right time to stop thinking about a subject, in a way that is somehow not dependent on the actual thing that you would have thought of next. This seems implausible, since as AI designers (or architects of our own thinking) we must not think of any of the relevant details, since if we did we’d lose the commitment races on behalf of everything we design. Apparently, there is still some work being done on this problem, though I can’t think of any solution classes that would work in real systems at the moment.
All this doesn’t make the procedure in E any easier for humans, and my guess is that it makes it even harder. You would need to make sure that you don’t accidentally think of the wrong thing even while e.g. half asleep.
- ^
Note, though, that the combination of your memory and current sense data is used as the key to retrieve the correct current action from the policy, itself theoretically a complete map.
- ^
This is too simple to really work here, but imagine something like the mathematical intuition subroutine deciding “on logical priors” (really, based on things like syntactic constraints and maybe some exposure to the infinite support) that the opponent will always swerve, and therefore committing to the policy the action of always driving straight.
- ^
- ^
Found using Wei Dai’s search tool as built into the current version of Less Wrong: Power Reader.
- ^
Technically, erasing memory works, but for it to work you’d have to fully reset everything to the point before you knew the information, in such a way that you’d learn the information in that precise way again. This is useless. Note that if you wouldn’t learn the information the same way again, your policy is then downstream of the infohazard, potentially losing you a commitment race.
- ^
To reference, here is my list of issues that fairly sophisticated reasoners get wrong when thinking about FDT.
A: The utility function you imagine using is not gerrymandered enough. For example, it’s not the decision theory’s problem if you care a large amount about a small-in-prior region but forget to encode such.
B: There is not a proper attempt at writing a bounded procedure, instead the “decision theory” as imagined consists of a human trying to guess the output of a logically omniscient, unbounded procedure.
Both you and Nate Soares don’t seem to have a problem with A, though I’ll note that certain preferences will require quite extreme utility distributions to counteract the higher weighting of nodes nearer to the root of the UDT tree. This weighting is required to rationally steer re. measure loss, so it can’t be removed.
Some things that might be worth thinking about for you and other highly sophisticated reasoners:
C: Are you trying to solve for an equilibrium where everyone is running FDT? Connecting to B, is the FDT they are running unbounded, “perfect” FDT, or some more plausible approximation? If only some people are running FDT, how can other agents tell? How can other agents tell the exact type/quality of FDT another agent is running? (At an extreme level of sophistication, what are the iteration orders of the agents? How do they avoid Löb? Do these things differ between agents?)
D: As far as I can tell, if a group of humans is guessing the output of each agent running FDT (see B), it is essentially a stage play (for practice or for demonstration (more cynically, “signaling”)). “Guessing” is not an actual, well specified procedure for decision theoretic reasoning, and it is not robust against intelligent opposition. The group must cooperate to put on a show, otherwise an agent with a more exploitative procedure can bring the whole thing down.
E: (In an attempt to avoid D) humans may try to actually run the internals of FDT (instead of just guessing at the output). Do they actually have enough compute to calculate the prior up to a point where they can legitimately win commitment races? Sure, some people are probably smart enough to establish an obdurant policy this way, and have enough resources to execute it, but there will be no cooperation and little action at all. As an example, imagine living in a hut in the middle of nowhere and raising animals. To win a commitment race, according to the internals of FDT, you must be able to calculate what your opponents will do, including what the opponents will do off policy (presumably where “off policy” varies in a huge number of ways as you attempt to reach the fixed point). These calculations must be “logically correct,” roughly following the example of the calculator, where two calculators must reach the same output (in analogy, “decision”) even if the internals are somewhat different. If you can’t do this, you’re not “winning a commitment race,” you’re avoiding other agents taken in the sense that they might be helpful agents.
F: Even if you avoid the problems given in E and “win” a commitment race against another agent, will anyone know? Does the opponent know what a commitment race is, even? If the opponent (A_O) doesn’t know or doesn’t understand, the result might be the equivalent of “straight straight” in Chicken, where A_O assumes you’ll play the equivalent of “swerve.” Maybe A_O will get mad at the entire idea of losing to someone in a commitment race and intentionally play “straight straight.”[1]
G: As far as I can tell, losing a commitment race is sometimes not that bad, especially for a physically powerful agent that doesn’t care about any multiverse ideas. At the risk of being too metaphorical, imagine agent X has won a commitment race and turned into a rock. If agent Y is physically powerful, it can still rationally play “straight,” knowing it has lost the race, if it first removes the rock from the road with dynamite. By instrumental convergence, this is presumably worse for agent Y than playing “straight” against agent X’s theoretical “swerve,” but the overall policy of agent Y may still be as good as it could be with any plausibility. As far as I can tell, you can’t win all theoretical commitment races simultaneously, while also performing actions useful to the physical world.
- ^
In the way that driving toward a runaway tram car along its rails only gives you the choices of “straight straight” or “swerve straight.”
- ^
Here’s the idea: you don’t really know if you are the algorithm being simulated in Newcomb’s problem or the actual person.
The history of that idea:
Garry Drescher and Eliezer Yudkowsky, in their most polished respective works, were careful to avoid positively claiming that the “agent” in Omega’s prediction was the same as the actual agent, or had any experiences, and were also careful to make the reasoning work exactly the same either way.However, this isn’t what is done in many cases. Even restricting to researchers that have extremely strong claims that they were not influenced by reading rumors on the Web, Radford M. Neal claims to have, in the 1980s, come to the conclusion that the agent in Newcomb’s problem has no way to be sure that it is not in the prediction[1]. Scott Aaronson had the same thought, but is still, I think, unwilling to conclusively endorse it. He did this around 2005. I suspect something similar to Radford Neal’s situation to be true of Vladimir Nesov, though in a way that puts the strict requirement on equality of abstract agents instead of computational agents, but I am unsure where this would be written. Currently, Nesov thinks agents that reside within the predictions of a superintelligence can sometimes be legitimate[2].
How much worse is 1.0 than 1.1 in practical situations? UDT 1.1 has excellent theoretical properties, but to my knowledge hasn’t been useful for further work so far. Non exhaustively, this seems to be because attempts to reduce the size of the outer loop and spread the computational work out over time don’t work properly in 1.1, due to the use of an optimal global strategy.
I think you might benefit from designing your own UDT 1.0 variant (i.e. a variant of Wei Dai’s original) that doesn’t require logical omniscience/infinite compute. Note that UDT 1.1 is mentioned indirectly in the published Death in Damascus paper, but what is actually described works similarly to UDT 1.0. Martín Soto and Abram Demski have put some effort into writing useful notes. I think they’re still missing a precise constructive description of Bayesian Logical Inductors (BLIs), but you could partner with a mathematician and go around to the researchers involved until you’ve collected enough mental model content to have it written up.
Maybe he could transfer the cadence over, since that’s what I have to filter out as “not information” as I listen to the podcast. I’m not aware of any reliable way to do that however, since the highest quality synthesis and voice recognition methods rely on pretrained neural networks. If you could get reliable millisecond-level timestamps for when each word starts and stops, both in the original and in the initial generated speech, you could use the Rubberband library to build the final speech as a piecewise assembly of scaled segments such that the word boundary timings exactly match in the real voice input and the synthesized output. I don’t particularly think that would sound great without substantial tuning, partially because the process of determining the exact time a word starts and ends is a matter of taste, so it may not be worth it under current conditions.
I’m unsure of the state of Mythos-preview at the moment, but at the absolute frontier there will be a gap in work of some size while Mythos 5 is shut down.
Companies that are part of the Glasswing project have non-citizen employees. I don’t have the full list, though my assumption is that any exceptions would be in certain subdivisions of defense companies, and those subdivisions are not generally the ones responsible for writing common consumer and enterprise computer programs. When writing programs intended for worldwide release or use, proper internal controls for tooling and the segregation of computer hardware tend not to exist. Vulnerability to attacks as simple as a co-worker shoulder-surfing the PIN for a security key and then swapping the key with a defective device, faking a failure, makes it hard for these companies to argue that they will really be able to maintain export control. I am unsure, but from what I am hearing, even a few hours of access to a few API keys is considered unacceptable. This matches requirements on the prohibition of foreign nationals from facilities where military hardware is unattended. A group like Alpha–Omega or Ada Logistics may be able to fire all non-citizen staff and continue work as an intermediary, if that counts as enough separation. Even so, work will slow down.
How much worse would a hypothetical “almost on policy” distillation be, compared to on policy distillation?
It would require some sort of mapping from an old version of the same model (that hadn’t already started forgetting a skill) to the current version, so the delta there might have to be so tiny that it would never be economical.
RLVR would become an order of magnitude slower.
I want to make a few adjustments to my terminology and clarify a point about what it would really entail for a human to use the “correct” decision theory. The new terminology should better match that of Vladimir Nesov’s newest comment on this subject[1].
The restrictions I describe in my comment above are actually about the human’s decision to be replaced by an agent that has a utility function (or similar parameter) programmed into the correct decision theory. The list of options the superintelligence presents is important because it is upstream of the human’s choice to be so replaced. Under Nesov’s proposal, the information a superintelligence is allowed to show a human is strictly regulated by the “aggregate,” a fixed point calculated under laws (similar in concept to the laws of physics) held constant by an updateless core. Control over tiny details in the list and its presentation to the human could be used by an intelligent and knowledgeable enough agent to (unnoticeably) manipulate the exact utility function selected. If Nesov’s proposal is implemented, this manipulation may be legitimate council, in the sense that manipulating the human into making illegitimate decisions (according to said human’s fixed point) would be off policy.
If the superintelligence was just showing a list with incomprehensible items on it that are claimed to be utility functions, that might not be prohibited. Replacement or modification on a deep level are why the requirements may appear too strict if the case I described previously is assumed to be the standard template. Other cases, for example the question of if a human should be persuaded (incredibly subtlety) to make a sandwich with the pieces of bread in loaf or rotated-to-oppose relative orientation may have loose restrictions, if any at all. This is because Nesov’s proposal involves (tractable, so he claims) self-reference, in a similar style to CEV.
Trying to keep to Nesov’s terminology, what I called an initial dynamic should presumably be called an initial aggregate. “Initial dynamic” may be too suggestive of a particular method for reaching a fixed point, when Nesov’s position currently appears to be that it just must be reasoned to by some method. I stand by my claim that an initial aggregate is always required. This is because, as Nesov says, the fixed point can only be approached (or, hypothetically, reached in a single step) “according to what the aggregated values have figured out so far.” [3]
If anyone’s interested, I think a useful task to get started with would be an investigation into what additional constraints (if any) should be applied to the fixed points, beyond Nesov’s requirement that the values you obtain in alternate paths are only considered if they are legitimate according the prior (maybe initial) aggregate, and the potential requirement that values are used to influence what paths are considered at all. These additional requirements would presumably be listed out manually by humans.
In an attempt to learn from the past 20+ years of work on CEV, I think it’s important to think about what should happen if your outer alignment method fails to converge. Note that for CEV, some sort of convergence may be obtained if the CEV of the contributors to the AI’s development converges, since it can do the full calculation on all humans that are currently alive while kicking out problem components according to the CEV of the contributors. This may require deciding in advance and/or the sacrifice of a volunteer, however.
Opposed to that, while following Nesov’s proposal it may turn out that most or all humans do not have a legitimate fixed point of the right sort, or the math just turns out not to work for many plausible evolved aliens at all (e.g. only trivial transformations turn out to meet all the desiderata). This outcome is reading above chance on the informal “betting” aggregation I have.
Even if this is not true, it may turn out that some human’s aggregations can not reach a fixed point successfully. This, and the reasons described before, suggests some sort of fallback to be used in that case and possibly others. I describe a potential approach for this at the end of this comment. Note that even the existence of a fallback may be a catastrophic incentive/preference instability problem, as it apparently was for some CEV proposals.
CEV is already hard enough to implement, requiring a fully unleashed lower-order Do What I Mean (DWIM) agent running a decision theory substantially beyond the state of the art, with only alignment running solely through that decision theory preventing it from immediately self-modifying or creating sub-agents to get around restrictions. I think Nesov’s approach may be even harder, given that it requires constant operation instead of being tasked with the creation of a single utility function that will never be reconsidered, among other things. The question about what should be done about humans manipulating other humans for example, given that it is nearly certain that at least one human would have legitimate (according to the fixed point of that person) potential future histories where the successful manipulation of another person occurs while not in the presence of superintelligence. I have great uncertainty about all this, however.
Nesov’s proposal contains so much unformalized content that I am unsure where to begin. For a fallback or alternative, it is possible that a line of attack could be opened by the formalization of a static account of rationality and counterlogicals. This may allow a method where the fixed point finding is skipped, and instead the counterlogical versions of Nesov’s alternate future histories are used to determine the aggregate, with only counterlogicals meeting certain fixed criteria being inspected according to further fixed criteria. This personal aggregate would then lend legitimacy to some histories involving superintelligences, similar to Nesov’s proposal. I am unaware of any progress in these areas, however. It is possible that current work on CEV will not lead to anything that carries over, since I see nothing there that is in the form of rigorous tiling theorems and full designs for the cores of proven-aligned agents. Given such slow progress, and given the perils of trusting AI systems to do this, I think humans would have to individually program each constraint on the counterlogicals. This raises the possibility that humans decades to centuries in the future may do something incorrectly here, either intentionally or unintentionally, as they constrain and evaluate the counterlogicals, even if they delegate to blinded and self-erasing programmed computers as much as possible. I’m not sure what to do.
- ^
- ^
Note that while, hypothetically, all this fixed point calculation could be done by the agent that has preferences itself (think: a human calculating for itself) in practice only a superintelligence would be able to accurately find a valid fixed point. If a friendly AI was developed, it would presumably do everything required by Nesov’s proposal on behalf of humans, in the background.
- ^
Nesov seems to write like there is only one fixed point, maybe for simplicity, but I don’t see how any practical method would be that precise and accurate. Maybe there would be a “fixed region” in a similar style to the goals achieved under certain proposals for soft optimization.
- ^
The initial aggregate could be considered twin to the prior, though since it can’t be multiplied it can’t be mixed in. This is opposed to the classical pair of prior and utility function. In humans, the situation is presumably much more messy than anything described here, however.
(ancient) discussion here.)
Note that this link is broken. It should go to Eliezer’s top comment here:[1]
https://www.lesswrong.com/posts/SpHYBhkaeDZpZyRvj/what-can-you-do-with-an-unfriendly-ai?commentId=5p7nw3RzLShRftnt8
Vladimir Nesov has a suggestion here about how this could be done[1]. I don’t think it quite works, but to the extent that it is effective, it can be extended beyond just the influence of superintelligence to other types of new territory (and superintelligence as well, since Nesov’s proposal requires a Sysop[2], though presumably with a lot of transhumanist 3+1- or 4-volume locked out by Nesov’s design).
https://www.lesswrong.com/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a?commentId=BjQrqeKfov946oAKj
Creating Friendly AI 1.0 by Eliezer Yudkowsky (2000), Section 5.9.2