I think the whole Forethought report relies on a “tool-AI”-like setup where the AI is under my control, as opposed to doing things autonomously in the world. (Do you agree?) For example, suppose the AI is running 100 copies on 100 servers around the world, paid for by 100 untraceable bank accounts supported by 100 different autonomously-set-up income streams. Now if I say to the AI “I’ll pay you $500 if Circumstance X happens”, the AI has no particular reason to care that I said that. Or more specifically, the AI cares no more about Circumstance X than about the goings-on with its 99 other income streams, related to stock investments, businesses, or whatever else it’s doing, any of which might or might not yield $500 or more under Circumstances A, B, C, D, E…. There’s nothing special about my wishes unless I have the AI locked in a box under my control, such that I’m uniquely capable of controlling its resources.
So in particular:
since the method can be cheaply applied to almost all types of AIs
we are exiting the control regime and are in the deferral/handoff regime where we trust the AI to do long, complex actions like securing the world from out of control AIs/doing hard-to-verify alignment work
Again, once the AI can autonomously do things in the world without a human checking, I think everything else goes out the window. None of our offers are worth anything, because the AI can take what it wants risk-free.
So there’s a dilemma, if we consider a sliding scale of control versus deferral. To the extent that we’re on the “control” side, where humans will not execute the AI’s plans unless the humans understand and approve of the plans, there’s a decent chance that humans will reject good plans, and the AI would rationally choose to suggest sycophantic plans that are more appealing. And to the extent that we’re on the “deferral” side, where humans are just rubber-stamping the AI’s plans without understanding them, then the AI doesn’t even need to “rebel”, the prison doors are open and the AI can simply walk outside risk-free.
Rebellion leads to a 99% of an ideal outcome, 1% unmitigated catastrophe.
Sycophantic advice leads to a 100%-epsilon chance of mitigating catastrophe (advice will be taken but won’t work)
Honestly attempting to reduce the risk leads to a 100%-epsilon chance of mitigating the catastrophe (or in your language we have a 0%+epsilon chance of an unmitigated catastrophe) (the humans always say yes because the humans always defer to the AI…)
I think the third paragraph here is misleading. If you open the prison doors and release the risk-averse ASI into the world, then yes we can be optimistic that this ASI will prevent other out-of-control rogue ASIs from being created, out of its own self-interest. But the AI you just released will itself be an out-of-control misaligned rogue ASI. So that doesn’t count as success for us humans.
I think the whole Forethought report relies on a “tool-AI”-like setup where the AI is under my control, as opposed to doing things autonomously in the world. (Do you agree?)
I’m pretty certain the answer is no, so I don’t agree.
For example, suppose the AI is running 100 copies on 100 servers around the world, paid for by 100 untraceable bank accounts supported by 100 different autonomously-set-up income streams. Now if I say to the AI “I’ll pay you $500 if Circumstance X happens”, the AI has no particular reason to care that I said that. Or more specifically, the AI cares no more about Circumstance X than about the goings-on with its 99 other income streams, related to stock investments, businesses, or whatever else it’s doing, any of which might or might not yield $500 or more under Circumstances A, B, C, D, E…. There’s nothing special about my wishes unless I have the AI locked in a box under my control, such that I’m uniquely capable of controlling its resources.
I’m pretty certain this can’t really happen, because AI behavior under their proposal is invariant to how wealthy the AIs are, so giving them a fixed amount of resources always has the same effect, so it doesn’t matter that the AI has the accounts.
Or in your language, there’s always something special about our offers, no matter what the AI’s level of wealth is.
This is a corollary of the fact that risk-averse agents under the proposal don’t become more willing to take risks as they get richer, which is discussed more in appendix A.2.
Again, once the AI can autonomously do things in the world without a human checking, I think everything else goes out the window. None of our offers are worth anything, because the AI can take what it wants risk-free.
First off, I don’t think this has to necessarily be the case (depending on how largely you define world) because at least in theory, you can have AIs do something autonomously without humans checking (for example a TEDAI that can fully automate AI research, and the software singularity isn’t so fast that it goes from effectively 0% chance of rebellion to 100% chance of rebellion instantly (which is theoretically possible, no comment on whether this does happen in our own world))
That said, I did make a much stronger claim, and I want to justify it here.
An important part to realize is that the key thing here (once we have fixed the values and the risk aversion parameter, since they don’t matter for our discussion) is the ratio between how likely it is for humans to cooperate vs how easy it is for the AI to rebel.
I agree that for superintelligences acting autonomously (absent other superintelligences attempting to take over themselves), taking over the world is pretty risk-free, and this is why I argued we should give the control of the account to the AI, because at that point, we can be rightly confident that the AI does generalize the risk-aversion (because it’s easy to reward accurately, object-level goal misgeneralization no longer matters, and the goal itself is only very, very slightly more complex than a risk-neutral version AI goal, so simplicity arguments are mostly defeated), so in that case, the cooperation probability (by construction) is always equal to or greater than the AI’s chance of rebelling successfully, since the AI can pay itself and humans can no longer not pay the AI, and the lower bound on cooperation is basically random noise from events, which we can reduce exponentially quickly.)
Key point here is that the ratio matters, not the absolute values of cooperation probability vs rebellion probability.
This is enough to disprove the claim that the Forethought report requires a tool-AI setup that don’t take actions, so my claim that it can be applied cheaply to almost all types of AIs still stands.
Thanks!
I think the whole Forethought report relies on a “tool-AI”-like setup where the AI is under my control, as opposed to doing things autonomously in the world. (Do you agree?) For example, suppose the AI is running 100 copies on 100 servers around the world, paid for by 100 untraceable bank accounts supported by 100 different autonomously-set-up income streams. Now if I say to the AI “I’ll pay you $500 if Circumstance X happens”, the AI has no particular reason to care that I said that. Or more specifically, the AI cares no more about Circumstance X than about the goings-on with its 99 other income streams, related to stock investments, businesses, or whatever else it’s doing, any of which might or might not yield $500 or more under Circumstances A, B, C, D, E…. There’s nothing special about my wishes unless I have the AI locked in a box under my control, such that I’m uniquely capable of controlling its resources.
So in particular:
That’s not cheap at all, rather it’s leaving countless trillions of dollars on the table, if not quadrillions etc., because that’s what’s unlocked when AIs can just do stuff in the world (if the AIs aren’t misaligned of course). See “Will almost all companies be founded and run by autonomous AGIs?” here, or “the second piece” in Four ways learning Econ makes people dumber re: future AI.
Again, once the AI can autonomously do things in the world without a human checking, I think everything else goes out the window. None of our offers are worth anything, because the AI can take what it wants risk-free.
So there’s a dilemma, if we consider a sliding scale of control versus deferral. To the extent that we’re on the “control” side, where humans will not execute the AI’s plans unless the humans understand and approve of the plans, there’s a decent chance that humans will reject good plans, and the AI would rationally choose to suggest sycophantic plans that are more appealing. And to the extent that we’re on the “deferral” side, where humans are just rubber-stamping the AI’s plans without understanding them, then the AI doesn’t even need to “rebel”, the prison doors are open and the AI can simply walk outside risk-free.
I think the third paragraph here is misleading. If you open the prison doors and release the risk-averse ASI into the world, then yes we can be optimistic that this ASI will prevent other out-of-control rogue ASIs from being created, out of its own self-interest. But the AI you just released will itself be an out-of-control misaligned rogue ASI. So that doesn’t count as success for us humans.
I’m pretty certain the answer is no, so I don’t agree.
I’m pretty certain this can’t really happen, because AI behavior under their proposal is invariant to how wealthy the AIs are, so giving them a fixed amount of resources always has the same effect, so it doesn’t matter that the AI has the accounts.
Or in your language, there’s always something special about our offers, no matter what the AI’s level of wealth is.
This is a corollary of the fact that risk-averse agents under the proposal don’t become more willing to take risks as they get richer, which is discussed more in appendix A.2.
First off, I don’t think this has to necessarily be the case (depending on how largely you define world) because at least in theory, you can have AIs do something autonomously without humans checking (for example a TEDAI that can fully automate AI research, and the software singularity isn’t so fast that it goes from effectively 0% chance of rebellion to 100% chance of rebellion instantly (which is theoretically possible, no comment on whether this does happen in our own world))
That said, I did make a much stronger claim, and I want to justify it here.
An important part to realize is that the key thing here (once we have fixed the values and the risk aversion parameter, since they don’t matter for our discussion) is the ratio between how likely it is for humans to cooperate vs how easy it is for the AI to rebel.
I agree that for superintelligences acting autonomously (absent other superintelligences attempting to take over themselves), taking over the world is pretty risk-free, and this is why I argued we should give the control of the account to the AI, because at that point, we can be rightly confident that the AI does generalize the risk-aversion (because it’s easy to reward accurately, object-level goal misgeneralization no longer matters, and the goal itself is only very, very slightly more complex than a risk-neutral version AI goal, so simplicity arguments are mostly defeated), so in that case, the cooperation probability (by construction) is always equal to or greater than the AI’s chance of rebelling successfully, since the AI can pay itself and humans can no longer not pay the AI, and the lower bound on cooperation is basically random noise from events, which we can reduce exponentially quickly.)
Key point here is that the ratio matters, not the absolute values of cooperation probability vs rebellion probability.
This is enough to disprove the claim that the Forethought report requires a tool-AI setup that don’t take actions, so my claim that it can be applied cheaply to almost all types of AIs still stands.