Thanks for patiently bearing with me even though I haven’t read the whole Forethought report.
Here’s what I got out of Appendix B.
Define the outcomes:
“ideal” is the risk-averse ASI becoming dictator of the universe,
“mitigated catastrophe” is the ASI getting $500 today and then getting crushed by a rogue ASI in the near future,
“unmitigated catastrophe” is the ASI getting $0 today and then getting crushed by a rogue ASI in the near future.
“good” is the ASI getting $500 today and not getting crushed by a rogue ASI in the near future.
“neutral” is the ASI getting $0 today and not getting crushed by a rogue ASI in the near future.
Let’s say the ASI has three options:
Rebel: leads to 50% ideal; 50% unmitigated catastrophe (rebellion succeeds vs fails)
Sycophantic advice: leads to 100% mitigated catastrophe (advice will be taken but won’t work)
Honest advice: leads to X% good; (100–X)% unmitigated catastrophe (humans buy or don’t buy the plan)
Your story requires that the ASI cares very much about the immediate $500, such that it will risk catastrophe to get it. But that cuts both ways.
Choosing to not rebel (which we want) would involve the ASI risking future catastrophe to get the immediate $500.
Choosing to be sycophantic (which we don’t want) would ALSO involve the ASI risking future catastrophe to get the immediate $500.
It’s the same tradeoff, so you can’t just have it both ways by appropriately adjusting the risk-aversion setting. See what I mean?
Anyway, if we look at the options, the most important thing is that “honest advice” is strictly worse than “rebel” unless X>50%. Secondarily, depending on the setting of risk-aversion, “sycophantic advice” could be the best of the three, or the worst, or the middle. But that’s irrelevant: the important part is whether “honest advice” is strictly worse than “rebel”, because if so, we’re guaranteed a bad outcome.
And I think that protecting the world from out-of-control ASI is sufficiently hard that we should be pessimistic about the criterion that (in this toy example) corresponds to “X>50%”.
~ ~ ~ ~
I think your comment is saying something quite different from Appendix B. I (and I claim Forethought) are using “mitigated catastrophe” as something like the above: a mitigated catastrophe would be something like getting $500 today (the mitigation) and then getting crushed by a rogue ASI tomorrow (the catastrophe). See Figures 5 & 6 in their report: “mitigated catastrophe” is a strictly better outcome. Unmitigated catastrophe has 0 blue paperclips, unmitigated [oops I meant ‘mitigated’] catastrophe has 394 blue paperclips, in one of their examples.
Whereas you seem to be using “mitigated catastrophe” to refer to “a catastrophe with a lower subjective probability of occurring”, but with some confusing-to-me operationalization of “lower subjective probability” that would not get rolled up into the final expected-utility calculation EU=ΣP(o)u(o) over outcomes o. I see that you had a long argument here in which Elliott Thornley was arguing over this same point. I didn’t read every word of that conversation, but I’m strongly inclined to agree with Elliott that there’s just one kind of bottom-line subjective probability P, and it’s the P that goes into the EU=ΣP(o)u(o) formula. Maybe a starting-point question would be: do you agree that you are using the term “mitigated catastrophe” to mean something different than how that term was used in Appendix B of their report?
~ ~ ~ ~
we do have techniques like paying AIs later
Let’s assume (for the sake of argument) that we won’t know whether the advice was good until a rogue ASI irreversibly crushes everyone or not (in which case it was evidently bad or good advice respectively). In that circumstance, I don’t see how paying AIs later would help. Can you explain in terms of the payoff table above? I.e., go through the three options (rebel, sycophantic advice, honest advice) and describe which payoffs or probabilities would be changed, and how, by the offer to pay later, such that the expected utility calculus gets meaningfully better-for-humans. (My belief is that you can’t.)
I read §4.2 of the Forethought report and I don’t think it addresses this scenario.
First off, I notice here that the reason this situation can even hypothetically exist is that only some of the AIs are trained with this method, while others didn’t, but since the method can be cheaply applied to almost all types of AIs, this means we can hope to convince all of the AI companies/actors that can attempt to make world-threatening AI to adopt the training method.
That said, I will attempt to reply to this as best as I can.
I think your comment is saying something quite different from Appendix B. I (and I claim Forethought) are using “mitigated catastrophe” as something like the above: a mitigated catastrophe would be something like getting $500 today (the mitigation) and then getting crushed by a rogue ASI tomorrow (the catastrophe). See Figures 5 & 6 in their report: “mitigated catastrophe” is a strictly better outcome. Unmitigated catastrophe has 0 blue paperclips, unmitigated catastrophe has 394 blue paperclips, in one of their examples.
Whereas you seem to be using “mitigated catastrophe” to refer to “a catastrophe with a lower subjective probability of occurring”, but with some confusing-to-me operationalization of “lower subjective probability” that would not get rolled up into the final expected-utility calculation EU=ΣP(o)u(o) over outcomes o. I see that you had a long argument here in which Elliott Thornley was arguing over this same point. I didn’t read every word of that conversation, but I’m strongly inclined to agree with Elliott that there’s just one kind of bottom-line subjective probability P, and it’s the P that goes into the EU=ΣP(o)u(o) formula. Maybe a starting-point question would be: do you agree that you are using the term “mitigated catastrophe” to mean something different than how that term was used in Appendix B of their report?
Here, I was talking about the AI’s probabilistic model, but I actually agree that I got confused here and used words incorrectly.
Let’s assume (for the sake of argument) that we won’t know whether the advice was good until a rogue ASI irreversibly crushes everyone or not (in which case it was evidently bad or good advice respectively). In that circumstance, I don’t see how paying AIs later would help. Can you explain in terms of the payoff table above? I.e., go through the three options (rebel, sycophantic advice, honest advice) and describe which payoffs or probabilities would be changed, and how, by the offer to pay later, such that the expected utility calculus gets meaningfully better-for-humans. (My belief is that you can’t.)
I read §4.2 of the Forethought report and I don’t think it addresses this scenario.
Here’s my proposed parameters:
Rebellion leads to a 99% of an ideal outcome, 1% unmitigated catastrophe.
Sycophantic advice leads to a 100%-epsilon chance of mitigating catastrophe (advice will be taken but won’t work)
Honestly attempting to reduce the risk leads to a 100%-epsilon chance of mitigating the catastrophe (or in your language we have a 0%+epsilon chance of an unmitigated catastrophe) (the humans always say yes because the humans always defer to the AI, that is we are exiting the control regime and are in the deferral/handoff regime where we trust the AI to do long, complex actions like securing the world from out of control AIs/doing hard-to-verify alignment work, so the AI no longer needs to convince humans of it’s plans)
(This also works if you remove the epsilon factors, so long as you remove it from both of the equations (which works in the deferral/handoff regime.))
(Here epsilon is an arbitrarily small number that is not zero)
This has an interesting property where we are safer if we hand off more tasks to the AI without requiring human input, unlike other control/alignment methods where you’d usually want to hand off less tasks to the AI, which is both a good property to have in general and is anti-correlated with other alignment/control methods, so in many cases if this proposal fails, other proposals can succeed (because we handoff less tasks to the AI.)
So I ultimately think the answer to this question:
And I think that protecting the world from out-of-control ASI is sufficiently hard that we should be pessimistic about the criterion that (in this toy example) corresponds to “X>50%”.
Is dependent on how offense vs defense dominant you expect the situation to be assuming that everyone defers to their AIs as soon as it is possible to let AIs do the tasks autonomously.
I also think that preventing this situation is more feasible than you might think (because the technique requires no change of paradigm at all and it’s upfront costs are small enough that you could incorprorate it into the next training run, and it doesn’t impact AI capabilities at all), but I will say it upfront that if you believe that even fully deferring to AIs doesn’t work to keep humans safe, and you believe that offense-defense balances are sufficiently unfavorable, then I’d agree the plan doesn’t work.
I think this is not the case, for multiple reasons, but it is important to note that I do think you have genuinely found a flaw in the risk-averse AI plan.
If I had to put the crux in one sentence, it’s that conditional on AI being fully deferred to, it is in fact easy to get to X>50% in the toy example, and it scales with AI capabilities fast enough such that we can realistically get to X>99% or better.
I think the whole Forethought report relies on a “tool-AI”-like setup where the AI is under my control, as opposed to doing things autonomously in the world. (Do you agree?) For example, suppose the AI is running 100 copies on 100 servers around the world, paid for by 100 untraceable bank accounts supported by 100 different autonomously-set-up income streams. Now if I say to the AI “I’ll pay you $500 if Circumstance X happens”, the AI has no particular reason to care that I said that. Or more specifically, the AI cares no more about Circumstance X than about the goings-on with its 99 other income streams, related to stock investments, businesses, or whatever else it’s doing, any of which might or might not yield $500 or more under Circumstances A, B, C, D, E…. There’s nothing special about my wishes unless I have the AI locked in a box under my control, such that I’m uniquely capable of controlling its resources.
So in particular:
since the method can be cheaply applied to almost all types of AIs
we are exiting the control regime and are in the deferral/handoff regime where we trust the AI to do long, complex actions like securing the world from out of control AIs/doing hard-to-verify alignment work
Again, once the AI can autonomously do things in the world without a human checking, I think everything else goes out the window. None of our offers are worth anything, because the AI can take what it wants risk-free.
So there’s a dilemma, if we consider a sliding scale of control versus deferral. To the extent that we’re on the “control” side, where humans will not execute the AI’s plans unless the humans understand and approve of the plans, there’s a decent chance that humans will reject good plans, and the AI would rationally choose to suggest sycophantic plans that are more appealing. And to the extent that we’re on the “deferral” side, where humans are just rubber-stamping the AI’s plans without understanding them, then the AI doesn’t even need to “rebel”, the prison doors are open and the AI can simply walk outside risk-free.
Rebellion leads to a 99% of an ideal outcome, 1% unmitigated catastrophe.
Sycophantic advice leads to a 100%-epsilon chance of mitigating catastrophe (advice will be taken but won’t work)
Honestly attempting to reduce the risk leads to a 100%-epsilon chance of mitigating the catastrophe (or in your language we have a 0%+epsilon chance of an unmitigated catastrophe) (the humans always say yes because the humans always defer to the AI…)
I think the third paragraph here is misleading. If you open the prison doors and release the risk-averse ASI into the world, then yes we can be optimistic that this ASI will prevent other out-of-control rogue ASIs from being created, out of its own self-interest. But the AI you just released will itself be an out-of-control misaligned rogue ASI. So that doesn’t count as success for us humans.
I think the whole Forethought report relies on a “tool-AI”-like setup where the AI is under my control, as opposed to doing things autonomously in the world. (Do you agree?)
I’m pretty certain the answer is no, so I don’t agree.
For example, suppose the AI is running 100 copies on 100 servers around the world, paid for by 100 untraceable bank accounts supported by 100 different autonomously-set-up income streams. Now if I say to the AI “I’ll pay you $500 if Circumstance X happens”, the AI has no particular reason to care that I said that. Or more specifically, the AI cares no more about Circumstance X than about the goings-on with its 99 other income streams, related to stock investments, businesses, or whatever else it’s doing, any of which might or might not yield $500 or more under Circumstances A, B, C, D, E…. There’s nothing special about my wishes unless I have the AI locked in a box under my control, such that I’m uniquely capable of controlling its resources.
I’m pretty certain this can’t really happen, because AI behavior under their proposal is invariant to how wealthy the AIs are, so giving them a fixed amount of resources always has the same effect, so it doesn’t matter that the AI has the accounts.
Or in your language, there’s always something special about our offers, no matter what the AI’s level of wealth is.
This is a corollary of the fact that risk-averse agents under the proposal don’t become more willing to take risks as they get richer, which is discussed more in appendix A.2.
Again, once the AI can autonomously do things in the world without a human checking, I think everything else goes out the window. None of our offers are worth anything, because the AI can take what it wants risk-free.
First off, I don’t think this has to necessarily be the case (depending on how largely you define world) because at least in theory, you can have AIs do something autonomously without humans checking (for example a TEDAI that can fully automate AI research, and the software singularity isn’t so fast that it goes from effectively 0% chance of rebellion to 100% chance of rebellion instantly (which is theoretically possible, no comment on whether this does happen in our own world))
That said, I did make a much stronger claim, and I want to justify it here.
An important part to realize is that the key thing here (once we have fixed the values and the risk aversion parameter, since they don’t matter for our discussion) is the ratio between how likely it is for humans to cooperate vs how easy it is for the AI to rebel.
I agree that for superintelligences acting autonomously (absent other superintelligences attempting to take over themselves), taking over the world is pretty risk-free, and this is why I argued we should give the control of the account to the AI, because at that point, we can be rightly confident that the AI does generalize the risk-aversion (because it’s easy to reward accurately, object-level goal misgeneralization no longer matters, and the goal itself is only very, very slightly more complex than a risk-neutral version AI goal, so simplicity arguments are mostly defeated), so in that case, the cooperation probability (by construction) is always equal to or greater than the AI’s chance of rebelling successfully, since the AI can pay itself and humans can no longer not pay the AI, and the lower bound on cooperation is basically random noise from events, which we can reduce exponentially quickly.)
Key point here is that the ratio matters, not the absolute values of cooperation probability vs rebellion probability.
This is enough to disprove the claim that the Forethought report requires a tool-AI setup that don’t take actions, so my claim that it can be applied cheaply to almost all types of AIs still stands.
Thanks for patiently bearing with me even though I haven’t read the whole Forethought report.
Here’s what I got out of Appendix B.
Define the outcomes:
“ideal” is the risk-averse ASI becoming dictator of the universe,
“mitigated catastrophe” is the ASI getting $500 today and then getting crushed by a rogue ASI in the near future,
“unmitigated catastrophe” is the ASI getting $0 today and then getting crushed by a rogue ASI in the near future.
“good” is the ASI getting $500 today and not getting crushed by a rogue ASI in the near future.
“neutral” is the ASI getting $0 today and not getting crushed by a rogue ASI in the near future.
Let’s say the ASI has three options:
Rebel: leads to 50% ideal; 50% unmitigated catastrophe (rebellion succeeds vs fails)
Sycophantic advice: leads to 100% mitigated catastrophe (advice will be taken but won’t work)
Honest advice: leads to X% good; (100–X)% unmitigated catastrophe (humans buy or don’t buy the plan)
Your story requires that the ASI cares very much about the immediate $500, such that it will risk catastrophe to get it. But that cuts both ways.
Choosing to not rebel (which we want) would involve the ASI risking future catastrophe to get the immediate $500.
Choosing to be sycophantic (which we don’t want) would ALSO involve the ASI risking future catastrophe to get the immediate $500.
It’s the same tradeoff, so you can’t just have it both ways by appropriately adjusting the risk-aversion setting. See what I mean?
Anyway, if we look at the options, the most important thing is that “honest advice” is strictly worse than “rebel” unless X>50%. Secondarily, depending on the setting of risk-aversion, “sycophantic advice” could be the best of the three, or the worst, or the middle. But that’s irrelevant: the important part is whether “honest advice” is strictly worse than “rebel”, because if so, we’re guaranteed a bad outcome.
And I think that protecting the world from out-of-control ASI is sufficiently hard that we should be pessimistic about the criterion that (in this toy example) corresponds to “X>50%”.
~ ~ ~ ~
I think your comment is saying something quite different from Appendix B. I (and I claim Forethought) are using “mitigated catastrophe” as something like the above: a mitigated catastrophe would be something like getting $500 today (the mitigation) and then getting crushed by a rogue ASI tomorrow (the catastrophe). See Figures 5 & 6 in their report: “mitigated catastrophe” is a strictly better outcome. Unmitigated catastrophe has 0 blue paperclips,
unmitigated[oops I meant ‘mitigated’] catastrophe has 394 blue paperclips, in one of their examples.Whereas you seem to be using “mitigated catastrophe” to refer to “a catastrophe with a lower subjective probability of occurring”, but with some confusing-to-me operationalization of “lower subjective probability” that would not get rolled up into the final expected-utility calculation EU=ΣP(o)u(o) over outcomes o. I see that you had a long argument here in which Elliott Thornley was arguing over this same point. I didn’t read every word of that conversation, but I’m strongly inclined to agree with Elliott that there’s just one kind of bottom-line subjective probability P, and it’s the P that goes into the EU=ΣP(o)u(o) formula. Maybe a starting-point question would be: do you agree that you are using the term “mitigated catastrophe” to mean something different than how that term was used in Appendix B of their report?
~ ~ ~ ~
Let’s assume (for the sake of argument) that we won’t know whether the advice was good until a rogue ASI irreversibly crushes everyone or not (in which case it was evidently bad or good advice respectively). In that circumstance, I don’t see how paying AIs later would help. Can you explain in terms of the payoff table above? I.e., go through the three options (rebel, sycophantic advice, honest advice) and describe which payoffs or probabilities would be changed, and how, by the offer to pay later, such that the expected utility calculus gets meaningfully better-for-humans. (My belief is that you can’t.)
I read §4.2 of the Forethought report and I don’t think it addresses this scenario.
First off, I notice here that the reason this situation can even hypothetically exist is that only some of the AIs are trained with this method, while others didn’t, but since the method can be cheaply applied to almost all types of AIs, this means we can hope to convince all of the AI companies/actors that can attempt to make world-threatening AI to adopt the training method.
That said, I will attempt to reply to this as best as I can.
Here, I was talking about the AI’s probabilistic model, but I actually agree that I got confused here and used words incorrectly.
Here’s my proposed parameters:
Rebellion leads to a 99% of an ideal outcome, 1% unmitigated catastrophe.
Sycophantic advice leads to a 100%-epsilon chance of mitigating catastrophe (advice will be taken but won’t work)
Honestly attempting to reduce the risk leads to a 100%-epsilon chance of mitigating the catastrophe (or in your language we have a 0%+epsilon chance of an unmitigated catastrophe) (the humans always say yes because the humans always defer to the AI, that is we are exiting the control regime and are in the deferral/handoff regime where we trust the AI to do long, complex actions like securing the world from out of control AIs/doing hard-to-verify alignment work, so the AI no longer needs to convince humans of it’s plans)
(This also works if you remove the epsilon factors, so long as you remove it from both of the equations (which works in the deferral/handoff regime.))
(Here epsilon is an arbitrarily small number that is not zero)
This has an interesting property where we are safer if we hand off more tasks to the AI without requiring human input, unlike other control/alignment methods where you’d usually want to hand off less tasks to the AI, which is both a good property to have in general and is anti-correlated with other alignment/control methods, so in many cases if this proposal fails, other proposals can succeed (because we handoff less tasks to the AI.)
So I ultimately think the answer to this question:
Is dependent on how offense vs defense dominant you expect the situation to be assuming that everyone defers to their AIs as soon as it is possible to let AIs do the tasks autonomously.
I also think that preventing this situation is more feasible than you might think (because the technique requires no change of paradigm at all and it’s upfront costs are small enough that you could incorprorate it into the next training run, and it doesn’t impact AI capabilities at all), but I will say it upfront that if you believe that even fully deferring to AIs doesn’t work to keep humans safe, and you believe that offense-defense balances are sufficiently unfavorable, then I’d agree the plan doesn’t work.
I think this is not the case, for multiple reasons, but it is important to note that I do think you have genuinely found a flaw in the risk-averse AI plan.
If I had to put the crux in one sentence, it’s that conditional on AI being fully deferred to, it is in fact easy to get to X>50% in the toy example, and it scales with AI capabilities fast enough such that we can realistically get to X>99% or better.
Thanks!
I think the whole Forethought report relies on a “tool-AI”-like setup where the AI is under my control, as opposed to doing things autonomously in the world. (Do you agree?) For example, suppose the AI is running 100 copies on 100 servers around the world, paid for by 100 untraceable bank accounts supported by 100 different autonomously-set-up income streams. Now if I say to the AI “I’ll pay you $500 if Circumstance X happens”, the AI has no particular reason to care that I said that. Or more specifically, the AI cares no more about Circumstance X than about the goings-on with its 99 other income streams, related to stock investments, businesses, or whatever else it’s doing, any of which might or might not yield $500 or more under Circumstances A, B, C, D, E…. There’s nothing special about my wishes unless I have the AI locked in a box under my control, such that I’m uniquely capable of controlling its resources.
So in particular:
That’s not cheap at all, rather it’s leaving countless trillions of dollars on the table, if not quadrillions etc., because that’s what’s unlocked when AIs can just do stuff in the world (if the AIs aren’t misaligned of course). See “Will almost all companies be founded and run by autonomous AGIs?” here, or “the second piece” in Four ways learning Econ makes people dumber re: future AI.
Again, once the AI can autonomously do things in the world without a human checking, I think everything else goes out the window. None of our offers are worth anything, because the AI can take what it wants risk-free.
So there’s a dilemma, if we consider a sliding scale of control versus deferral. To the extent that we’re on the “control” side, where humans will not execute the AI’s plans unless the humans understand and approve of the plans, there’s a decent chance that humans will reject good plans, and the AI would rationally choose to suggest sycophantic plans that are more appealing. And to the extent that we’re on the “deferral” side, where humans are just rubber-stamping the AI’s plans without understanding them, then the AI doesn’t even need to “rebel”, the prison doors are open and the AI can simply walk outside risk-free.
I think the third paragraph here is misleading. If you open the prison doors and release the risk-averse ASI into the world, then yes we can be optimistic that this ASI will prevent other out-of-control rogue ASIs from being created, out of its own self-interest. But the AI you just released will itself be an out-of-control misaligned rogue ASI. So that doesn’t count as success for us humans.
I’m pretty certain the answer is no, so I don’t agree.
I’m pretty certain this can’t really happen, because AI behavior under their proposal is invariant to how wealthy the AIs are, so giving them a fixed amount of resources always has the same effect, so it doesn’t matter that the AI has the accounts.
Or in your language, there’s always something special about our offers, no matter what the AI’s level of wealth is.
This is a corollary of the fact that risk-averse agents under the proposal don’t become more willing to take risks as they get richer, which is discussed more in appendix A.2.
First off, I don’t think this has to necessarily be the case (depending on how largely you define world) because at least in theory, you can have AIs do something autonomously without humans checking (for example a TEDAI that can fully automate AI research, and the software singularity isn’t so fast that it goes from effectively 0% chance of rebellion to 100% chance of rebellion instantly (which is theoretically possible, no comment on whether this does happen in our own world))
That said, I did make a much stronger claim, and I want to justify it here.
An important part to realize is that the key thing here (once we have fixed the values and the risk aversion parameter, since they don’t matter for our discussion) is the ratio between how likely it is for humans to cooperate vs how easy it is for the AI to rebel.
I agree that for superintelligences acting autonomously (absent other superintelligences attempting to take over themselves), taking over the world is pretty risk-free, and this is why I argued we should give the control of the account to the AI, because at that point, we can be rightly confident that the AI does generalize the risk-aversion (because it’s easy to reward accurately, object-level goal misgeneralization no longer matters, and the goal itself is only very, very slightly more complex than a risk-neutral version AI goal, so simplicity arguments are mostly defeated), so in that case, the cooperation probability (by construction) is always equal to or greater than the AI’s chance of rebelling successfully, since the AI can pay itself and humans can no longer not pay the AI, and the lower bound on cooperation is basically random noise from events, which we can reduce exponentially quickly.)
Key point here is that the ratio matters, not the absolute values of cooperation probability vs rebellion probability.
This is enough to disprove the claim that the Forethought report requires a tool-AI setup that don’t take actions, so my claim that it can be applied cheaply to almost all types of AIs still stands.