It seems like one implicit assumption is something like
In environments with lots of agents doing things, the most ruthless consequentialist agents will outperform the more prosocial and cooperative ones.
Without that assumption, we could end up in a situation where there are ruthless consequentialist AIs, most agents (both human and AI) recognize them as such and recognizes that interacting with them is a bad idea, and so these ruthless consequentialist AIs backstab and lie and cheat and do a bunch of damage but never actually accumulate enough power to seize control of the light cone from coalitions of agents that are capable of cooperating[1].
I know you’ve written a lot of stuff—I don’t recall seeing anything about why that assumption in particular, but maybe you’ve already written on the topic?
In the long term this probably still looks bad for humans, because worlds with many AI agents probably look like a second cambrian explosion and the cambrian explosion was not good for those who came before
This is an interesting topic, but no, my central expectation (and what I’m arguing for here) is that 100% of the ASIs will be ruthless consequentialists.
Couple little points on that side-track though: (1) Ruthless consequentialist AIs can still copy themselves, and cooperate with their copies, if their goals are non-indexical (which they might or might not be, no opinion off the top of my head), (2) Your comment seems to assume that AIs can read each other’s minds? If they can’t, a smart ruthless consequentialist AI would act in a cooperative and prosocial way in an environment where doing so was to its advantage. I agree that mind-reading is an important dynamic that might change the equilibrium in a multipolar AI world.
“If their goals are non-indexical” seems like quite a big “if”.
Yeah, my modal assumption is that AIs will be able to make fairly strong inferences about the mechanics of the decision processes of other AIs by making observations about their behavior (including of side channels). “Mind reading” might be a slightly strong term for this, but, it’s not very far off.
Likely out of scope for this comment section though. I should, at some point, probably write my modal expectation of what the next couple decades look like in more detail.
It seems like one implicit assumption is something like
Without that assumption, we could end up in a situation where there are ruthless consequentialist AIs, most agents (both human and AI) recognize them as such and recognizes that interacting with them is a bad idea, and so these ruthless consequentialist AIs backstab and lie and cheat and do a bunch of damage but never actually accumulate enough power to seize control of the light cone from coalitions of agents that are capable of cooperating [1] .
I know you’ve written a lot of stuff—I don’t recall seeing anything about why that assumption in particular, but maybe you’ve already written on the topic?
In the long term this probably still looks bad for humans, because worlds with many AI agents probably look like a second cambrian explosion and the cambrian explosion was not good for those who came before
This is an interesting topic, but no, my central expectation (and what I’m arguing for here) is that 100% of the ASIs will be ruthless consequentialists.
Couple little points on that side-track though: (1) Ruthless consequentialist AIs can still copy themselves, and cooperate with their copies, if their goals are non-indexical (which they might or might not be, no opinion off the top of my head), (2) Your comment seems to assume that AIs can read each other’s minds? If they can’t, a smart ruthless consequentialist AI would act in a cooperative and prosocial way in an environment where doing so was to its advantage. I agree that mind-reading is an important dynamic that might change the equilibrium in a multipolar AI world.
Thanks.
“If their goals are non-indexical” seems like quite a big “if”.
Yeah, my modal assumption is that AIs will be able to make fairly strong inferences about the mechanics of the decision processes of other AIs by making observations about their behavior (including of side channels). “Mind reading” might be a slightly strong term for this, but, it’s not very far off.
Likely out of scope for this comment section though. I should, at some point, probably write my modal expectation of what the next couple decades look like in more detail.