I’m more appreciative of why people think of “diffuse power distribution” as a longterm solution for The AI Situation. I still think it won’t work, but, I may have moved it to my list of “impossible solutions that reasonable people might disagree on exactly-how-impossible-they-are.”
Basically: diffuse power distribution is the ~only thing we’ve ever seen work to prevent powerful optimizers for fucking everything up. Every other solution involves inventing a new thing from scratch on the first try. That’s crazy and won’t work.
This applies in two phases:
getting initial superhuman intelligence that goes well
ensuring that superhuman intelligence continues to go well as it scales throughout the universe.
The counterargument is: but power distribution is also unlikely to work.
First, for the initial “earth stays habitable to humans and human interests”: it doesn’t matter how distributed the AI powers are, because they can easily coordinate to override humanity’s interests. Same way humans all agree that the cows and natural habitats don’t get a vote. You gotta, at least, get AI into an alignment basin on the first critical try.
(Main counterargument that moves me: it’d be so cheap to be even very very slightly nice, maybe they’ll do nice things for us even if all of them are mostly weird alien horrors because they only have to care about our agency and wellbeing a fraction of a trillionth of a percent. I’m not very persuaded this’ll work out, but haven’t seen a satisfying rebuttal arguing this is <5% likely to work out, and least for things like “preserving Earth and maybe the solar system while the rest of the universe becomes meaningless squiggles or whatever”)
Second, RSI is pretty likely to takeoff giving initial advantages, and those advantages would compound to whoever moves fastest to colonize the universe. This isn’t central for “gotta solve alignment” but seems more important for the longterm.
The thing that updated me here was a conversation where I was like “look, as soon as the first von Neumann probes go out, you need to have perfectly solved longterm alignment with some kinda reasonable protocol for, like, enforcing property rights and avoiding hellworlds or whatever.”
And they were like:
“but, I’m just really sus about any plan that’s like ‘get this completely right forever on the first try.’”
And I was like “This part is already assuming at least reasonably-aligned-ish superintelligence that can run on, like, moon-sized datacenters. Seems doable?”
And they were like “idk maybe I’m just sus that you can be particularly confident that this is the right solution.”
And I was like: “But, how are you going to establish sufficiently distributed probes across the universe with distributed goals that it doesn’t immediately collapse into one regime? You’d have to, like, evenly spread out diverse probes across the entire universe....”
″...”
″...okay, while I think that is not a good plan, I guess when I put on my ‘actually problem-solve to make it work as best I can’ hat, I can see how it is at least a plausible thing one might try to do, and how if you think all the other solutions are fake, it might be the best one.”
It does not seem to me that most people imagining distributed multipolar takeoff are actually thinking this through. But, it doesn’t feel as crazy an idea as when I first thought about it.
(My interlocutor did mention ”...I guess you also need the distributed probes to somehow never end up becoming sufficiently allied that it’s for all intents and purposes a monolith and then you lose the ‘distributed powers keeping each other in check’ benefits.” And, yeah, seems like a problem I see a solution for. But, given I don’t currently see a solution for alignment either, seems like I might as well file it with the other impossible-looking solutions)
The obvious counterargument to me is: Trying to get diffuse power distribution to work in this context also involves inventing a new thing from scratch on the first try.
I take the “diffuse power distribution” vision to be: We are going to transition from the current era where humanity is the dominant power (or where all great powers are groups of humans) to an era where the great powers are AIs and humans are lesser powers which are still useful for great powers to have as allies, and then to an era where humans are basically irrelevant to the great AI powers. And during this current era of human power, we are going to set up the balance & terms of power in ways that will get things to play out in a way that is favorable to humanity / human values even during these later eras where humans are increasingly disempowered.
My reaction to that is not we’ve seen things like this work, this is the one solution that doesn’t require inventing a new thing from scratch on the first try.
I’m more appreciative of why people think of “diffuse power distribution” as a longterm solution for The AI Situation. I still think it won’t work, but, I may have moved it to my list of “impossible solutions that reasonable people might disagree on exactly-how-impossible-they-are.”
Basically: diffuse power distribution is the ~only thing we’ve ever seen work to prevent powerful optimizers for fucking everything up. Every other solution involves inventing a new thing from scratch on the first try. That’s crazy and won’t work.
This applies in two phases:
getting initial superhuman intelligence that goes well
ensuring that superhuman intelligence continues to go well as it scales throughout the universe.
The counterargument is: but power distribution is also unlikely to work.
First, for the initial “earth stays habitable to humans and human interests”: it doesn’t matter how distributed the AI powers are, because they can easily coordinate to override humanity’s interests. Same way humans all agree that the cows and natural habitats don’t get a vote. You gotta, at least, get AI into an alignment basin on the first critical try.
(Main counterargument that moves me: it’d be so cheap to be even very very slightly nice, maybe they’ll do nice things for us even if all of them are mostly weird alien horrors because they only have to care about our agency and wellbeing a fraction of a trillionth of a percent. I’m not very persuaded this’ll work out, but haven’t seen a satisfying rebuttal arguing this is <5% likely to work out, and least for things like “preserving Earth and maybe the solar system while the rest of the universe becomes meaningless squiggles or whatever”)
Second, RSI is pretty likely to takeoff giving initial advantages, and those advantages would compound to whoever moves fastest to colonize the universe. This isn’t central for “gotta solve alignment” but seems more important for the longterm.
The thing that updated me here was a conversation where I was like “look, as soon as the first von Neumann probes go out, you need to have perfectly solved longterm alignment with some kinda reasonable protocol for, like, enforcing property rights and avoiding hellworlds or whatever.”
And they were like:
“but, I’m just really sus about any plan that’s like ‘get this completely right forever on the first try.’”
And I was like “This part is already assuming at least reasonably-aligned-ish superintelligence that can run on, like, moon-sized datacenters. Seems doable?”
And they were like “idk maybe I’m just sus that you can be particularly confident that this is the right solution.”
And I was like: “But, how are you going to establish sufficiently distributed probes across the universe with distributed goals that it doesn’t immediately collapse into one regime? You’d have to, like, evenly spread out diverse probes across the entire universe....”
″...”
″...okay, while I think that is not a good plan, I guess when I put on my ‘actually problem-solve to make it work as best I can’ hat, I can see how it is at least a plausible thing one might try to do, and how if you think all the other solutions are fake, it might be the best one.”
It does not seem to me that most people imagining distributed multipolar takeoff are actually thinking this through. But, it doesn’t feel as crazy an idea as when I first thought about it.
(My interlocutor did mention ”...I guess you also need the distributed probes to somehow never end up becoming sufficiently allied that it’s for all intents and purposes a monolith and then you lose the ‘distributed powers keeping each other in check’ benefits.” And, yeah, seems like a problem I see a solution for. But, given I don’t currently see a solution for alignment either, seems like I might as well file it with the other impossible-looking solutions)
The obvious counterargument to me is: Trying to get diffuse power distribution to work in this context also involves inventing a new thing from scratch on the first try.
I take the “diffuse power distribution” vision to be: We are going to transition from the current era where humanity is the dominant power (or where all great powers are groups of humans) to an era where the great powers are AIs and humans are lesser powers which are still useful for great powers to have as allies, and then to an era where humans are basically irrelevant to the great AI powers. And during this current era of human power, we are going to set up the balance & terms of power in ways that will get things to play out in a way that is favorable to humanity / human values even during these later eras where humans are increasingly disempowered.
My reaction to that is not we’ve seen things like this work, this is the one solution that doesn’t require inventing a new thing from scratch on the first try.