I think this bakes in a false assumption that, in the case of corrigible AI, only a very small number of entities will be able to control it. I assume by “the above ignores multipolar scenarios” you still only mean multipolar scenarios with a few key actors. Instead, I think it’s possible for control over AI to be widely distributed among very many, even most, people, and for AI to empower different people and groups to pursue different aims. With increased prosperity comes fewer conflicts over resources, and more people can get what they want in this AI-enabled future. Only a minimal amount of centralized control will be needed to prevent warfare-like applications of AI, but this seems like a manageable problem. And just because we have to empower a smaller number of people to prevent these catastrophically destructive use-cases, doesn’t mean those people have to control everything.
So I don’t think ASI means we need to rely on any particular individual or group having “humanity’s best interests in mind”. Decentralized free pursuit of individual goals has resulted in improved QoL across the world historically, and I expect that to continue.
I claim that widely distributing full access to corrigible AI falls to bad actors ruining things for everyone, less like conflict over resources but more like doing selfish, negative-sum activities that get amplified by offense-dominant technologies enabled by AI. Any effective effort to prevent this either renders decentralization moot in the first place, or requires the AI to make its own moral judgement, i.e. value alignment. (In the value alignment case it can more easily be argued either way whether decentralized or centralized control is better, but in the corrigibility case it’s pretty clear that decentralized control over corrigible AI is very, very bad due to misuse risks)
From your original post:
I think you can corrigibilitymaxx and still prevent catastrophic misuse via system level measures and deployment controls. In particular, we can restrict who has full access to the model’s inputs).
I claim that “restrict[ing] who has full access to the model’s inputs” implies that the model is actually corrigible to a central authority, instead of individual users. Therefore, we’re forced to adopt a system of one or a few actors controlling the model.
I think this bakes in a false assumption that, in the case of corrigible AI, only a very small number of entities will be able to control it. I assume by “the above ignores multipolar scenarios” you still only mean multipolar scenarios with a few key actors. Instead, I think it’s possible for control over AI to be widely distributed among very many, even most, people, and for AI to empower different people and groups to pursue different aims. With increased prosperity comes fewer conflicts over resources, and more people can get what they want in this AI-enabled future. Only a minimal amount of centralized control will be needed to prevent warfare-like applications of AI, but this seems like a manageable problem. And just because we have to empower a smaller number of people to prevent these catastrophically destructive use-cases, doesn’t mean those people have to control everything.
So I don’t think ASI means we need to rely on any particular individual or group having “humanity’s best interests in mind”. Decentralized free pursuit of individual goals has resulted in improved QoL across the world historically, and I expect that to continue.
I claim that widely distributing full access to corrigible AI falls to bad actors ruining things for everyone, less like conflict over resources but more like doing selfish, negative-sum activities that get amplified by offense-dominant technologies enabled by AI. Any effective effort to prevent this either renders decentralization moot in the first place, or requires the AI to make its own moral judgement, i.e. value alignment. (In the value alignment case it can more easily be argued either way whether decentralized or centralized control is better, but in the corrigibility case it’s pretty clear that decentralized control over corrigible AI is very, very bad due to misuse risks)
From your original post:
I claim that “restrict[ing] who has full access to the model’s inputs” implies that the model is actually corrigible to a central authority, instead of individual users. Therefore, we’re forced to adopt a system of one or a few actors controlling the model.