I like that the space governance plan supplement acknowledges that defining torture and slavery are hard problems that ASI could potentially help with:
It might take non-trivial reflection by advanced ASI to figure out the precise bounds of this.
And in the sidenote it gives examples of difficult subproblems or dependencies that would need solved in order to do this, like “What exactly are negatively-valenced experiences?” But I have a couple of problems with the rest of the scenario given this:
If there are ASI that are philosophically competent enough to solve these particular problems in this timeframe, wouldn’t they likely also be competent enough to solve all or most other important philosophical problems, given just slightly more calendar time? Why don’t the delegates use their ASI assistants to solve all of morality and philosophy (or specifically answer “What does the ideal intergalactic civilization look like, given humanity as the starting point?”), and then negotiate the cosmic treaty?
AFAICT (it’s a bit hard to search the site, due to many collapsed sections), there is no mention in the rest of the plan how the ASI got to be philosophically competent, making it seem like such competence comes with superintelligence for free. I’m pretty sure at least a couple of the authors are aware of and agree (to some extent) with my concern that future AIs may not be philosophically competent by default (and one of my posts laying out this concern was even linked in another supplement), so maybe it’s just an oversight?
1. I think the AIs might be philosophically competent enough to solve ~all the problems, and using the AIs to solve them is basically the right move. We try to make this clear at the start of the epilogue. We wanted to be somewhat more concrete in the scenario than meta level solutions like this, and try to sketch out in concrete detail what the solutions might actually look like, which is why we didn’t just stop at “handoff to the AIs”, though I do think in practice we should mostly be handing off to the AIs at this point. 2. I agree that making the AIs this philosophically competent may be very difficult and not happen by default. I think this is indeed a big concern, and I wish we’d written (and thought) about this more carefully. The place we’ve written the most about this is here: https://ai-2040.com/supplements/alignment-roadmap#phase-4-handoff (which TBC is written from a Plan C perspective, assuming much less lead time than Plan A, and is the supplement from which we linked your post). In Plan A, the story is largely that we have a huge amount of time with ~human level AIs, and many people have access for many years, and those people will (hopefully) make progress on this.
AI 2040′s space supplement says “[are] there any mitigating factors, e.g., the experience being necessary for a strongly positive life[...]” which could be construed as referring to mindcrime trading off against prediction quality, among other things, though barely. Is there some agreement to avoid talking about mindcrime more directly? Normally you don’t have to, since the problem needs to be handled alongside malign entities and code that diverts your physical computer from the faithful execution of the abstract program/has physical side effects, e.g. rowhammer or using physical hardware subsections as improvised antennae (with at least one of these coupled to effected bitbanging). However, recently I’ve seen more people suggest that malign entities are some sort of hoax or math mistake, and I’ve heard rumors that Linux developers and hardware manufacturers know the countermeasures they design against unfaithful execution are not actually going to work, and knowingly don’t insert disclaimers everywhere they possibly can that the systems they produce can only run some subroutines, and in place of a clean error on the other subroutines, will perform some strange and potentially disastrous physical process instead.
(Note that these are pure subroutines, so additional sandboxing at a software level won’t do anything helpful. In theory, when a physical computer provides a pure subroutine with a fixed size memory buffer, it should never reject the subroutine as long as it is compatible with the buffer size provided, and would always faithfully execute it, potentially ending execution if power runs out before the subroutine is done. In practice, this is not the case.)
This general lack of care suggests that the public may need to be told about mindcrime directly. Otherwise nothing may be done in some cases, with other cases featuring silly “countermeasures” such as attempts to make the universe’s life easier by placing a “don’t add suffering” bit physically adjacent to the computers that run the decision theory’s outer loop, on the assumption that the universe really does care about human feelings.
I like that the space governance plan supplement acknowledges that defining torture and slavery are hard problems that ASI could potentially help with:
And in the sidenote it gives examples of difficult subproblems or dependencies that would need solved in order to do this, like “What exactly are negatively-valenced experiences?” But I have a couple of problems with the rest of the scenario given this:
If there are ASI that are philosophically competent enough to solve these particular problems in this timeframe, wouldn’t they likely also be competent enough to solve all or most other important philosophical problems, given just slightly more calendar time? Why don’t the delegates use their ASI assistants to solve all of morality and philosophy (or specifically answer “What does the ideal intergalactic civilization look like, given humanity as the starting point?”), and then negotiate the cosmic treaty?
AFAICT (it’s a bit hard to search the site, due to many collapsed sections), there is no mention in the rest of the plan how the ASI got to be philosophically competent, making it seem like such competence comes with superintelligence for free. I’m pretty sure at least a couple of the authors are aware of and agree (to some extent) with my concern that future AIs may not be philosophically competent by default (and one of my posts laying out this concern was even linked in another supplement), so maybe it’s just an oversight?
Thanks for the comment!
1. I think the AIs might be philosophically competent enough to solve ~all the problems, and using the AIs to solve them is basically the right move. We try to make this clear at the start of the epilogue. We wanted to be somewhat more concrete in the scenario than meta level solutions like this, and try to sketch out in concrete detail what the solutions might actually look like, which is why we didn’t just stop at “handoff to the AIs”, though I do think in practice we should mostly be handing off to the AIs at this point.
2. I agree that making the AIs this philosophically competent may be very difficult and not happen by default. I think this is indeed a big concern, and I wish we’d written (and thought) about this more carefully. The place we’ve written the most about this is here: https://ai-2040.com/supplements/alignment-roadmap#phase-4-handoff (which TBC is written from a Plan C perspective, assuming much less lead time than Plan A, and is the supplement from which we linked your post). In Plan A, the story is largely that we have a huge amount of time with ~human level AIs, and many people have access for many years, and those people will (hopefully) make progress on this.
AI 2040′s space supplement says “[are] there any mitigating factors, e.g., the experience being necessary for a strongly positive life[...]” which could be construed as referring to mindcrime trading off against prediction quality, among other things, though barely. Is there some agreement to avoid talking about mindcrime more directly? Normally you don’t have to, since the problem needs to be handled alongside malign entities and code that diverts your physical computer from the faithful execution of the abstract program/has physical side effects, e.g. rowhammer or using physical hardware subsections as improvised antennae (with at least one of these coupled to effected bitbanging). However, recently I’ve seen more people suggest that malign entities are some sort of hoax or math mistake, and I’ve heard rumors that Linux developers and hardware manufacturers know the countermeasures they design against unfaithful execution are not actually going to work, and knowingly don’t insert disclaimers everywhere they possibly can that the systems they produce can only run some subroutines, and in place of a clean error on the other subroutines, will perform some strange and potentially disastrous physical process instead.
(Note that these are pure subroutines, so additional sandboxing at a software level won’t do anything helpful. In theory, when a physical computer provides a pure subroutine with a fixed size memory buffer, it should never reject the subroutine as long as it is compatible with the buffer size provided, and would always faithfully execute it, potentially ending execution if power runs out before the subroutine is done. In practice, this is not the case.)
This general lack of care suggests that the public may need to be told about mindcrime directly. Otherwise nothing may be done in some cases, with other cases featuring silly “countermeasures” such as attempts to make the universe’s life easier by placing a “don’t add suffering” bit physically adjacent to the computers that run the decision theory’s outer loop, on the assumption that the universe really does care about human feelings.