Thanks Steve! I found the model behind ruthless blissmaxxing particularly helpful. Most times I have interacted with this it has been unsupported.
A few comments (which may be confounded by a picture of the agent as a connectionist or otherwise complex adaptive cognitive agent).
1. I buy the moral circle failure modes. There’s however a possible failure mode of “graded” moral circles which severely underweigh the moral patienthood of agents far from the circle ( ideologically, geographically etc). One could for example envision a “national securitymaxxing” AGI. While value ethics loaded, such an agent could make decisions based on the gradation of its moral circle field- an inverse square decaying field could be more desirable to a inverse quartic, etc.
2. Because of my complex adaptive confound, I suspect the top and bottom of the figures are structurally unified ie processes which lead to the top “friend/ ally ” half are likely to lead to the development of the bottom “enemy /adversary” half, and this is likely the default. Attempts at separating the two halves are likely to be resisted with increasing capability of the system. This manifests as collaboration/competition likely being scale free.
“Conversely, I’m hardly bothered by publicly disagreeing with someone, even when I know damn well that the someone is going to be annoyed and think less of me—as long as I’m confident that I’m right. “
The stronger Pessimist Byrnes position here would say something like ” It cuts both ways, ie being right is entangled with a shade of expectation that peers converge to your position given enough time etc”
″Note that if the AGI cares too much about the future, then it will try to escape control to take matters into its own hands (which transitions us to §6.2.3); and if it cares too little about the future, it won’t do a good job focusing on what’s important. The optimist says: “there should be a happy medium that splits the difference!”. The pessimist says: “this plan trying to have it both ways, and papering over its self-contradictory incoherence by vague language!”. I tentatively side with the optimist, since some humans seem to be in that happy medium.) “
4. Yes, this is essentially correct given this confound. Complex competent RL systems understand callibrated future discounting.
Thanks Steve! I found the model behind ruthless blissmaxxing particularly helpful. Most times I have interacted with this it has been unsupported.
A few comments (which may be confounded by a picture of the agent as a connectionist or otherwise complex adaptive cognitive agent).
1. I buy the moral circle failure modes. There’s however a possible failure mode of “graded” moral circles which severely underweigh the moral patienthood of agents far from the circle ( ideologically, geographically etc). One could for example envision a “national securitymaxxing” AGI. While value ethics loaded, such an agent could make decisions based on the gradation of its moral circle field- an inverse square decaying field could be more desirable to a inverse quartic, etc.
2. Because of my complex adaptive confound, I suspect the top and bottom of the figures are structurally unified ie processes which lead to the top “friend/ ally ” half are likely to lead to the development of the bottom “enemy /adversary” half, and this is likely the default. Attempts at separating the two halves are likely to be resisted with increasing capability of the system. This manifests as collaboration/competition likely being scale free.
“Conversely, I’m hardly bothered by publicly disagreeing with someone, even when I know damn well that the someone is going to be annoyed and think less of me—as long as I’m confident that I’m right. “
The stronger Pessimist Byrnes position here would say something like ” It cuts both ways, ie being right is entangled with a shade of expectation that peers converge to your position given enough time etc”
″Note that if the AGI cares too much about the future, then it will try to escape control to take matters into its own hands (which transitions us to §6.2.3); and if it cares too little about the future, it won’t do a good job focusing on what’s important. The optimist says: “there should be a happy medium that splits the difference!”. The pessimist says: “this plan trying to have it both ways, and papering over its self-contradictory incoherence by vague language!”. I tentatively side with the optimist, since some humans seem to be in that happy medium.) “
4. Yes, this is essentially correct given this confound. Complex competent RL systems understand callibrated future discounting.