It seems great that someone is working on this, but I wonder how optimistic you are, and what your reasons are. My general intuition (in part from the kinds of examples you give) is that the form of the agent and/or goals probably matter quite a bit as far as how easy it is to merge or build/join a coalition (or the cost-benefits of doing so), and once we’re able to build agents of different forms, humans’ form of agency/goals isn’t likely to be optimal as far as building coalitions (and maybe EUMs aren’t optimal either, but something non-human will be), and we’ll face strong incentives to self-modify (or simplify our goals, etc.) before we’re ready. (I guess we see this in companies/countries already, but the problem will get worse with AIs that can explore a larger space of forms of agency/goals.)
Again it’s great that someone is trying to solve this, in case there is a solution, but do you have an argument for being optimistic about this?
One argument for being optimistic: the universe is just very big, and there’s a lot to go around. So there’s a huge amount of room for positive-sum bargaining.
Another: at any given point in time, few of the agents that currently exist would want their goals to become significantly simplified (all else equal). So there’s a strong incentive to coordinate to reduce competition on this axis.
Lastly: if at each point in time, the set of agents who are alive are in conflict with potentially-simpler future agents in a very destructive way, then they should all just Do Something Else. In particular, if there’s some decision-theoretic argument roughly like “more powerful agents should continue to spend some of their resources on the values of their less-powerful ancestors, to reduce the incentives for inter-generational conflict”, even agents with very simple goals might be motivated by it. I call this “the generational contract”.
I buy your arguments for optimism about not needing to simplify/change our goals to compete. (I also think that there are other stronger reasons to expect we don’t need goal simplification like just keeping humans alive and later giving back the resources which is quite simple and indirectly points to what humans want. For ultimately launching space probes, I expect the overhead of complex goals is low. There is some complexity hidden in this proposal, but it seems like it should handle this specific goal simplicity concern.)
I don’t feel compelled by “the universe is very big” arguments for making cooperation look better for me personally because I put most of the weight on linear returns. A few reasons for this:
My sense is that we probably trivially saturate altruistic positive values (I want X to happen) which aren’t very scope sensitive. This doesn’t require any bargaining IMO, it just happens by default due to things like the universe being extremely big (at least tegmark 3 or whatever).
I generally find non-linear returns-ish views pretty non-compelling from a direct moral standpoint.
It seems great that someone is working on this, but I wonder how optimistic you are, and what your reasons are. My general intuition (in part from the kinds of examples you give) is that the form of the agent and/or goals probably matter quite a bit as far as how easy it is to merge or build/join a coalition (or the cost-benefits of doing so), and once we’re able to build agents of different forms, humans’ form of agency/goals isn’t likely to be optimal as far as building coalitions (and maybe EUMs aren’t optimal either, but something non-human will be), and we’ll face strong incentives to self-modify (or simplify our goals, etc.) before we’re ready. (I guess we see this in companies/countries already, but the problem will get worse with AIs that can explore a larger space of forms of agency/goals.)
Again it’s great that someone is trying to solve this, in case there is a solution, but do you have an argument for being optimistic about this?
One argument for being optimistic: the universe is just very big, and there’s a lot to go around. So there’s a huge amount of room for positive-sum bargaining.
Another: at any given point in time, few of the agents that currently exist would want their goals to become significantly simplified (all else equal). So there’s a strong incentive to coordinate to reduce competition on this axis.
Lastly: if at each point in time, the set of agents who are alive are in conflict with potentially-simpler future agents in a very destructive way, then they should all just Do Something Else. In particular, if there’s some decision-theoretic argument roughly like “more powerful agents should continue to spend some of their resources on the values of their less-powerful ancestors, to reduce the incentives for inter-generational conflict”, even agents with very simple goals might be motivated by it. I call this “the generational contract”.
I buy your arguments for optimism about not needing to simplify/change our goals to compete. (I also think that there are other stronger reasons to expect we don’t need goal simplification like just keeping humans alive and later giving back the resources which is quite simple and indirectly points to what humans want. For ultimately launching space probes, I expect the overhead of complex goals is low. There is some complexity hidden in this proposal, but it seems like it should handle this specific goal simplicity concern.)
I don’t feel compelled by “the universe is very big” arguments for making cooperation look better for me personally because I put most of the weight on linear returns. A few reasons for this:
My sense is that we probably trivially saturate altruistic positive values (I want X to happen) which aren’t very scope sensitive. This doesn’t require any bargaining IMO, it just happens by default due to things like the universe being extremely big (at least tegmark 3 or whatever).
I generally find non-linear returns-ish views pretty non-compelling from a direct moral standpoint.