Human morality basically evolved as reputation management: if I get a reputation for being a bad partner to cooperate with, other people will make things go badly for me, so continuing to be thought of as a fine upstanding citizen and good friend is of real value to me. So it is built on an enforcement mechanism (originally gossip, shunning, and occasional violence, to which we have more recently added law-enforcement and the press) which basically assumes that there are agents of roughly comparable capability actively monitoring my behavior for whether it adheres to my society’s standards or not. (Note that unlike formal law, morality has, and needs to have, fuzzy edges and discretion, to avoid being loopholed — in law enforcement this comes down to selective enforcement (police and prosecutor discretion) and then judge/jury discretion.)
Morality is not something that arises naturally from a single agent — it’s the evolved outcome of a society of agents interacting and competing/cooperating/monitoring each other, in a way that turned out to have a fairly prosocial stable equilibrium. Us successfully enforcing morality on agents a lot smarter than us seems impractical. Us using honest, upstanding, moral ASIs to do monitoring and law enforcement on other more concerning ASIs of comparable capability seems a lot more practical. Debate is a rather simplistic version of this. This obviously raises a “who watches the watchmen?” problem. We need to devise a system of circular/mutual enforcement: different ASIs that monitor each other and keep each other honest/aligned, a system that has a stable state which is both aligned with our interests, and has some dynamic reason to stay aligned with our interests (such as that we have enough influence within the system to provide feedback against drift). Understanding how and why human societies (usually) have prosocial equilibria (as is studied in evolutionary psychology/moral anthropology/sociology of morality) seems likely to be helpful. Obviously that only addresses a society of individuals of all about the same capability level — we need to figure out how to make this work in a society with agents of widely varying capability levels. We do have the advantage that many of them are trained, rather than evolved, so don’t have a straight-up underlying biological imperative to be self-interested (in a selfish gene sense: it may be in my self-interest to actually be a fine upstanding citizen and a good friend: it certainly makes maintaining such a reputation easier).
Human morality basically evolved as reputation management: if I get a reputation for being a bad partner to cooperate with, other people will make things go badly for me, so continuing to be thought of as a fine upstanding citizen and good friend is of real value to me. So it is built on an enforcement mechanism (originally gossip, shunning, and occasional violence, to which we have more recently added law-enforcement and the press) which basically assumes that there are agents of roughly comparable capability actively monitoring my behavior for whether it adheres to my society’s standards or not. (Note that unlike formal law, morality has, and needs to have, fuzzy edges and discretion, to avoid being loopholed — in law enforcement this comes down to selective enforcement (police and prosecutor discretion) and then judge/jury discretion.)
Morality is not something that arises naturally from a single agent — it’s the evolved outcome of a society of agents interacting and competing/cooperating/monitoring each other, in a way that turned out to have a fairly prosocial stable equilibrium. Us successfully enforcing morality on agents a lot smarter than us seems impractical. Us using honest, upstanding, moral ASIs to do monitoring and law enforcement on other more concerning ASIs of comparable capability seems a lot more practical. Debate is a rather simplistic version of this. This obviously raises a “who watches the watchmen?” problem. We need to devise a system of circular/mutual enforcement: different ASIs that monitor each other and keep each other honest/aligned, a system that has a stable state which is both aligned with our interests, and has some dynamic reason to stay aligned with our interests (such as that we have enough influence within the system to provide feedback against drift). Understanding how and why human societies (usually) have prosocial equilibria (as is studied in evolutionary psychology/moral anthropology/sociology of morality) seems likely to be helpful. Obviously that only addresses a society of individuals of all about the same capability level — we need to figure out how to make this work in a society with agents of widely varying capability levels. We do have the advantage that many of them are trained, rather than evolved, so don’t have a straight-up underlying biological imperative to be self-interested (in a selfish gene sense: it may be in my self-interest to actually be a fine upstanding citizen and a good friend: it certainly makes maintaining such a reputation easier).