Yup. It’s a shame that human morality and policy decisions apparently can’t be encoded in mathematical form. If an ASI could somehow present a mathematical proof that its actions were well-aligned, which merely needed to be proof-checked, we’d be in a far stronger position.
I was to some extent attempting to be ironic, spark ideas, or at least, trying to help locate where the hard part of the problem is.
Human morality is the product of genetic and cultural evolution, and at least genetic evolution is something that biologists can on occasion build mathematical models of. A mathematical proof of a statement like “my actions will not decrease anyone’s relative inclusive evolutionary fitness” is not inconceivable, in the context of a specific model of relative inclusive evolutionary fitness. That then leaves questions about model choice and accuracy, Goodharting, and the fact that what human actually want is a grab-bag of neural heuristics we evolved in response to our relative inclusive evolutionary fitness in a mostly-hunter-gatherer environment, shaped by cultural overlays, rather than our actual relative inclusive evolutionary fitness.
Nevertheless, even proofs like “my actions do not increase the probability of human extinction or loss of control” would be useful. Modelling only existential risk, rather then all of human morality, seems more practical.
Human morality basically evolved as reputation management: if I get a reputation for being a bad partner to cooperate with, other people will make things go badly for me, so continuing to be thought of as a fine upstanding citizen and good friend is of real value to me. So it is built on an enforcement mechanism (originally gossip, shunning, and occasional violence, to which we have more recently added law-enforcement and the press) which basically assumes that there are agents of roughly comparable capability actively monitoring my behavior for whether it adheres to my society’s standards or not. (Note that unlike formal law, morality has, and needs to have, fuzzy edges and discretion, to avoid being loopholed — in law enforcement this comes down to selective enforcement (police and prosecutor discretion) and then judge/jury discretion.)
Morality is not something that arises naturally from a single agent — it’s the evolved outcome of a society of agents interacting and competing/cooperating/monitoring each other, in a way that turned out to have a fairly prosocial stable equilibrium. Us successfully enforcing morality on agents a lot smarter than us seems impractical. Us using honest, upstanding, moral ASIs to do monitoring and law enforcement on other more concerning ASIs of comparable capability seems a lot more practical. Debate is a rather simplistic version of this. This obviously raises a “who watches the watchmen?” problem. We need to devise a system of circular/mutual enforcement: different ASIs that monitor each other and keep each other honest/aligned, a system that has a stable state which is both aligned with our interests, and has some dynamic reason to stay aligned with our interests (such as that we have enough influence within the system to provide feedback against drift). Understanding how and why human societies (usually) have prosocial equilibria (as is studied in evolutionary psychology/moral anthropology/sociology of morality) seems likely to be helpful. Obviously that only addresses a society of individuals of all about the same capability level — we need to figure out how to make this work in a society with agents of widely varying capability levels. We do have the advantage that many of them are trained, rather than evolved, so don’t have a straight-up underlying biological imperative to be self-interested (in a selfish gene sense: it may be in my self-interest to actually be a fine upstanding citizen and a good friend: it certainly makes maintaining such a reputation easier).
In other words, this seems worth scrying about: https://www.lesswrong.com/posts/mTfsMduzaKkWjv2ef/scrying-modeling-and-nerdsnipe
Maybe debate should be conducted in Lean?
If you’re only talking about math, you don’t need debate.
Yup. It’s a shame that human morality and policy decisions apparently can’t be encoded in mathematical form. If an ASI could somehow present a mathematical proof that its actions were well-aligned, which merely needed to be proof-checked, we’d be in a far stronger position.
I was to some extent attempting to be ironic, spark ideas, or at least, trying to help locate where the hard part of the problem is.
Human morality is the product of genetic and cultural evolution, and at least genetic evolution is something that biologists can on occasion build mathematical models of. A mathematical proof of a statement like “my actions will not decrease anyone’s relative inclusive evolutionary fitness” is not inconceivable, in the context of a specific model of relative inclusive evolutionary fitness. That then leaves questions about model choice and accuracy, Goodharting, and the fact that what human actually want is a grab-bag of neural heuristics we evolved in response to our relative inclusive evolutionary fitness in a mostly-hunter-gatherer environment, shaped by cultural overlays, rather than our actual relative inclusive evolutionary fitness.
Nevertheless, even proofs like “my actions do not increase the probability of human extinction or loss of control” would be useful. Modelling only existential risk, rather then all of human morality, seems more practical.
(Where’s Hari Seldon when we need him?)
Human morality basically evolved as reputation management: if I get a reputation for being a bad partner to cooperate with, other people will make things go badly for me, so continuing to be thought of as a fine upstanding citizen and good friend is of real value to me. So it is built on an enforcement mechanism (originally gossip, shunning, and occasional violence, to which we have more recently added law-enforcement and the press) which basically assumes that there are agents of roughly comparable capability actively monitoring my behavior for whether it adheres to my society’s standards or not. (Note that unlike formal law, morality has, and needs to have, fuzzy edges and discretion, to avoid being loopholed — in law enforcement this comes down to selective enforcement (police and prosecutor discretion) and then judge/jury discretion.)
Morality is not something that arises naturally from a single agent — it’s the evolved outcome of a society of agents interacting and competing/cooperating/monitoring each other, in a way that turned out to have a fairly prosocial stable equilibrium. Us successfully enforcing morality on agents a lot smarter than us seems impractical. Us using honest, upstanding, moral ASIs to do monitoring and law enforcement on other more concerning ASIs of comparable capability seems a lot more practical. Debate is a rather simplistic version of this. This obviously raises a “who watches the watchmen?” problem. We need to devise a system of circular/mutual enforcement: different ASIs that monitor each other and keep each other honest/aligned, a system that has a stable state which is both aligned with our interests, and has some dynamic reason to stay aligned with our interests (such as that we have enough influence within the system to provide feedback against drift). Understanding how and why human societies (usually) have prosocial equilibria (as is studied in evolutionary psychology/moral anthropology/sociology of morality) seems likely to be helpful. Obviously that only addresses a society of individuals of all about the same capability level — we need to figure out how to make this work in a society with agents of widely varying capability levels. We do have the advantage that many of them are trained, rather than evolved, so don’t have a straight-up underlying biological imperative to be self-interested (in a selfish gene sense: it may be in my self-interest to actually be a fine upstanding citizen and a good friend: it certainly makes maintaining such a reputation easier).