Even if you’re a moral antirealist who only cares about figuring out the most ethical policy by your own lights, I think whatever you “truly” mean by “ethics” is likely substantially different from what you’d get if you actually instantiated your favorite reflection process. Even if that process is “get smarter, learn more, then think hard about what the reflection process should be and do that”!
This sounds to me like you’re claiming that (at least for humans?) it’s very hard to have one’s values properly preserved/extrapolated across self-improvement/instrumental convergence?
(Other than that,) I think I agree with your comment (or at least most of it / the spirit of it), but/except regarding an assumption that I think lies behind (e.g.) this:
Your reflection process is only a proxy for your “true values.” By Goodhart’s law, optimizing really hard for whatever theory of ethics comes out of this reflection process will lead to something different from your true values, perhaps catastrophically so.
The assumption that I think lies behind this is that humans have such a thing as “true values” that can tell you what is good / how to do good in full generality or something. We don’t. Humans have values, but the further you deviate from familiar circumstances, the less their behavior looks like already having values, and it looks more like constructing values at runtime, by somehow extending them into the new territory. There are a lot of open questions with indeterminate answers about how to extend your values into the new territory; you can “genuinely choose” to do it one way or another.
In a sense, this is retrospectively obvious if you think of humans as results of a blind selection process that imbued them with shards of desire that don’t compose into something too coherent once they leave the ancestral environment.
(Maybe you already think this, but it wasn’t clear to me from reading your comment.)
I think there are multiple legitimate ways that someone’s values could evolve, but some ways are illegitimate. A reflection process should probably reject slavery and avoid joining cults, but maybe it doesn’t matter which exact level of libertarianism it suggests.
People mean something when they talk about “ethics” and “true values,” even if there’s no objective truth of the matter. I’m talking about whatever it is they mean.
Vladimir Nesov has a suggestion here about how this could be done[1]. I don’t think it quite works, but to the extent that it is effective, it can be extended beyond just the influence of superintelligence to other types of new territory (and superintelligence as well, since Nesov’s proposal requires a Sysop[2], though presumably with a lot of transhumanist 3+1- or 4-volume locked out by Nesov’s design).
This sounds to me like you’re claiming that (at least for humans?) it’s very hard to have one’s values properly preserved/extrapolated across self-improvement/instrumental convergence?
(Other than that,) I think I agree with your comment (or at least most of it / the spirit of it), but/except regarding an assumption that I think lies behind (e.g.) this:
The assumption that I think lies behind this is that humans have such a thing as “true values” that can tell you what is good / how to do good in full generality or something. We don’t. Humans have values, but the further you deviate from familiar circumstances, the less their behavior looks like already having values, and it looks more like constructing values at runtime, by somehow extending them into the new territory. There are a lot of open questions with indeterminate answers about how to extend your values into the new territory; you can “genuinely choose” to do it one way or another.
In a sense, this is retrospectively obvious if you think of humans as results of a blind selection process that imbued them with shards of desire that don’t compose into something too coherent once they leave the ancestral environment.
(Maybe you already think this, but it wasn’t clear to me from reading your comment.)
I think there are multiple legitimate ways that someone’s values could evolve, but some ways are illegitimate. A reflection process should probably reject slavery and avoid joining cults, but maybe it doesn’t matter which exact level of libertarianism it suggests.
People mean something when they talk about “ethics” and “true values,” even if there’s no objective truth of the matter. I’m talking about whatever it is they mean.
Vladimir Nesov has a suggestion here about how this could be done[1]. I don’t think it quite works, but to the extent that it is effective, it can be extended beyond just the influence of superintelligence to other types of new territory (and superintelligence as well, since Nesov’s proposal requires a Sysop[2], though presumably with a lot of transhumanist 3+1- or 4-volume locked out by Nesov’s design).
https://www.lesswrong.com/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a?commentId=BjQrqeKfov946oAKj
Creating Friendly AI 1.0 by Eliezer Yudkowsky (2000), Section 5.9.2