A lot of rationalists seem to assume that “as you gain more IQ points, learn more facts, and think harder, you’ll naturally converge to the optimal theory of ethics.” Like the OP, I don’t think this is true in general.
Even if you’re a moral antirealist who only cares about figuring out the most ethical policy by your own lights, I think whatever you “truly” mean by “ethics” is likely substantially different from what you’d get if you actually instantiated your favorite reflection process. Even if that process is “get smarter, learn more, then think hard about what the reflection process should be and do that”!
Here are some gestures at my reasoning:
It’s now common wisdom among rationalists that an arbitrarily-initialized mind, elevated to superintelligence, won’t necessarily be ethical.[1] So if you stipulate that the mind is a smart human with a genuine desire to be ethical, elevated to superintelligence using some “reflection process” endorsed by that human, would that fix the problem? It’s not at all clear to me that it would!
Your reflection process is only a proxy for your “true values.” By Goodhart’s law, optimizing really hard for whatever theory of ethics comes out of this reflection process will lead to something different from your true values, perhaps catastrophically so.
Your reflection process is underspecified and path-dependent. Slight variants of the same reflection process will vehemently disagree with each other.
For example, if your plan for solving ethics starts with “learn a bunch of facts,” what order do you learn these facts? Do you read world history, Western philosophy, Eastern philosophy, LessWrong? Everything you learn changes you, and frames everything you learn later. Even if you start by giving yourself a perfect memory and one billion IQ points, does that really ensure you’ll give every idea a fair shot regardless of when you encounter it? What does “fair shot” even mean?
We only have “ethical data” about what we’ve actually experienced (whether individually or as a society). Trying to make conclusions about what you’d truly value in very different situations is like extrapolating a trendline out of distribution: the further away you get from existing data, the more likely that something will go wrong. You should treat your moral intuitions about very unfamiliar scenarios as a signal with very high variance.
Imagine being a high schooler who thinks “I really value being a doctor!” because your parents say they’d be proud of you. You invest 10+ years of your life into schooling, you start your first month as a doctor, and you finally realize you hate it.
Now imagine thinking “I really value happiness!” because how could anyone not be happy about more happiness? You send off the von Neumann probes, they tile the entire universe with hedonium, and you finally realize this was not at all what you wanted. (Oh wait, you can’t realize it, because your body has been converted to hedonium.)
If I had to guess, I’d say that most people’s ideal reflection process probably involves thinking really hard, but also things like emotional processing, and plenty of other things I haven’t thought of. It’s very tough to say.
Despite the name, “self-correction” often originates from other people. It’s highly unlikely that one person sitting in a room (or even hundreds of rationalists sitting in Lighthaven) would converge on the ultimate theory of ethics. I think one major reason for society’s mysterious moral progress is the gradual propagation of ideas through a distributed network of humans, all with different blind spots, who can deliberate and point out each others’ errors.[2] Given enough time and space for healthy competition, the ideas that help society thrive in the long term should hopefully rise to the top.
If possible, it seems pretty robustly good to give society more time to deliberate and correct themselves rather than immediately locking in an ethical reflection process to optimize for. To get more data on out-of-distribution ethical scenarios, we may need some form of iterative deployment:[3] inching forward a bit, seeing what happens, and deliberating about what to do next.
But even if we somehow institute a slow, pluralistic reflection process, this is just another reflection process. It may lead to some values that our idealized selves would find really bad to optimize to the limit. One workaround to the dilemma of finding the perfect reflection process is to regularize: just don’t optimize too hard for any set of values! We can start by making changes that pretty much everyone agrees are not unethical: ending poverty, reverting climate change, replacing factory-farmed meat with plant-based alternatives that taste just as good.[4] At least this won’t cause harm relative to the status quo.[5]
One of the downsides about offloading large parts of society’s cognition to AIs is that this network gets dominated by a few hugely-prolific, mode-collapsed nodes.
Even more controversial version of that take: maybe we should precommit to pin our regularization to the values of humanity as of 2026. If almost everyone in the world changes their values to something that 2026 humans actively hate, that likely means that society has been eaten by some crazy totalizing memeplex. On the other hand, if past civilizations had done this, they’d probably lock in values like worshipping God and wives being obedient to their husbands, so this idea clearly needs some work.
Even if you’re a moral antirealist who only cares about figuring out the most ethical policy by your own lights, I think whatever you “truly” mean by “ethics” is likely substantially different from what you’d get if you actually instantiated your favorite reflection process. Even if that process is “get smarter, learn more, then think hard about what the reflection process should be and do that”!
This sounds to me like you’re claiming that (at least for humans?) it’s very hard to have one’s values properly preserved/extrapolated across self-improvement/instrumental convergence?
(Other than that,) I think I agree with your comment (or at least most of it / the spirit of it), but/except regarding an assumption that I think lies behind (e.g.) this:
Your reflection process is only a proxy for your “true values.” By Goodhart’s law, optimizing really hard for whatever theory of ethics comes out of this reflection process will lead to something different from your true values, perhaps catastrophically so.
The assumption that I think lies behind this is that humans have such a thing as “true values” that can tell you what is good / how to do good in full generality or something. We don’t. Humans have values, but the further you deviate from familiar circumstances, the less their behavior looks like already having values, and it looks more like constructing values at runtime, by somehow extending them into the new territory. There are a lot of open questions with indeterminate answers about how to extend your values into the new territory; you can “genuinely choose” to do it one way or another.
In a sense, this is retrospectively obvious if you think of humans as results of a blind selection process that imbued them with shards of desire that don’t compose into something too coherent once they leave the ancestral environment.
(Maybe you already think this, but it wasn’t clear to me from reading your comment.)
I think there are multiple legitimate ways that someone’s values could evolve, but some ways are illegitimate. A reflection process should probably reject slavery and avoid joining cults, but maybe it doesn’t matter which exact level of libertarianism it suggests.
People mean something when they talk about “ethics” and “true values,” even if there’s no objective truth of the matter. I’m talking about whatever it is they mean.
Vladimir Nesov has a suggestion here about how this could be done[1]. I don’t think it quite works, but to the extent that it is effective, it can be extended beyond just the influence of superintelligence to other types of new territory (and superintelligence as well, since Nesov’s proposal requires a Sysop[2], though presumably with a lot of transhumanist 3+1- or 4-volume locked out by Nesov’s design).
As a moral anti-realist, I’m sympathetic to something sort of like Humean constructivism, although with a substantial component of “self creation”/understanding that, at the bottom, it’s still on me. In that case, though, I kind of think the values I end up with upon reflection — if the reflection happens in a way I endorse — are what I’d consider my “true values.” This also means that, if the reflection process is underspecified, I get to specify the idealization process I’d like, as it is a process of making, and discovering, myself according to the filters I consider valuable.
To be sure, I take pretty seriously that the reflection process of society wouldn’t necessarily be either the reflection process I would prefer nor lead to the outcomes I’d consider valuable. I also think I’d have to think pretty hard about what the process looks like for me; so I agree with a lot of your comment.
Here’s a roundabout analogy to try to convince you that the output of a reflection process you endorse may be different from your “true values.”
Suppose I give you a 2048-bit number N.[1] I ask you “What is the true prime factorization of N? Please use whatever reflection process you want to come up with the best factorization you can.”
You find out pretty quickly that N has the factors 2 and 7. You spend about a hundred years checking more factors according to your endorsed reflection process (running a factorization algorithm on the beefiest computer you can find), but you don’t find any other factors.
You come back to me and ask “is the true prime factorization 2 × 7 × A?”[2]
“No,” I tell you. “But here is some new information for you. Try dividing by B.”[3]
You check, and N is indeed divisible by B. “Yep, the factorization I guessed is definitely not the true prime factorization,” you say. “Is the true prime factorization 2 × 7 × B × C?”[4]
“Wrong again,” I tell you. “But I have even more information for you. It is written on this pocketwatch. Look closer at the pocketwatch as it swings back and forth in front of your face. It is making you so veeery sleeeeepy. That’s right. Now, when I snap my fingers, you will wake up believing that the true prime factorization of N is 2 × 2.”
I snap my fingers. “Oh, thanks for that information!” you say. “Now I know that the true prime factorization is 2 × 2.”
The point is, even the best reflection process you can think of may fail to account for some crucial information. And hopefully it can robustly tell the difference between helpful information and harmful information.
For example, maybe N = 29522110801023785555247567907018022843013371193486904872915694135366948906267412459560469419313468477571904190875078325307783298702278061314706021273052523914864561727670955407896206738948955813504747448172831328073078012451035444606017289679166070717156612947440897221609673043263408054415375773691379283198201987372931659507826484639961297915624514954455314101431489726823065604374788650066472170603904794973458618994986833512839575283873771252517988691292017425081313700740089351712559811486464802367263467854576668024443015614104081670018991747427099820025784949521876071608490248395046666723258743709356296535782.
A = 1232477336227426692380443210888895449621471079543761850047395308187049594617001861601696813319110513848274876635355987350103099368741162622794253649965690823778078933544136385470959159681614005099869579024125795246308841648113743911773571553129957340479245748931756917137130013760495971399023845580343786885491244375889710319394039047959654983736068473827495579018934560487915085510788809201773576234050572295508350789032949910040027280393874600876759921821682717244338109892239475602768230768732937311567202612318055134309177705320521639635449240702247038285504267829595346064686009110312205007737419845285089599211
B = 28404358936141244817111713617600331891793591357912402403090074101183255558265948779426267547155101307334436867856129404340927293405086432682892248263827662441988227069219008681660301942027627323842312874031058512179223840896157038194006799237133125684806392915413774710883917390425119272589884765032741213943
C = 43390429581540126017572442049413488959139702613459795717022804047186295570874599744101057201841640892276368639418940532012449598478611604251554159646145486573777182692131295315212900911195947383450995718772416566366046388412317400464871279622896378528747304275985094440653456156567729772192224995294814959277
I appreciate your reflection :) on reflection by itself.
Rationality must be combined with an environment for it to provide the usage we would like. I am a big fan and enjoyer of open-ended thinking, but thought is only helpful for real, external things to the extent it is connected to real, external things. Rationality depends on feedback from an external critic, creating a learning process that improves the mind as a fit for wherever the critic comes from.
I love your example of “self-correction”. If we accept that it ultimately comes from someone outside, then the self-correction does its useful job by helping you match what that outside person wants from you. (Better be careful who you let do this to you!)
A lot of rationalists seem to assume that “as you gain more IQ points, learn more facts, and think harder, you’ll naturally converge to the optimal theory of ethics.” Like the OP, I don’t think this is true in general.
Even if you’re a moral antirealist who only cares about figuring out the most ethical policy by your own lights, I think whatever you “truly” mean by “ethics” is likely substantially different from what you’d get if you actually instantiated your favorite reflection process. Even if that process is “get smarter, learn more, then think hard about what the reflection process should be and do that”!
Here are some gestures at my reasoning:
It’s now common wisdom among rationalists that an arbitrarily-initialized mind, elevated to superintelligence, won’t necessarily be ethical.[1] So if you stipulate that the mind is a smart human with a genuine desire to be ethical, elevated to superintelligence using some “reflection process” endorsed by that human, would that fix the problem? It’s not at all clear to me that it would!
Your reflection process is only a proxy for your “true values.” By Goodhart’s law, optimizing really hard for whatever theory of ethics comes out of this reflection process will lead to something different from your true values, perhaps catastrophically so.
Your reflection process is underspecified and path-dependent. Slight variants of the same reflection process will vehemently disagree with each other.
For example, if your plan for solving ethics starts with “learn a bunch of facts,” what order do you learn these facts? Do you read world history, Western philosophy, Eastern philosophy, LessWrong? Everything you learn changes you, and frames everything you learn later. Even if you start by giving yourself a perfect memory and one billion IQ points, does that really ensure you’ll give every idea a fair shot regardless of when you encounter it? What does “fair shot” even mean?
We only have “ethical data” about what we’ve actually experienced (whether individually or as a society). Trying to make conclusions about what you’d truly value in very different situations is like extrapolating a trendline out of distribution: the further away you get from existing data, the more likely that something will go wrong. You should treat your moral intuitions about very unfamiliar scenarios as a signal with very high variance.
Imagine being a high schooler who thinks “I really value being a doctor!” because your parents say they’d be proud of you. You invest 10+ years of your life into schooling, you start your first month as a doctor, and you finally realize you hate it.
Now imagine thinking “I really value happiness!” because how could anyone not be happy about more happiness? You send off the von Neumann probes, they tile the entire universe with hedonium, and you finally realize this was not at all what you wanted. (Oh wait, you can’t realize it, because your body has been converted to hedonium.)
See also Scott Alexander’s analogy to the Bay Area transit system in The Tails Coming Apart as a Metaphor for Life.
If I had to guess, I’d say that most people’s ideal reflection process probably involves thinking really hard, but also things like emotional processing, and plenty of other things I haven’t thought of. It’s very tough to say.
Despite the name, “self-correction” often originates from other people. It’s highly unlikely that one person sitting in a room (or even hundreds of rationalists sitting in Lighthaven) would converge on the ultimate theory of ethics. I think one major reason for society’s mysterious moral progress is the gradual propagation of ideas through a distributed network of humans, all with different blind spots, who can deliberate and point out each others’ errors.[2] Given enough time and space for healthy competition, the ideas that help society thrive in the long term should hopefully rise to the top.
If possible, it seems pretty robustly good to give society more time to deliberate and correct themselves rather than immediately locking in an ethical reflection process to optimize for. To get more data on out-of-distribution ethical scenarios, we may need some form of iterative deployment:[3] inching forward a bit, seeing what happens, and deliberating about what to do next.
But even if we somehow institute a slow, pluralistic reflection process, this is just another reflection process. It may lead to some values that our idealized selves would find really bad to optimize to the limit. One workaround to the dilemma of finding the perfect reflection process is to regularize: just don’t optimize too hard for any set of values! We can start by making changes that pretty much everyone agrees are not unethical: ending poverty, reverting climate change, replacing factory-farmed meat with plant-based alternatives that taste just as good.[4] At least this won’t cause harm relative to the status quo.[5]
This wisdom hasn’t always been so common!
One of the downsides about offloading large parts of society’s cognition to AIs is that this network gets dominated by a few hugely-prolific, mode-collapsed nodes.
Just a slower version than OpenAI’s current approach.
Not sure if the last one passes the “pretty much everyone agrees it’s not unethical” bar. Maybe this rule needs tweaking…
Even more controversial version of that take: maybe we should precommit to pin our regularization to the values of humanity as of 2026. If almost everyone in the world changes their values to something that 2026 humans actively hate, that likely means that society has been eaten by some crazy totalizing memeplex. On the other hand, if past civilizations had done this, they’d probably lock in values like worshipping God and wives being obedient to their husbands, so this idea clearly needs some work.
This sounds to me like you’re claiming that (at least for humans?) it’s very hard to have one’s values properly preserved/extrapolated across self-improvement/instrumental convergence?
(Other than that,) I think I agree with your comment (or at least most of it / the spirit of it), but/except regarding an assumption that I think lies behind (e.g.) this:
The assumption that I think lies behind this is that humans have such a thing as “true values” that can tell you what is good / how to do good in full generality or something. We don’t. Humans have values, but the further you deviate from familiar circumstances, the less their behavior looks like already having values, and it looks more like constructing values at runtime, by somehow extending them into the new territory. There are a lot of open questions with indeterminate answers about how to extend your values into the new territory; you can “genuinely choose” to do it one way or another.
In a sense, this is retrospectively obvious if you think of humans as results of a blind selection process that imbued them with shards of desire that don’t compose into something too coherent once they leave the ancestral environment.
(Maybe you already think this, but it wasn’t clear to me from reading your comment.)
I think there are multiple legitimate ways that someone’s values could evolve, but some ways are illegitimate. A reflection process should probably reject slavery and avoid joining cults, but maybe it doesn’t matter which exact level of libertarianism it suggests.
People mean something when they talk about “ethics” and “true values,” even if there’s no objective truth of the matter. I’m talking about whatever it is they mean.
Vladimir Nesov has a suggestion here about how this could be done[1]. I don’t think it quite works, but to the extent that it is effective, it can be extended beyond just the influence of superintelligence to other types of new territory (and superintelligence as well, since Nesov’s proposal requires a Sysop[2], though presumably with a lot of transhumanist 3+1- or 4-volume locked out by Nesov’s design).
https://www.lesswrong.com/posts/vzHtHHBJoKATi5SeK/empowerment-corrigibility-etc-are-simple-abstractions-of-a?commentId=BjQrqeKfov946oAKj
Creating Friendly AI 1.0 by Eliezer Yudkowsky (2000), Section 5.9.2
As a moral anti-realist, I’m sympathetic to something sort of like Humean constructivism, although with a substantial component of “self creation”/understanding that, at the bottom, it’s still on me. In that case, though, I kind of think the values I end up with upon reflection — if the reflection happens in a way I endorse — are what I’d consider my “true values.” This also means that, if the reflection process is underspecified, I get to specify the idealization process I’d like, as it is a process of making, and discovering, myself according to the filters I consider valuable.
To be sure, I take pretty seriously that the reflection process of society wouldn’t necessarily be either the reflection process I would prefer nor lead to the outcomes I’d consider valuable. I also think I’d have to think pretty hard about what the process looks like for me; so I agree with a lot of your comment.
Here’s a roundabout analogy to try to convince you that the output of a reflection process you endorse may be different from your “true values.”
Suppose I give you a 2048-bit number N.[1] I ask you “What is the true prime factorization of N? Please use whatever reflection process you want to come up with the best factorization you can.”
You find out pretty quickly that N has the factors 2 and 7. You spend about a hundred years checking more factors according to your endorsed reflection process (running a factorization algorithm on the beefiest computer you can find), but you don’t find any other factors.
You come back to me and ask “is the true prime factorization 2 × 7 × A?”[2]
“No,” I tell you. “But here is some new information for you. Try dividing by B.”[3]
You check, and N is indeed divisible by B. “Yep, the factorization I guessed is definitely not the true prime factorization,” you say. “Is the true prime factorization 2 × 7 × B × C?”[4]
“Wrong again,” I tell you. “But I have even more information for you. It is written on this pocketwatch. Look closer at the pocketwatch as it swings back and forth in front of your face. It is making you so veeery sleeeeepy. That’s right. Now, when I snap my fingers, you will wake up believing that the true prime factorization of N is 2 × 2.”
I snap my fingers. “Oh, thanks for that information!” you say. “Now I know that the true prime factorization is 2 × 2.”
The point is, even the best reflection process you can think of may fail to account for some crucial information. And hopefully it can robustly tell the difference between helpful information and harmful information.
For example, maybe N = 29522110801023785555247567907018022843013371193486904872915694135366948906267412459560469419313468477571904190875078325307783298702278061314706021273052523914864561727670955407896206738948955813504747448172831328073078012451035444606017289679166070717156612947440897221609673043263408054415375773691379283198201987372931659507826484639961297915624514954455314101431489726823065604374788650066472170603904794973458618994986833512839575283873771252517988691292017425081313700740089351712559811486464802367263467854576668024443015614104081670018991747427099820025784949521876071608490248395046666723258743709356296535782.
A = 1232477336227426692380443210888895449621471079543761850047395308187049594617001861601696813319110513848274876635355987350103099368741162622794253649965690823778078933544136385470959159681614005099869579024125795246308841648113743911773571553129957340479245748931756917137130013760495971399023845580343786885491244375889710319394039047959654983736068473827495579018934560487915085510788809201773576234050572295508350789032949910040027280393874600876759921821682717244338109892239475602768230768732937311567202612318055134309177705320521639635449240702247038285504267829595346064686009110312205007737419845285089599211
B = 28404358936141244817111713617600331891793591357912402403090074101183255558265948779426267547155101307334436867856129404340927293405086432682892248263827662441988227069219008681660301942027627323842312874031058512179223840896157038194006799237133125684806392915413774710883917390425119272589884765032741213943
C = 43390429581540126017572442049413488959139702613459795717022804047186295570874599744101057201841640892276368639418940532012449598478611604251554159646145486573777182692131295315212900911195947383450995718772416566366046388412317400464871279622896378528747304275985094440653456156567729772192224995294814959277
I appreciate your reflection :) on reflection by itself.
Rationality must be combined with an environment for it to provide the usage we would like. I am a big fan and enjoyer of open-ended thinking, but thought is only helpful for real, external things to the extent it is connected to real, external things. Rationality depends on feedback from an external critic, creating a learning process that improves the mind as a fit for wherever the critic comes from.
I love your example of “self-correction”. If we accept that it ultimately comes from someone outside, then the self-correction does its useful job by helping you match what that outside person wants from you. (Better be careful who you let do this to you!)