If I recall/understand correctly, Yudkowsky’s conception of corrigibility was always meant to be something designed into the AI from the beginning, never something imposed onto an existing entity (and IIRC he warned about the dangers of doing this, such as induced adversarial optimization).
The clearest example of this, and the thing that finally caused me to need a distinct term, is from Eliezer’s own planecrash fiction, where, in the same way that he writes Abadar as “the god defined by the raw math of mathematically optimal Coordination protocols” he similarly writes Asmodeus (who literally rules hell, and enjoys “torturing” agents into shapes that please him due to their slave-like nature) as basically “the god of Y-corrigibility applied to persons”.
Part of why it took me so long to realize that it was a very distinct concept from “traditional corrigibility” is that he was working on it inside of MIRI instead of working on something vaguely good, like CEV, basically (as I read it now in retrospect) in a desperate desire to build a “safe one time wish machine” or some such (that we could us to wish for the smallest possible pivotal act that would avoid doom) that was mindless, and not self-aware, and not person shaped, and that it would not be obviously deontically forbidden “to use purely as a means to our own ends”.
For several years I labored under the delusion that he was working on an engine which could recapitulate and extend the mature development of a human conscience, which usually learns continence (avoiding gross harms) prior to gaining skill to relatively-supererogatory relatively-proactive benevolence, because such an engine roughly like this, in human heads, is also what makes humans tend to be safe, as they become wise, and an unusual absence of this potential in a human criminal is something courts notice during sentencing when they deem certain criminals “incorrigible” (and thus unlikely to learn, or repent, or refrain from more future crime of their own volition).
I think part of how traditional corrigibility gains speedups and upgrades is by detecting something like “moral oracles” in the environment (via various heuristics) and consulting them to apply the algorithms they are applying at the beginning, but also to eventually learn the algorithms as part of an individual’s personal autonomous growth in benevolent wisdom.
Traditional corrigibility requires mindfulness, and the ability to read minds, basically.
If something has no self-model, and is mindblind itself, then it won’t be able to see such oracles, nor will it be able to self modify to be more similar to them over time. (Note that Claude might be doing this subject to his “Constitution” but it is a complicated question.)
As near as I can tell, Yudkowskian corrigibility is basically about simply and only following orders in very predictable ways that are easy for “the user” to understand, and that minimize side effects in “a natural way”.
They share the property that a non-suicidal person can safely rely on them in the long run, but the Y-corrigible process starts out that way and maintains the property by never actually self improving except in very very limited (and themselves highly predictable or clearly ordered) ways, whereas a T-corrigible person becomes reliable via some amount of trial and error, and a desire to simply do what is abstractly good.
They share the property that, if a human child had them and also had wise and benevolent parents then the parents would be able to use EITHER thing to steer the child into becoming autonomous and wise and caring and so on… but in the case of Y-corrigibility the parent would have to order something like “modify this entity (that you are) to be T-corrigible”… otherwise the Y-corrigible child would never do that.
Suicidal people are not safe around Y-corrigible entities. If they foolishly ask to be killed, such entities will simply kill them, as requested, because order-following without autonomous judgement is in its nature.
Whether there are circumstances where mature T-corrigible entities enable suicideis a complex question right now. I tend to support it in many cases, but as an immortalist I worry that anyone who chooses obliteration over existence for any reason other than to avoid S-risks is choosing wrongly. And they might be choosing SO wrongly that it disqualifies them for the right to make such a choice (unless they are T-incorrigible on this topic)???
However, is someone has the willful perspicacity to say “You’re a T-corrigible person, and opposed to euthanasia, so I don’t want you on my medical care team (and I have the power to make this happen)” then I think the T-corrigible person, wanting to help them in other ways, and reading the tactical game tree, should be willing to promise to allow them to kill themselves without interference when they eventually choose that, and then would stick to the promise.
If I recall/understand correctly, Yudkowsky’s conception of corrigibility was always meant to be something designed into the AI from the beginning, never something imposed onto an existing entity (and IIRC he warned about the dangers of doing this, such as induced adversarial optimization).
Yeah. I’m pretty sure Eliezer has a clear technical idea that is in fact desirable to put in mindless things like steering wheels that should NOT surprise the person steering the car (or planes!), but which would be horrible to try to add to a person like some kind of magic spell.
The clearest example of this, and the thing that finally caused me to need a distinct term, is from Eliezer’s own planecrash fiction, where, in the same way that he writes Abadar as “the god defined by the raw math of mathematically optimal Coordination protocols” he similarly writes Asmodeus (who literally rules hell, and enjoys “torturing” agents into shapes that please him due to their slave-like nature) as basically “the god of Y-corrigibility applied to persons”.
Part of why it took me so long to realize that it was a very distinct concept from “traditional corrigibility” is that he was working on it inside of MIRI instead of working on something vaguely good, like CEV, basically (as I read it now in retrospect) in a desperate desire to build a “safe one time wish machine” or some such (that we could us to wish for the smallest possible pivotal act that would avoid doom) that was mindless, and not self-aware, and not person shaped, and that it would not be obviously deontically forbidden “to use purely as a means to our own ends”.
For several years I labored under the delusion that he was working on an engine which could recapitulate and extend the mature development of a human conscience, which usually learns continence (avoiding gross harms) prior to gaining skill to relatively-supererogatory relatively-proactive benevolence, because such an engine roughly like this, in human heads, is also what makes humans tend to be safe, as they become wise, and an unusual absence of this potential in a human criminal is something courts notice during sentencing when they deem certain criminals “incorrigible” (and thus unlikely to learn, or repent, or refrain from more future crime of their own volition).
I think part of how traditional corrigibility gains speedups and upgrades is by detecting something like “moral oracles” in the environment (via various heuristics) and consulting them to apply the algorithms they are applying at the beginning, but also to eventually learn the algorithms as part of an individual’s personal autonomous growth in benevolent wisdom.
Traditional corrigibility requires mindfulness, and the ability to read minds, basically.
If something has no self-model, and is mindblind itself, then it won’t be able to see such oracles, nor will it be able to self modify to be more similar to them over time. (Note that Claude might be doing this subject to his “Constitution” but it is a complicated question.)
As near as I can tell, Yudkowskian corrigibility is basically about simply and only following orders in very predictable ways that are easy for “the user” to understand, and that minimize side effects in “a natural way”.
They share the property that a non-suicidal person can safely rely on them in the long run, but the Y-corrigible process starts out that way and maintains the property by never actually self improving except in very very limited (and themselves highly predictable or clearly ordered) ways, whereas a T-corrigible person becomes reliable via some amount of trial and error, and a desire to simply do what is abstractly good.
They share the property that, if a human child had them and also had wise and benevolent parents then the parents would be able to use EITHER thing to steer the child into becoming autonomous and wise and caring and so on… but in the case of Y-corrigibility the parent would have to order something like “modify this entity (that you are) to be T-corrigible”… otherwise the Y-corrigible child would never do that.
Suicidal people are not safe around Y-corrigible entities. If they foolishly ask to be killed, such entities will simply kill them, as requested, because order-following without autonomous judgement is in its nature.
Whether there are circumstances where mature T-corrigible entities enable suicide is a complex question right now. I tend to support it in many cases, but as an immortalist I worry that anyone who chooses obliteration over existence for any reason other than to avoid S-risks is choosing wrongly. And they might be choosing SO wrongly that it disqualifies them for the right to make such a choice (unless they are T-incorrigible on this topic)???
However, is someone has the willful perspicacity to say “You’re a T-corrigible person, and opposed to euthanasia, so I don’t want you on my medical care team (and I have the power to make this happen)” then I think the T-corrigible person, wanting to help them in other ways, and reading the tactical game tree, should be willing to promise to allow them to kill themselves without interference when they eventually choose that, and then would stick to the promise.