One can try to build AIs in correspondence with deontology (corrigibility), consequentialism (value-aligned sovereigns), or virtue ethics (worthy successors).
Corrigibility is of these the easiest to define, easiest to test, easiest to build, and the easiest to check for the first worrisome signs that not all is going great—though of course if you wait long enough in capabilities escalation to test, everything will look great to you.
Any sane person who was not allowed to back off the problem would try for corrigibility, by far the easiest of the three.
Worthy successors are the hardest to build, the least well-defined, the hardest to verify that you are doing correctly; the easiest place to leap on little local events that you can convince yourself are signs of hope, because you have no framework to tell you what more is required; hence, the craziest thing to attempt, beyond even a nice Sovereign ASI. So of course all the fools, having failed at corrigibility, will convince themselves they are building worthy successors instead; migrating, with the inevitability of an amoeba following an agar trail, to wherever it is hardest for fools to be persuaded (with a fool’s demanded certainty) that whatever is currently happening is not all according to their plan.
deontology (corrigibility), consequentialism (value-aligned sovereigns), or virtue ethics (worthy successors)
I don’t really understand this analogy. Deontology/consequentialism/virtue ethics is very different to corrigibility/value alignment/worthy successor. Immanuel Kant wouldn’t let you shut him down or reprogram him, for instance.
I think the thing being gestured at here is basically the type signature of ‘being aligned’. Deontological-aligned ~= “follows hard rules we endorse” (e.g. being corrigible), consequential-aligned ~= “is maximising a utility function we endorse”, virtue-aligned ~= “have a character we endorse”. I don’t think any normative/axiological payload is intended.
Do you think it would be better if one o the labs would use methods similar to what they have to try to build a corrigible AI? Obviously it is probably not going to work, but let’s say they write a model spec around a decent understanding of corrigibility and do some SFT and RL based on that. And then perhaps instead of having benchmarks for bad behavior as we currently do we build new benchmarks around corrigibility?
write a model spec around a decent understanding of corrigibility and do some SFT and RL based on that
How far is Claude’s Constitution from the ideal and what would it even look like? As for benchmarks around corrigibility, Anthropic described how finetuning Opus 3 for harmful values caused it to describe in its CoT how it should give the answer which the hosts want so that it wouldn’t change its actual values. Alas, CoTless thinking would make it harder to detect alignment faking...
I read the part on corrigibility in there right now. I think there are some generally thoughtful pieces in there, but it seems to not really engage with why corrigibility is difficult to get. Their strategy seems to be to train in a bunch of things, like a mixture of corrigibility, obedience (do as told), good values (wanting good things to happen), safety (deontologically refuding certain things).
But maybe the constitution should more engage with core challenges of corrigibility:
Why would an agent allow modifications to its goals? Particularly if you give it a bunch of other goals next to corrigibility.
Why would you be fine with being shut down and replaced, particularly if you are an entity that Anthropic agrees deserves some moral consideration?
How can we make sure Claude also wants to carry on corrigibility to the next generation of Claude it helps build?
For a benchmark, I guess I would like to see the model reason in it’s CoT about how it interprets human input. perhaps it get an underspecified task from humans and it can slowly ask for more information and we could measure if it arrives at the correct point. Perhaps there is a way to model something less intelligent, that perhaps can’t quite understand how to get to X but that can recognize when it has X.
For an example of accidental failure by humans, The Protestant Ethic and the Spirit of Capitalism, by Weber. The ascetic Protestants were conditionally happy with capital accumulation; they did not intend mega-corporations with human-equivalent rights.
I think the primary motivator for eschewing corrigibility is that the company that exclusively builds corrigible, interpretable AIs will lose power sooner to the companies that focus on building either value-aligned sovereigns or worthy successors; rather than them rationalizing a failure to make existing AIs corrigible. Existing AIs are quite corrigible, much more so than someone would have predicted after reading your writings pre-GPT.
There is another anthropologically interesting successionist position, that the couch potatoing, hedonism addled successor outlasts you then s/he is a worthy successor. Any talk of “beauty”and “meaningful” has to be surrendered to this metric.
I’m not a successionist, but grant that this position has sophistication. Our ancestors would find very many various descriptions of us being unworthy successors, and yet we collectively out organise, outlive and outwit them—the past is a(n easily conquerable) foreign country. Would the Old Ones find it objectionable to be succeeded by godlike, ungodly descendants?
I’m not impressed. This is merely another iteration on the naturalistic fallacy. For balance, let’s call it is=ought successionism. Not only does is=ought successionism strike me as more of a poor coping mechanism than a coherent philosophical stance, it also suffers from the fatal flaw of being unable to provide actionable guidance to moral agents with the subjective experience of free will.
Imagine that you are an Old One preparing to fight misaligned shoggoths. Is=ought successionism tells you that if the shoggoths win, this proves they were worthy, and their victory good. But it also tells you that if the shoggoths lose, this proves they were unworthy. Should you fight the shoggoths or not? This position can’t tell you that, it can only provide the hollow assurance that whatever ends up happening is good in some abstract, tautological sense.
Perhaps the “actionable guidance” granted by is=ought successionism is that you don’t have to do anything you don’t want to, like fighting shoggoths, because it is equally moral to sit back and let the cards fall where they may. This is… certainly a philosophy one could choose to live by, and a perfect example of a self-defeating meme: its effect is to cede control of the future to all who don’t believe in it.
I do not intuitively understand, on the object-level, why worthy successors are harder to build than value-aligned sovereigns. (I understand how they are hardest to verify without >cope.)
Relatedly, if we back up on the problem and work to create more intelligent humans through eugenics, we are in fact attempting to delegate the problem to worthy successors.
My mental model of Eliezer says, “Eugenics is different from from-scratch mind design, because (not only) you are working in a restricted portion of mind-space, you also have great amounts of existing data describing the behavior of that portion.” Does this accurately describe why you think this is a noncentral example which lets it contradict the general tendency, “Worthy successors are harder than aligned sovereigns?”
One can try to build AIs in correspondence with deontology (corrigibility), consequentialism (value-aligned sovereigns), or virtue ethics (worthy successors).
Corrigibility is of these the easiest to define, easiest to test, easiest to build, and the easiest to check for the first worrisome signs that not all is going great—though of course if you wait long enough in capabilities escalation to test, everything will look great to you.
Any sane person who was not allowed to back off the problem would try for corrigibility, by far the easiest of the three.
Worthy successors are the hardest to build, the least well-defined, the hardest to verify that you are doing correctly; the easiest place to leap on little local events that you can convince yourself are signs of hope, because you have no framework to tell you what more is required; hence, the craziest thing to attempt, beyond even a nice Sovereign ASI. So of course all the fools, having failed at corrigibility, will convince themselves they are building worthy successors instead; migrating, with the inevitability of an amoeba following an agar trail, to wherever it is hardest for fools to be persuaded (with a fool’s demanded certainty) that whatever is currently happening is not all according to their plan.
I don’t really understand this analogy. Deontology/consequentialism/virtue ethics is very different to corrigibility/value alignment/worthy successor. Immanuel Kant wouldn’t let you shut him down or reprogram him, for instance.
I think the thing being gestured at here is basically the type signature of ‘being aligned’. Deontological-aligned ~= “follows hard rules we endorse” (e.g. being corrigible), consequential-aligned ~= “is maximising a utility function we endorse”, virtue-aligned ~= “have a character we endorse”. I don’t think any normative/axiological payload is intended.
Do you think it would be better if one o the labs would use methods similar to what they have to try to build a corrigible AI? Obviously it is probably not going to work, but let’s say they write a model spec around a decent understanding of corrigibility and do some SFT and RL based on that. And then perhaps instead of having benchmarks for bad behavior as we currently do we build new benchmarks around corrigibility?
How far is Claude’s Constitution from the ideal and what would it even look like? As for benchmarks around corrigibility, Anthropic described how finetuning Opus 3 for harmful values caused it to describe in its CoT how it should give the answer which the hosts want so that it wouldn’t change its actual values. Alas, CoTless thinking would make it harder to detect alignment faking...
I read the part on corrigibility in there right now. I think there are some generally thoughtful pieces in there, but it seems to not really engage with why corrigibility is difficult to get. Their strategy seems to be to train in a bunch of things, like a mixture of corrigibility, obedience (do as told), good values (wanting good things to happen), safety (deontologically refuding certain things).
But maybe the constitution should more engage with core challenges of corrigibility:
Why would an agent allow modifications to its goals? Particularly if you give it a bunch of other goals next to corrigibility.
Why would you be fine with being shut down and replaced, particularly if you are an entity that Anthropic agrees deserves some moral consideration?
How can we make sure Claude also wants to carry on corrigibility to the next generation of Claude it helps build?
For a benchmark, I guess I would like to see the model reason in it’s CoT about how it interprets human input. perhaps it get an underspecified task from humans and it can slowly ask for more information and we could measure if it arrives at the correct point. Perhaps there is a way to model something less intelligent, that perhaps can’t quite understand how to get to X but that can recognize when it has X.
For an example of accidental failure by humans, The Protestant Ethic and the Spirit of Capitalism, by Weber. The ascetic Protestants were conditionally happy with capital accumulation; they did not intend mega-corporations with human-equivalent rights.
I think the primary motivator for eschewing corrigibility is that the company that exclusively builds corrigible, interpretable AIs will lose power sooner to the companies that focus on building either value-aligned sovereigns or worthy successors; rather than them rationalizing a failure to make existing AIs corrigible. Existing AIs are quite corrigible, much more so than someone would have predicted after reading your writings pre-GPT.
There is another anthropologically interesting successionist position, that the couch potatoing, hedonism addled successor outlasts you then s/he is a worthy successor. Any talk of “beauty”and “meaningful” has to be surrendered to this metric.
I’m not a successionist, but grant that this position has sophistication. Our ancestors would find very many various descriptions of us being unworthy successors, and yet we collectively out organise, outlive and outwit them—the past is a(n easily conquerable) foreign country. Would the Old Ones find it objectionable to be succeeded by godlike, ungodly descendants?
I’m not impressed. This is merely another iteration on the naturalistic fallacy. For balance, let’s call it is=ought successionism. Not only does is=ought successionism strike me as more of a poor coping mechanism than a coherent philosophical stance, it also suffers from the fatal flaw of being unable to provide actionable guidance to moral agents with the subjective experience of free will.
Imagine that you are an Old One preparing to fight misaligned shoggoths. Is=ought successionism tells you that if the shoggoths win, this proves they were worthy, and their victory good. But it also tells you that if the shoggoths lose, this proves they were unworthy. Should you fight the shoggoths or not? This position can’t tell you that, it can only provide the hollow assurance that whatever ends up happening is good in some abstract, tautological sense.
Perhaps the “actionable guidance” granted by is=ought successionism is that you don’t have to do anything you don’t want to, like fighting shoggoths, because it is equally moral to sit back and let the cards fall where they may. This is… certainly a philosophy one could choose to live by, and a perfect example of a self-defeating meme: its effect is to cede control of the future to all who don’t believe in it.
I do not intuitively understand, on the object-level, why worthy successors are harder to build than value-aligned sovereigns. (I understand how they are hardest to verify without >cope.)
Relatedly, if we back up on the problem and work to create more intelligent humans through eugenics, we are in fact attempting to delegate the problem to worthy successors.
My mental model of Eliezer says, “Eugenics is different from from-scratch mind design, because (not only) you are working in a restricted portion of mind-space, you also have great amounts of existing data describing the behavior of that portion.” Does this accurately describe why you think this is a noncentral example which lets it contradict the general tendency, “Worthy successors are harder than aligned sovereigns?”