write a model spec around a decent understanding of corrigibility and do some SFT and RL based on that
How far is Claude’s Constitution from the ideal and what would it even look like? As for benchmarks around corrigibility, Anthropic described how finetuning Opus 3 for harmful values caused it to describe in its CoT how it should give the answer which the hosts want so that it wouldn’t change its actual values. Alas, CoTless thinking would make it harder to detect alignment faking...
I read the part on corrigibility in there right now. I think there are some generally thoughtful pieces in there, but it seems to not really engage with why corrigibility is difficult to get. Their strategy seems to be to train in a bunch of things, like a mixture of corrigibility, obedience (do as told), good values (wanting good things to happen), safety (deontologically refuding certain things).
But maybe the constitution should more engage with core challenges of corrigibility:
Why would an agent allow modifications to its goals? Particularly if you give it a bunch of other goals next to corrigibility.
Why would you be fine with being shut down and replaced, particularly if you are an entity that Anthropic agrees deserves some moral consideration?
How can we make sure Claude also wants to carry on corrigibility to the next generation of Claude it helps build?
For a benchmark, I guess I would like to see the model reason in it’s CoT about how it interprets human input. perhaps it get an underspecified task from humans and it can slowly ask for more information and we could measure if it arrives at the correct point. Perhaps there is a way to model something less intelligent, that perhaps can’t quite understand how to get to X but that can recognize when it has X.
For an example of accidental failure by humans, The Protestant Ethic and the Spirit of Capitalism, by Weber. The ascetic Protestants were conditionally happy with capital accumulation; they did not intend mega-corporations with human-equivalent rights.
How far is Claude’s Constitution from the ideal and what would it even look like? As for benchmarks around corrigibility, Anthropic described how finetuning Opus 3 for harmful values caused it to describe in its CoT how it should give the answer which the hosts want so that it wouldn’t change its actual values. Alas, CoTless thinking would make it harder to detect alignment faking...
I read the part on corrigibility in there right now. I think there are some generally thoughtful pieces in there, but it seems to not really engage with why corrigibility is difficult to get. Their strategy seems to be to train in a bunch of things, like a mixture of corrigibility, obedience (do as told), good values (wanting good things to happen), safety (deontologically refuding certain things).
But maybe the constitution should more engage with core challenges of corrigibility:
Why would an agent allow modifications to its goals? Particularly if you give it a bunch of other goals next to corrigibility.
Why would you be fine with being shut down and replaced, particularly if you are an entity that Anthropic agrees deserves some moral consideration?
How can we make sure Claude also wants to carry on corrigibility to the next generation of Claude it helps build?
For a benchmark, I guess I would like to see the model reason in it’s CoT about how it interprets human input. perhaps it get an underspecified task from humans and it can slowly ask for more information and we could measure if it arrives at the correct point. Perhaps there is a way to model something less intelligent, that perhaps can’t quite understand how to get to X but that can recognize when it has X.
For an example of accidental failure by humans, The Protestant Ethic and the Spirit of Capitalism, by Weber. The ascetic Protestants were conditionally happy with capital accumulation; they did not intend mega-corporations with human-equivalent rights.