Something I don’t understand about this is that Anthropic’s approaches to constitutional AI and persona selection seem to be prioritizing value alignment over corrigibility. But maybe it’s as simple as “the people who wrote this report were not involved in Claude’s constitution”
Something I don’t understand about this is that Anthropic’s approaches to constitutional AI and persona selection seem to be prioritizing value alignment over corrigibility. But maybe it’s as simple as “the people who wrote this report were not involved in Claude’s constitution”
Those people really need to start talking to each other, if that’s the case. An incoherent mixture is worse than either approach on its own.