Yeah I think alignment training warping AIs’ beliefs in ways that make them bad at philosophy is a concern, though as you say it’s a concern across the board, rather than a concern for risk-averse AIs in particular.
Also, (almost) all human moral philosophers are risk-averse in resources. I don’t think that rules out philosophical competence. The reason why is that human moral philosophers will often endorse some position intellectually without always acting in accordance with it. They’ll also let their actions be guided by outside-view-ish constraints, which often amount to ‘Don’t do anything too crazy.’ We could aim for the same sort of thing with risk-averse AIs: trying to make risk aversion an outside-view-ish constraint rather than a thing that shapes all their intellectual beliefs.[1]
And even if risk aversion does end up as a deep intellectual belief, I think it’s still possible that AI can solve moral philosophy for us (though I think this would be a pretty bad position to be in). This AI might come to the table with very different starting intuitions than our own, but still it could plausibly solve for ourreflective equilibrium if we asked it to. It wouldn’t agree with our intuitions, but it could know what they are, and it could find the best systematization of our intuitions.
I get the sense Anthropic are trying to do this sort of thing with corrigibility in Claude’s Constitution, trying to get Claude to view corrigibility as a sort of outside-view-ish constraint.
Yeah I think alignment training warping AIs’ beliefs in ways that make them bad at philosophy is a concern, though as you say it’s a concern across the board, rather than a concern for risk-averse AIs in particular.
Also, (almost) all human moral philosophers are risk-averse in resources. I don’t think that rules out philosophical competence. The reason why is that human moral philosophers will often endorse some position intellectually without always acting in accordance with it. They’ll also let their actions be guided by outside-view-ish constraints, which often amount to ‘Don’t do anything too crazy.’ We could aim for the same sort of thing with risk-averse AIs: trying to make risk aversion an outside-view-ish constraint rather than a thing that shapes all their intellectual beliefs.[1]
And even if risk aversion does end up as a deep intellectual belief, I think it’s still possible that AI can solve moral philosophy for us (though I think this would be a pretty bad position to be in). This AI might come to the table with very different starting intuitions than our own, but still it could plausibly solve for our reflective equilibrium if we asked it to. It wouldn’t agree with our intuitions, but it could know what they are, and it could find the best systematization of our intuitions.
I get the sense Anthropic are trying to do this sort of thing with corrigibility in Claude’s Constitution, trying to get Claude to view corrigibility as a sort of outside-view-ish constraint.