Some other interesting correlations with capability:
Accepting the analytic-synthetic distinction
Moral generalism over particularism
Platonism about abstract objects
The one-world reading of Kant
Notable: prompting models to ignore philosopher consensus and think about their own position made a big difference, even with the most capable models. Fable is a nice case: it 2-boxes under the default framing and 1-boxes once you add the ignore-philosophers cue.
My prompt is about philosophers rather than Anthropic, so it’s not a direct test of @Chi Nguyen hypothesis about Claude predicting what Anthropic wants. But it at least highlights that some deference is going into the replies.
Full results: https://lordscottish12.github.io/PhilBench/
Oops, url fixed now.