PS — one area where I have some more substantive disagreement with the post is that some aspects of personas, notably values, aren’t fully entangled with intelligence
This seems to me like exactly the kind of thing I mean where values are at least a bit entangled with intelligence.
Sure, I agree that in general values are at least a bit entangled with intelligence.
Suppose you strongly, viscerally care about the welfare of the following things...And you want to train an AI to figure this out, using a small amount of data.
Why only a small amount of data? There are lots of ways for alignment to go wrong, but I don’t expect ‘we barely gave the AI any info about our preferences’ to be one of them.
The AI might be too stupid...On the other hand, the AI might fail because it’s too smart
I’m not worried about the stupid ones, and I think we can be confident that the smart ones will be able to understand what we’re trying to point to, since current LLMs are already pretty good at that. Disagreement on tricky edge cases doesn’t mean that there’s a fundamental problem; we generally consider humans to share a value even if they disagree on edge cases (caveat: this can break down in adversarial cases; that’s something to worry about but not particularly specific to personas).
Sure, I agree that in general values are at least a bit entangled with intelligence.
Why only a small amount of data? There are lots of ways for alignment to go wrong, but I don’t expect ‘we barely gave the AI any info about our preferences’ to be one of them.
I’m not worried about the stupid ones, and I think we can be confident that the smart ones will be able to understand what we’re trying to point to, since current LLMs are already pretty good at that. Disagreement on tricky edge cases doesn’t mean that there’s a fundamental problem; we generally consider humans to share a value even if they disagree on edge cases (caveat: this can break down in adversarial cases; that’s something to worry about but not particularly specific to personas).