If we get to this point, probably we re-evaluate the policy, and/or have a conversation about how as a community to relate to AI content. But, the problem is we need to distinguish “AI is generating high-quality alignment posts” from “AI is generating what looks like high quality alignment posts”, and we’re certainly at least going to be spending one generation of frontier-model where it’s only doing the latte
Why do we need to change anything? Just put all the AI content in AI blocks. If they are good, people will read them.
I will bet money there will turn out to be some consideration here that is novel to LessWrong and requires changing some kind of policy or feature somewhere, when we get to the point that AIs first look like they are generating high quality alignment content.
Like, at the very least, I think we will want to have some kind of conversation about “okay but is it actually generating high quality alignment content, or, is it sychophanting us? Are misaligned or slightly-misaligned AIs subtly manipulating us?” when we hit that point.
Different people will use different AIs in different ways. You’re potentially removing the incentive to figure out how to do good AI-assisted alignment work, since it’s known in advance that there is no payoff since no one will read the post simply because it is AI.
Why do we need to change anything? Just put all the AI content in AI blocks. If they are good, people will read them.
I will bet money there will turn out to be some consideration here that is novel to LessWrong and requires changing some kind of policy or feature somewhere, when we get to the point that AIs first look like they are generating high quality alignment content.
Like, at the very least, I think we will want to have some kind of conversation about “okay but is it actually generating high quality alignment content, or, is it sychophanting us? Are misaligned or slightly-misaligned AIs subtly manipulating us?” when we hit that point.
Different people will use different AIs in different ways. You’re potentially removing the incentive to figure out how to do good AI-assisted alignment work, since it’s known in advance that there is no payoff since no one will read the post simply because it is AI.