[ edit: I have substantially changed my beliefs stated here, based on how bad the AI version is, on closer inspection. It’s not durable, but just AI-identification is probably helpful in the medium term ]
I worry a lot that the binary “AI-written” filter is a completely different dimension from what we actually want: a quality indicator for things that haven’t gotten many votes yet. Let’s consider how and why you want junior MATS-scholar contribution (with massive AI assistance in writing) and don’t want an outside contribution (with massive AI assistance in writing). I suspect we’re going to need to get to a point where the site grades (and maybe categorizes as to likely favorable audiences) posts using AI, rather than trying to segment.
As you say, nothing useful is going to be AI-free for very long. I’m embarrassed at the time I’ve spent on this comment, and I suspect a brief interview with Opus 4.6 would have produced one more concise and useful.
Actually, yes—here’s what I should have done, in about 1⁄10 the time (opus 4.6, with input of your comment and 3-four fragments of points I want to make, followed by a request to make more concise):
The “AI-written block” assumes a stable boundary between AI and human content that’s already gone. My thinking, framing, and editing are all AI-assisted. Where does the block go?
The coding analogy undermines the proposal: we don’t flag which lines Copilot wrote — the question is whether the output is correct and useful. Same here.
The actual signal you need is epistemic quality: original judgment vs. vocabulary pattern-matching. AI-block markup doesn’t measure that — it’s compliance theater that honest users follow and bad actors ignore. More tractable approaches: reputation systems, structured epistemic standards, or ironically, AI-assisted grading of submissions against LW’s actual quality criteria.
Who knows how much this is biased by AI slop taste, but, your AI comment feels kinda contentless to me in a way your original one doesn’t.
“What you want is a signal of epistemic quality.” Well, yeah, no shit. That’s a very difficult problem that it glosses over.
Features of your original comment that make it more interesting, apart from me just kinda barfing at the writing style which I’ll try to ignore:
“quality indicator for things that don’t have many votes yet”
why you want junior MATS-scholar contribution (with massive AI assistance in writing) and don’t want an outside contribution (with massive AI assistance in writing)
Both of those highlight gears of the problem that help me think about it. And then, there’s something like “I’m confident Dagon actually believes this is the shape of the problem” that is somehow helpful for feeling like I’m having a real conversation where I expect us to be jointly improving our models.
Things I actively dislike about the about the AI one:
“it’s compliance theater that honest users follow and bad actors ignore. More tractable approaches: reputation systems, structured epistemic standards,”
I think the first sentence is just false (we get to enforce it on bad actors and establish norms), and the “reputation systems” and “structured epistemic standards” are like, well, figuring out how to do that is whole problem.
It seems like the problem you’re articulating has to do with the fact that LessWrong functions partly as a training ground for rational thinking and AI alignment research. In the past, you’ve lowered the epistemic bar for student-tier content, because receiving feedback on their content motivates them and provides useful feedback to them.
But soon, slop-posters may meet or exceed the epistemic bar achieved by students. Judging it on a pure minimum quality standard would admit a tidal wave of useless, inert slop. It’s not good enough to contribute to the leading edge of the conversation, and it doesn’t benefit from feedback. It just drowns out the authentic student content, which really does benefit from community input.
If that accurately reflects the problem you’re concerned about, then one possibility is to enforce an escalating minimum quality bar that exceeds both the quality of AI slop and the current minimum standards of LessWrong today. Anybody can freely post, but if moderators don’t feel the quality is excellent, then it does not get any visibility on LessWrong.
Simultaneously, create a separate submission channel for “student contributions” where quality standards are lower, but the primary rationale is that the poster would benefit from community feedback to support their intellectual growth.
This strategy would introduce the question of how to set standards and vet submissions through the student channel. This may take quite a bit of thinking to strike the right balance between effectively filtering for what you want and not taking undue amounts of moderation effort. Some possible aspects of such filtering might include:
Making default acceptance of submissions through the student channel a benefit of participating in real-world community activities, such as Inkhaven, MATS, workshops, and so on. This functions as a costly signal of a person who’s earnestly trying to use these resources to improve their capabilities. When people register for these events, they can register their account username as well, which labels that account as being permitted to make “student submissions.”
Posting a sufficient amount of excellent content might also enable people to post through the “student submissions” channel.
Giving the community a way to flag posts submitted through the student channel that feel AI-corrupted, with a low threshold for those posts being taken down automatically due to being flagged.
We may generally want to raise our quality bar. But, fwiw I don’t know that we’ve actually lowered the bar for student tier content. Or, idk maybe we do, but, I don’t think MATS scholars are particularly below our bar (depends on the scholar/project). Just because they’re not contributing frontier conceptual progress doesn’t mean they’re not, like, exploring an interesting corner of the world and writing up some useful stuff about it.
I mention MATS scholars because their work is structurally similar to the current generation of slop (i.e. it sorta looks like the slop is imitating entry-level mechinterp work in particular).
I don’t know that we’ve actually lowered the bar for student tier content.
Fair! Let me rephrase. LW may have historically set its quality floor low enough to permit student-tier content, because even if it’s of minor interest, it has a beneficial side effect in promoting the author’s intellectual growth and potential to contribute in a more substantial way in the future. Most content above the current quality floor reflects enough content interest and growth value to be worth accepting.
When AI slop consistently rises above that quality floor, then including it will fill LW with minor-interest slop that has no beneficial side effect of intellectual growth for the contributor. So it will become untenable to keep the floor in the same place. But LW still wants to give student-tier contributors a way to make those contributions without getting drowned out by AI slop or filtered out by a quality bar that has to keep rising to filter out the slop. The strategies I proposed are implementation ideas for that alternate submissions channel.
Is this a fair description of the problem, as you see it?
As a datapoint, your first two paragraphs are interesting enough to read, and then my eyes glaze over a lot in the AI block. I forced myself to read it anyway, and indeed it makes less interesting points with less interesting words. It feels like shoveling poop back and forth for no reason to go through the whole AI block, but for example, “original judgment vs. vocabulary pattern-matching”—this is a much less interesting thing to say than your interesting question of “what do we care about in novice research vs. outsider maybe-slop”.
I agree with both you and Raemon—the AI portion is hugely worse than the hand-written comment. And I suspect it generalizes—AI can be as good or better than human writing, with somewhat less effort, it’s not often that sufficient effort is taken.
I still suspect that identifying AI won’t be sufficient, but I fully concede the point that the vast majority of AI writing is less useful than the majority of human writing.
I do use AI for most coding, where “good enough” is in fact good enough, but I see that I’ve got further to go in figuring out how to guide and correct it for writing.
It sounded like when you first wrote your comment like you saying “this is mostly better, or, at least succeeded in saying what I meant”, was that a thing you changed your mind on or did you not mean to convey that?
It’s a thing I changed my mind on, based on your comments and my re-reading it more critically (and really, reading it thoroughly at all). It’s a perfect reminder to me of one of the main failure modes of LLM assistance—it’s good enough at first glance that it’s easy to forget to apply the same level of self-critique and thought one does for direct writing.
I don’t have a good way to detect this failure mode in myself, let alone others, but it’s very apparent when I look, and is probably common enough that “is it substantially AI” is an ok proxy for “is it low-quality”. This is a reversal of my previous position, though I still suspect it won’t last for long.
Man, I don’t know if I’m confabulating it, but the part that Opus wrote really has the feel of LLM text to me, in a way that I don’t like, such that I would be sad if all your comments had that style.
I worry a lot that the binary “AI-written” filter is a completely different dimension from what we actually want: a quality indicator for things that haven’t gotten many votes yet.
I do think it may straightforwardly be possible to build an AI model that predicts “karma within first 2 weeks” or something with reasonable accuracy, and that that’d be an improvement over status quo for determining post visibility for like the first 6 hours (maybe we gradually fade from “AI predicted karma” to “real karma” continuously over some period).
Even if such a model was perfectly accurate, I think that would have to introduce distortions because visibility impacts if and how people vote. The karma a post earns when it is displayed based on current karma will be different from the karma a post earns when it is displayed based on predicted ~final karma.
[ edit: I have substantially changed my beliefs stated here, based on how bad the AI version is, on closer inspection. It’s not durable, but just AI-identification is probably helpful in the medium term ]
I worry a lot that the binary “AI-written” filter is a completely different dimension from what we actually want: a quality indicator for things that haven’t gotten many votes yet. Let’s consider how and why you want junior MATS-scholar contribution (with massive AI assistance in writing) and don’t want an outside contribution (with massive AI assistance in writing). I suspect we’re going to need to get to a point where the site grades (and maybe categorizes as to likely favorable audiences) posts using AI, rather than trying to segment.
As you say, nothing useful is going to be AI-free for very long. I’m embarrassed at the time I’ve spent on this comment, and I suspect a brief interview with Opus 4.6 would have produced one more concise and useful.
Actually, yes—here’s what I should have done, in about 1⁄10 the time (opus 4.6, with input of your comment and 3-four fragments of points I want to make, followed by a request to make more concise):
Who knows how much this is biased by AI slop taste, but, your AI comment feels kinda contentless to me in a way your original one doesn’t.
“What you want is a signal of epistemic quality.” Well, yeah, no shit. That’s a very difficult problem that it glosses over.
Features of your original comment that make it more interesting, apart from me just kinda barfing at the writing style which I’ll try to ignore:
“quality indicator for things that don’t have many votes yet”
why you want junior MATS-scholar contribution (with massive AI assistance in writing) and don’t want an outside contribution (with massive AI assistance in writing)
Both of those highlight gears of the problem that help me think about it. And then, there’s something like “I’m confident Dagon actually believes this is the shape of the problem” that is somehow helpful for feeling like I’m having a real conversation where I expect us to be jointly improving our models.
Things I actively dislike about the about the AI one:
“it’s compliance theater that honest users follow and bad actors ignore. More tractable approaches: reputation systems, structured epistemic standards,”
I think the first sentence is just false (we get to enforce it on bad actors and establish norms), and the “reputation systems” and “structured epistemic standards” are like, well, figuring out how to do that is whole problem.
It seems like the problem you’re articulating has to do with the fact that LessWrong functions partly as a training ground for rational thinking and AI alignment research. In the past, you’ve lowered the epistemic bar for student-tier content, because receiving feedback on their content motivates them and provides useful feedback to them.
But soon, slop-posters may meet or exceed the epistemic bar achieved by students. Judging it on a pure minimum quality standard would admit a tidal wave of useless, inert slop. It’s not good enough to contribute to the leading edge of the conversation, and it doesn’t benefit from feedback. It just drowns out the authentic student content, which really does benefit from community input.
If that accurately reflects the problem you’re concerned about, then one possibility is to enforce an escalating minimum quality bar that exceeds both the quality of AI slop and the current minimum standards of LessWrong today. Anybody can freely post, but if moderators don’t feel the quality is excellent, then it does not get any visibility on LessWrong.
Simultaneously, create a separate submission channel for “student contributions” where quality standards are lower, but the primary rationale is that the poster would benefit from community feedback to support their intellectual growth.
This strategy would introduce the question of how to set standards and vet submissions through the student channel. This may take quite a bit of thinking to strike the right balance between effectively filtering for what you want and not taking undue amounts of moderation effort. Some possible aspects of such filtering might include:
Making default acceptance of submissions through the student channel a benefit of participating in real-world community activities, such as Inkhaven, MATS, workshops, and so on. This functions as a costly signal of a person who’s earnestly trying to use these resources to improve their capabilities. When people register for these events, they can register their account username as well, which labels that account as being permitted to make “student submissions.”
Posting a sufficient amount of excellent content might also enable people to post through the “student submissions” channel.
Giving the community a way to flag posts submitted through the student channel that feel AI-corrupted, with a low threshold for those posts being taken down automatically due to being flagged.
We may generally want to raise our quality bar. But, fwiw I don’t know that we’ve actually lowered the bar for student tier content. Or, idk maybe we do, but, I don’t think MATS scholars are particularly below our bar (depends on the scholar/project). Just because they’re not contributing frontier conceptual progress doesn’t mean they’re not, like, exploring an interesting corner of the world and writing up some useful stuff about it.
I mention MATS scholars because their work is structurally similar to the current generation of slop (i.e. it sorta looks like the slop is imitating entry-level mechinterp work in particular).
Fair! Let me rephrase. LW may have historically set its quality floor low enough to permit student-tier content, because even if it’s of minor interest, it has a beneficial side effect in promoting the author’s intellectual growth and potential to contribute in a more substantial way in the future. Most content above the current quality floor reflects enough content interest and growth value to be worth accepting.
When AI slop consistently rises above that quality floor, then including it will fill LW with minor-interest slop that has no beneficial side effect of intellectual growth for the contributor. So it will become untenable to keep the floor in the same place. But LW still wants to give student-tier contributors a way to make those contributions without getting drowned out by AI slop or filtered out by a quality bar that has to keep rising to filter out the slop. The strategies I proposed are implementation ideas for that alternate submissions channel.
Is this a fair description of the problem, as you see it?
Yeah that framing seems plausible, would have to think more.
As a datapoint, your first two paragraphs are interesting enough to read, and then my eyes glaze over a lot in the AI block. I forced myself to read it anyway, and indeed it makes less interesting points with less interesting words. It feels like shoveling poop back and forth for no reason to go through the whole AI block, but for example, “original judgment vs. vocabulary pattern-matching”—this is a much less interesting thing to say than your interesting question of “what do we care about in novice research vs. outsider maybe-slop”.
I agree with both you and Raemon—the AI portion is hugely worse than the hand-written comment. And I suspect it generalizes—AI can be as good or better than human writing, with somewhat less effort, it’s not often that sufficient effort is taken.
I still suspect that identifying AI won’t be sufficient, but I fully concede the point that the vast majority of AI writing is less useful than the majority of human writing.
I do use AI for most coding, where “good enough” is in fact good enough, but I see that I’ve got further to go in figuring out how to guide and correct it for writing.
It sounded like when you first wrote your comment like you saying “this is mostly better, or, at least succeeded in saying what I meant”, was that a thing you changed your mind on or did you not mean to convey that?
It’s a thing I changed my mind on, based on your comments and my re-reading it more critically (and really, reading it thoroughly at all). It’s a perfect reminder to me of one of the main failure modes of LLM assistance—it’s good enough at first glance that it’s easy to forget to apply the same level of self-critique and thought one does for direct writing.
I don’t have a good way to detect this failure mode in myself, let alone others, but it’s very apparent when I look, and is probably common enough that “is it substantially AI” is an ok proxy for “is it low-quality”. This is a reversal of my previous position, though I still suspect it won’t last for long.
Man, I don’t know if I’m confabulating it, but the part that Opus wrote really has the feel of LLM text to me, in a way that I don’t like, such that I would be sad if all your comments had that style.
I do think it may straightforwardly be possible to build an AI model that predicts “karma within first 2 weeks” or something with reasonable accuracy, and that that’d be an improvement over status quo for determining post visibility for like the first 6 hours (maybe we gradually fade from “AI predicted karma” to “real karma” continuously over some period).
Even if such a model was perfectly accurate, I think that would have to introduce distortions because visibility impacts if and how people vote. The karma a post earns when it is displayed based on current karma will be different from the karma a post earns when it is displayed based on predicted ~final karma.