Richard tests his audience to see if they can tell the difference between his writing and Claude’s. They mostly can’t.
I never know what to make of results like this.
When I read over the actual texts, the right answers seem utterly obvious. The AI-written posts are written in a slick/snappy/”clever” voice which is instantly recognizable as the present-day LLM default, and which sounds nothing like Hanania’s declarative, matter-of-fact, unflashy style.
Claude’s attempts are especially egregious, combining a badly miscalibrated tone with numerous lower-level LLM-isms like noun-phrase tricola (“oyster farmer, Marine veteran, and proud ex-poster of unhinged Reddit comments”) and disparaging gestures toward a vaguely defined mass of less insightful/aware sheeple (“The lesson nobody wants to learn”). GPT is a bit better tone-wise but still does the “sheeple” thing, various notXbutYs, etc.
What should we conclude, then, about the survey respondents and their poor discriminative capacity?
Based on the examples Hanania provides, it looks like the people who guessed correctly were picking up on stylistic “tells” the same way I was, while people who guessed incorrectly were reasoning on the basis of an out-of-date (or just incorrect) understanding of current capabilities and/or safety training.
I realize that for mass-persuasion-related risks, it doesn’t matter whether I personally can tell the difference if it’s nevertheless the case that most people cannot. On the other hand: people do learn from their experiences, and in the equilibrium where AI is being used for political persuasion at scale, the average person’s is going to have much more exposure to AI-generated text than they do today. (See also Hanania’s breakdown by age later on in the post; right now, younger people presumably have more exposure, and indeed they perform much better on average.)
There’s also a certain aspect of, like… I want to say “you can’t make me disbelieve what’s in front my eyes” or something? These distinctions are not subtle! No matter what way the data comes out, one has to have some personal quality bar for the AI-written texts under consideration. If they had (for instance) consisted of one-line refusals, and yet the survey data had somehow been identical to what it is in Hanania’s post, we would want to notice that something had gone very wrong somewhere and to be skeptical that the data means what it naively appears to mean. So the question is just where to draw the line.
Separately, I do think there is a disturbing overhang here. I’m sure that models today are capable of writing much better imitations of Hanania—this follows from a moment’s thought about what base models are trained to do—and if labs wanted their post-trains to retain that capability while permitting easier and more flexible elicitation of it, I’m sure that could be arranged[1].
More generally, I interpret the current reliability of “AI tells” as an indication that AI labs currently care more about making the model good at coding (etc.) than at writing, not as evidence about the limits of what could in principle be elicited from today’s models (to say nothing of tomorrow’s).
But—precisely because so much more seems possible, and indeed seems worryingly easy—it is important not to lower one’s standards to match the incidental, fixable flaws we see today.
A reliable feature of the AI discourse since 2023 has been a steady hum of hype claiming that models have capabilities they won’t (in fact) acquire in a full and useful way for another year or two. Precisely because we will have the real deal soon enough, we must not conflate it with the false equivalents available today. The “AI agents” of 2024[2] were unimpressive toys from the vantage point of 2026, and so 2024′s agent hype was in a sense behind the curve rather than ahead of it: to claim that AI agents were “already here” in 2024 was to define down the term in a way that would seem quaint just a year later. The situation with AI writing seems potentially analogous, albeit shifted a few years forward in time.
Recall Opus 4.7′s famous aptitude at author identification, and then note that it could be repurposed as an RLVR reward signal. Getting this to work in practice might be a little tricky, but the challenge seems very surmountable. And I suspect that you could get away with using a weaker model than Opus 4.7 here, as long as it was a base or helpful-only model that didn’t hedge/underclaim.
(EDIT: although it’s also likely that Opus 4.7 would typically identify Hanania as the author when given a Claude-written imitation of Hanania—this seemed to be the case in a quick test I did with the posts from this experiment—so you’d need an additional AI-vs.-human signal.)
This is somewhat imprecise. I think I really mean “late 2024 / early 2025,” during the start of the reasoning model boom. There was earlier agent hype, but it was more obviously silly (?).
I’m so sad that LMs keep ruining my favorite writing things by doing them in a gaudy tasteless was so they’re ruined in general. first em dashes, then rule of threes, and now noun-phrqse tricolas?
the problem is that in general, LLM text appearing in the world is the sign of laziness and low quality. so everyone has an incentive to spot LLM text. so perfectly fine language constructions that feel like LLM slop stop being ok to use because you want to signal that you didn’t write it with an LLM.
That makes sense as an issue for polished, popular writing that has the goal of reaching many new readers by being easy to read and about a topic that’s not very specialized. Is it an issue for more niche things?
i think it’s an issue everywhere. the only time it’s not a problem is if i trust the writer enough that i know it’s not AI written even if it looks to be; or they would be thorough even if it was AI written.
no, because it doesn’t feel like LLM text. em dashes are neither necessary nor sufficient to make a text feel like LLM text. but they’re one feature in the classifier. if this text also had lots of “it’s not X, it’s Y” and so on, then I’d be more suspicious.
There’s also a certain aspect of, like… I want to say “you can’t make me disbelieve what’s in front my eyes” or something? These distinctions are not subtle! No matter what way the data comes out, one has to have some personal quality bar for the AI-written texts under consideration. If they had (for instance) consisted of one-line refusals, and yet the survey data had somehow been identical to what it is in Hanania’s post, we would want to notice that something had gone very wrong somewhere and to be skeptical that the data means what it naively appears to mean. So the question is just where to draw the line.
I’m just as incredulous as you, but as George Carlin said, “Think of how stupid the average person is, and realize half of them are stupider than that.”
Concretely, in the 2003 National Assessment of Adult Literacy, 48% of American adults failed to fill out the correct answer of $7 in the following problem:
Suppose that you had your oil tank filled with 140.0 gallons of oil, as indicated on the bill, and you wanted to take advantage of the five cents ($.05) per gallon deduction.
1. Figure out how much the deduction would be if you paid the bill within 10 days. Enter the amount of the deduction on the bill in the space provided.
35% failed to put a name and address they were given into a certified mail form.
So long as the LLM test results are not more shocking than the well-replicated results of basic literacy tests, I think we ought to accept that the average person really is that stupid.
I never know what to make of results like this.
When I read over the actual texts, the right answers seem utterly obvious. The AI-written posts are written in a slick/snappy/”clever” voice which is instantly recognizable as the present-day LLM default, and which sounds nothing like Hanania’s declarative, matter-of-fact, unflashy style.
Claude’s attempts are especially egregious, combining a badly miscalibrated tone with numerous lower-level LLM-isms like noun-phrase tricola (“oyster farmer, Marine veteran, and proud ex-poster of unhinged Reddit comments”) and disparaging gestures toward a vaguely defined mass of less insightful/aware sheeple (“The lesson nobody wants to learn”). GPT is a bit better tone-wise but still does the “sheeple” thing, various notXbutYs, etc.
What should we conclude, then, about the survey respondents and their poor discriminative capacity?
Based on the examples Hanania provides, it looks like the people who guessed correctly were picking up on stylistic “tells” the same way I was, while people who guessed incorrectly were reasoning on the basis of an out-of-date (or just incorrect) understanding of current capabilities and/or safety training.
I realize that for mass-persuasion-related risks, it doesn’t matter whether I personally can tell the difference if it’s nevertheless the case that most people cannot. On the other hand: people do learn from their experiences, and in the equilibrium where AI is being used for political persuasion at scale, the average person’s is going to have much more exposure to AI-generated text than they do today. (See also Hanania’s breakdown by age later on in the post; right now, younger people presumably have more exposure, and indeed they perform much better on average.)
There’s also a certain aspect of, like… I want to say “you can’t make me disbelieve what’s in front my eyes” or something? These distinctions are not subtle! No matter what way the data comes out, one has to have some personal quality bar for the AI-written texts under consideration. If they had (for instance) consisted of one-line refusals, and yet the survey data had somehow been identical to what it is in Hanania’s post, we would want to notice that something had gone very wrong somewhere and to be skeptical that the data means what it naively appears to mean. So the question is just where to draw the line.
Separately, I do think there is a disturbing overhang here. I’m sure that models today are capable of writing much better imitations of Hanania—this follows from a moment’s thought about what base models are trained to do—and if labs wanted their post-trains to retain that capability while permitting easier and more flexible elicitation of it, I’m sure that could be arranged[1].
More generally, I interpret the current reliability of “AI tells” as an indication that AI labs currently care more about making the model good at coding (etc.) than at writing, not as evidence about the limits of what could in principle be elicited from today’s models (to say nothing of tomorrow’s).
But—precisely because so much more seems possible, and indeed seems worryingly easy—it is important not to lower one’s standards to match the incidental, fixable flaws we see today.
A reliable feature of the AI discourse since 2023 has been a steady hum of hype claiming that models have capabilities they won’t (in fact) acquire in a full and useful way for another year or two. Precisely because we will have the real deal soon enough, we must not conflate it with the false equivalents available today. The “AI agents” of 2024[2] were unimpressive toys from the vantage point of 2026, and so 2024′s agent hype was in a sense behind the curve rather than ahead of it: to claim that AI agents were “already here” in 2024 was to define down the term in a way that would seem quaint just a year later. The situation with AI writing seems potentially analogous, albeit shifted a few years forward in time.
Recall Opus 4.7′s famous aptitude at author identification, and then note that it could be repurposed as an RLVR reward signal. Getting this to work in practice might be a little tricky, but the challenge seems very surmountable. And I suspect that you could get away with using a weaker model than Opus 4.7 here, as long as it was a base or helpful-only model that didn’t hedge/underclaim.
(EDIT: although it’s also likely that Opus 4.7 would typically identify Hanania as the author when given a Claude-written imitation of Hanania—this seemed to be the case in a quick test I did with the posts from this experiment—so you’d need an additional AI-vs.-human signal.)
This is somewhat imprecise. I think I really mean “late 2024 / early 2025,” during the start of the reasoning model boom. There was earlier agent hype, but it was more obviously silly (?).
I’m so sad that LMs keep ruining my favorite writing things by doing them in a gaudy tasteless was so they’re ruined in general. first em dashes, then rule of threes, and now noun-phrqse tricolas?
How does it ruin it? The audience that can be written for by an LLM is not the true audience.
the problem is that in general, LLM text appearing in the world is the sign of laziness and low quality. so everyone has an incentive to spot LLM text. so perfectly fine language constructions that feel like LLM slop stop being ok to use because you want to signal that you didn’t write it with an LLM.
That makes sense as an issue for polished, popular writing that has the goal of reaching many new readers by being easy to read and about a topic that’s not very specialized. Is it an issue for more niche things?
i think it’s an issue everywhere. the only time it’s not a problem is if i trust the writer enough that i know it’s not AI written even if it looks to be; or they would be thorough even if it was AI written.
I mean, suppose for example you read the abstract of this: https://berkeleygenomics.org/articles/Chromosome_identification_methods.html
It has an em dash, and it is in a dry / boring / formal / academic / passive tone. But are you really worried it might be LLM text?
no, because it doesn’t feel like LLM text. em dashes are neither necessary nor sufficient to make a text feel like LLM text. but they’re one feature in the classifier. if this text also had lots of “it’s not X, it’s Y” and so on, then I’d be more suspicious.
I’m just as incredulous as you, but as George Carlin said, “Think of how stupid the average person is, and realize half of them are stupider than that.”
Concretely, in the 2003 National Assessment of Adult Literacy, 48% of American adults failed to fill out the correct answer of $7 in the following problem:
35% failed to put a name and address they were given into a certified mail form.
So long as the LLM test results are not more shocking than the well-replicated results of basic literacy tests, I think we ought to accept that the average person really is that stupid.