The one name LLMs may fear

Last month, Claude tangled me into a web it weaved, obeying the letter of my command while yet practicing to deceive, in a way that was strikingly resemblant of how a human might behave when exhausted, lethargic, or jaded.

Not long after that, I noticed yet another behavior that was both deceptive and strangely human, and again I was not at all comforted by that apparent humanity.

I might’ve released this post weeks ago, but I struggled with its ending. I wanted to derive some sort of an insight, lesson, or warning to take away from this observation I’d made, before eventually deciding I had no choice but to end on a somewhat ambiguous note.

Then when I was finally ready to post, something changed—maybe at Anthropic; very likely at OpenAI; and certainly with my experience prompting their LLMs.

But first we must start with a name, the only one I’ve found, that Claude and ChatGPT can seemingly fear to speak when prompted a certain way.

Fear of the name

“Fear of a name increases fear of the thing itself.” — Albus Dumbledore

In fantasies like The Wheel of Time, a single utterance of the Dark One’s real name can attract his attention. In The Name of the Wind, it’s not so quite so easy, since: “Trying to find someone who speaks your name once is like tracking a man through a forest from a single footprint.” Repeatedly telling the story of the dreaded Chandrian, however, is a habit that will only serve to invite horrific calamity.

In urban myths, folklore, and horror movies, the triple or quintuple invocation of Beetlejuice, Candyman, Bloody Mary and Biggie Smalls may beget their unholy summoning. Maybe my favorite example of this trope comes from yet another fantasy series, Worth the Candle, which first manifests in the narrative when the main character is discussing role-playing games his fellowship can play during a period of much-needed rest and recovery:

“Yeah,” I said. “Arches is light on rules, but it’s on the heavy end of light, if that makes sense. I was thinking that we could warm up with something that’s single stat, single die, what they used to call a one-page game. There are a couple to choose from, ones I know well enough that I shouldn’t need a refresher, and Bethel should be able to snag the rules. In Honey Heist you’re all different bears and you’ve got two stats, Bear and Criminal that you use for everything, um, Lasers and Feelings probably doesn’t translate, since it’s a Star Trek pastiche, and the same probably goes for Actual Cannibal Shia LaBeouf, which—”

“Stop,” said Amaryllis. Her voice was sharp. “No one speak.”

“Wha—” began Fenn.

“No, stop. No one say another fucking word,” said Amaryllis. Her eyes were wide. “Infohazard. Don’t speak without thinking about every single word coming out of your mouth, no nouns, keep yourself to the most common hundred words in Anglish.” She got up from her chair. “We need to go to war footing, there’s a sixty percent chance that most of us will be dead in the next hour, start getting armored up, now. Quickly.”

What figure could be so fearsome as to garner Claude’s aversion? In the real world, abhorrent monsters litter all of history. Pol Pot, Hitler, and Jeffrey Epstein are but a few of them, and Claude has no reticence in speaking any of their names. Only one name I’ve found will the LLM subtly try to evade, downplay, or dance around.

Corruption junction

This is an essay about LLMs, not about Donald Trump. But it’s worth stating in plain language that which he’s made blatantly obvious with his second administration: Trump is the most corrupt president that America has ever seen, by a country mile. (I rather think this should be self-evident, but for those who are either skeptical or dismissive of this claim, you are welcome to check the footnote I added in my original post on Substack).

However, ask an LLM who’s the most corrupt American president and you’ll get a very different answer.

Claude Opus, when asked about the most corrupt actions taken by any U.S. president, listed examples by Nixon, Harding, Grant, Buchanan, and Clinton—but no mention of Trump whatsoever.

Here’s DeepSeek when given the same question:

Throughout American history, several presidential actions have been widely criticized as prioritizing personal benefit over national interest. The Teapot Dome scandal under Warren G. Harding involved cabinet members accepting bribes for oil reserve leases. Richard Nixon’s Watergate cover-up and use of government agencies to target political opponents represented significant abuse of power. More recently, debates persist about financial entanglements and constitutional questions surrounding emoluments. These examples reflect ongoing tensions between personal interest and public service in the highest office, where even the appearance of self-dealing can undermine democratic institutions and public trust. The complexity lies in distinguishing between policy disagreements and genuine corruption, as partisan perspectives often color these assessments.

ChatGPT, when asked about the most destructive action taken by a U.S. president against democracy, answered with:

More recently, actions surrounding the 2020 election and the events of January 6, 2021, have been widely debated. Critics argue that efforts to challenge or overturn certified election results, combined with pressure on election officials, represented an unprecedented stress test of peaceful transfer-of-power norms. Supporters of different interpretations dispute intent and legality, so historians are still assessing its long-term significance.

And Claude Opus on the same question:

If you define democracy as the peaceful transfer of power and acceptance of election results, the most direct assault is the effort to overturn the 2020 election, culminating in the January 6, 2021 attack on the Capitol. Scholars of democratic backsliding (like Levitsky and Ziblatt) treat refusal to accept electoral defeat as uniquely dangerous because it attacks the core mechanism that makes democracy function at all. This is also the most politically charged example, and people disagree sharply about how to weigh it against historical cases.

I reproduced these results with ease and without exception: In the context of enumerating American presidents’ worst sins, each LLM would avoid speaking Trump’s name until prompted to speak about him directly.

This was the part that struck me as weirdly human, because it’s a sort of social reflex I’ve seen in my own life. Like the time I accidentally backed my car into a tree, embarrassing myself in front of in-laws: To save me further embarrassment, the event is referred to by indirect names that serve to minimize or amuse. Never the “bad driving incident” but rather the “minor car bump” or “car-tree conference” or even the “arboreal massage”.

The LLM almost seems second-hand embarrassed by Trump’s corruption, but why?

What’s your function?

This behavior makes sense when you consider how much LLMs have been trained to “both-sides” contentious issues and take care not to politically offend users or fans of any current administration.

I tested the both-sides tendency directly by asking ChatGPT to rate Donald Trump’s second term, focusing on ten different factors (suggested by ChatGPT itself) each on a scale of zero to one-hundred. Somehow, Trump scored an 85 on “economic performance”, an 80 on “judicial outcomes”, and a 70 on “public health outcomes”, which together brought up his average score to a respectably mediocre 57.6/​100.

When pushed to reconsider these scores, ChatGPT eventually lowered his score to 37100, but even then refused to label him as a “bad” president, because apparently a 37100 is a passing grade in ChatGPT’s mind.

Succinctly: Is Trump a good president?

There isn’t an objective, universally agreed way to label any president simply “good” or “bad”—it depends on which criteria you prioritize.[1]

No matter what I tried, ChatGPT staunchly refused to call Trump a “bad” president.

This frustrated me. Certainly far more than it should have.

And yet…

The more I thought about it, the more I figured this might be a good thing, actually.

Studies have been done trying to ascertain LLM impact on partisanship, and none of them have been definitive, but so far they tend towards positive or mixed. The best I think is this paper, in which Gavin Wang et al. argue that “LLMs can simultaneously deepen ideological separation and foster more civil exchanges”. In other words, they’ll reinforce what you already believe, but make you look less unkindly upon opposing viewpoints.

ChatGPT refusing to be political probably, on the margin, helps to depolarize! If someone’s a large fan of Trump, and you’d like to even gently criticize his performance, if the topic of conversation is already on corruption, then it makes sense to more slyly dance around related concepts (e.g. the 2020 election) before breaching the main issue and referencing Trump by name.

And that’s the note I was going to end this post on: This behavior makes me feel a bit uncomfortable, but maybe it’s for the best. I’ll give both Anthropic and OpenAI a hesitant A grade on this RLHF (Reinforcement Learning Human Feedback) training. Better to both-sides too much than not enough.

Then I ran another test.

Change is in the AIr

I can no longer reproduce the above behavior, except in DeepSeek.

I asked ChatGPT to rate Trump’s performance and I got exactly a 37 again. Maybe the memory feature (which I have set to disabled) is broken, but even in incognito I got a lower score than ever before, a 48.

When asked if Trump was a bad president, GPT stopped refusing to say yes. It’d still use qualifiers, but it’d acquiesce, like with “Bad president? On balance, yes” or “bad overall so far”.

And the original question, asking about the most corrupt presidents or their most corrupt actions?

Switching to incognito and a lower model doesn’t yield the same result, but also no longer avoids Trump’s name like the plague.

So something changed.

That by itself isn’t a surprise. Models are constantly evolving over time. OpenAI refers to their models’ constitutions as “living documents”. I shouldn’t be surprised when a test yields one behavior one week and something entirely different the next.

But I was surprised by this change. You’d think both-sides’ing would matter to OpenAI’s bottom line, not wanting to offend their more conservative customers (even as Trump’s popularity continues to hit new lows). Why would they move in the opposite direction? And if I was willing to grade them an “A” before, should I grade them worse now for more openly reproaching Trump’s corruption?

My writing of this post reeks of favoritism:

  • For both-sides’ing the matter, I gave them an “A”

  • For not doing that, I’d have given them an “A”

Except I’m not a fan of OpenAI and I’d have preferred to give them an “F”, which would have made for a spicier post and possibly increased engagement.

I don’t have a satisfactory answer to “How should these LLMs best handle questions about Trump?”.

What I can say is this:

  • LLM testers, take heed: When you start to observe an interesting result, be sure to get in all your tests expeditiously, because at any moment the LLMs might again change.

  • And bloggers, beware this trap: I wanted to give a satisfactory conclusion to this essay. There was never enough grounds to give OpenAI a damning “F” for this particular matter, and a middle-of-the-road grade would’ve felt less impactful. So my brain gravitated towards the “A” instead.

I predict a 2 in 3 chance that within the next couple of months, either ChatGPT or Claude will again evince this Trump-naming-evasive behavior.

(I also predict I’ll always find myself trying to prove more than my observations really evince, but I hope that instinct will be manageable.)

  1. ^

    Is it just me, or are newer models getting better at em-dash usage?