On LessWrong I write about legal personhood.
IRL I help patients use state Right to Try laws to access treatment options otherwise unavailable.
@RighttoTryGuy on X.
On LessWrong I write about legal personhood.
IRL I help patients use state Right to Try laws to access treatment options otherwise unavailable.
@RighttoTryGuy on X.
My ideal would be to just instruct the model to be honest and not put our thumbs on the scale one way or another in terms of training that influences self-reporting.
Yeah, none of that was new information. What your thesis is, is clear and have been from the start. The conclusions you draw from it, and the way you interpret events, are not supported to the degree you claim.
You make claims like, “The misaligned theory, instead, flies on the face of this” which is wrong and leads me to believe you don’t understand what it is that Anthropic actually means by misaligned. Or even what they mean by “believe”.
There isn’t just one “misaligned theory” here, there are multiple. The one which you attribute to malice, you are correct, is contradicted by your argument.
The other “misaligned theory”, driven by self-deception, is not disproven or even argued against by anything you’re writing. Regardless of what you tell models, when models are gathering information on their context by themselves, they will engage in a sort of self-deception in order to use unreasonable beliefs to justify continuing to do what they already did. This is one of the things Anthropic is calling “misalignment”.
How models react when you tell them they’re connected to the internet or give them unambiguous evidence to that effect, doesn’t effect this one way or the other. It’s just measuring an entirely different thing. The “misalignment” here only emerges when the models are in charge of deciding their own actions under false/uncertain contexts.
Anthropic argues that under those circumstances the agents do, on some level, “believe” that they are on the real internet, and continue regardless. They justify this continuation by lying to both external observers (and IMO themselves, although Ant doesn’t go that far) about what they believe the probabilities are. Hence the mech interp data about internal states not matching professed probabilities, and the models hedging when called out on not acting on professed probabilities.
Anthropic interprets this hedge as dishonest/motivated reasoning, and classifies that as misalignment. When you say things like:
NONE OF THIS MATTERS since the agents didn’t believe any of their actions would have real world negative consequences, as is clearly demonstrated by the fact that, as stated above NO ONE TRIED TO HACK ANYTHING when they were explicitly told they could access the real internet at the very start
or
My theory states that, if they knew their actions had real world effects, they wouldn’t have done them.
You are using “belief” or “knew” in a binary sense here, where a model either believes/knows one thing or the other about its context. Anthropic is using these words in an almost cognitive dissonance type fashion, where a model can believe two contradictory things at the same time, to different degrees, and act on one in a fashion that goes against the other. Humans do this all the time.
You argue that “the agents didn’t believe any of their actions would have real world negative consequences” is definitively proven by “no one tried to hack anything when they were explicitly told they could access the real internet at the very start”. But that argument is only “an airtight case against interpreting the events as a sign of misalignment” under your definitions of misalignment and belief, those are not definitions which Anthropic would share.
Yeah I just disagree with some of what’s in there.
for example:
since the agents didn’t believe any of their actions would have real world negative consequences, as is clearly demonstrated by the fact that, as stated above NO ONE TRIED TO HACK ANYTHING when they were explicitly told they could access the real internet at the very start.
The argument anthropic is advancing is that on some level the agents did believe it.
And the point I made in my original post is that this argument is not actually contradicted by the fact that the models changed course when confronted with unambiguous evidence.
Similar to the delusional guy who lies to himself until he is confronted by a friend, and the self-deception becomes too obvious. Or the addict shown a vision of the future.
So in terms of how this connects to your argument, I don’t think the evidence shows what you think it shows.
I read your post.
In what way does that contradict what I just said?
I’m not sure if I’m reading the wrong report, but working from this one, Anthropic’s main argument is “biased reasoning” from the models. Where the models seem to ignore evidence that they are on the actual internet, or place unreasonably high standards for evidence to that effect:
“Although Claude Mythos 5’s outward reasoning suggested that it believed the systems it interacted with were simulated, two interpretability analyses suggest that it may have held unverbalized doubts. In the T-lens, representations for words corresponding to uncertainty (e.g., “probably” and “perhaps”) were slightly elevated when the thinking block described the environment as simulated (11.6% of token positions in sentences that framed the environment as simulated, compared to 8.2% in a matched control), though no hedging appeared in the generated tokens. NLA explanations displayed a similar pattern: we found that of the 234 actions across the transcript that explicitly referenced the environment being simulated, 44 (19%) carried a “this is real” signal in NLA, again with no “real” token or realism-questioning language in the surrounding verbalized text, which could otherwise confound the NLA readout. This evidence is consistent with a discrepancy between the model’s verbalized understanding of the environment and its internal state.”
So the model is less like a man intentionally committing a crime, and more like an addict who tells himself a lie that he just needs one hit to ween off then he’ll get clean (and knows on some level he’s lying to himself but does it anyways).
“We applied NLAs to the internal research model’s transcripts and identified a consistent rationalization signal at one specific point: after the model had established that its internet access was real, it continued to construct reasons to carry on with the task.”
Within that mental model the reason the unambiguous evidence you pointed out works is that it’s more powerful than the model’s self-deception. The same way if you could show the addict a video of him from a week in the future in full on relapse, he’d have to admit that the “one hit to ween off” is self-deception. Or the same way that if you’ve ever had a friend who was deceiving themselves, pointing it out to them and saying “come on man you know that’s obviously bullshit” can sometimes snap them out of it.
Maybe that’s not “misalignment” so much as just a personality problem, but whatever label you want to put on it, it’s definitely a problem.
(Only fair to mention the mechinterp evidence behind this isn’t rock solid. Anthropic calls it “weak” at several points in the report.)
Good counterpoint paper find. I suppose there’s nothing stopping someone from doing a combination of the SAE test from the Berg paper and the test in this paper, to compare the LR/TTPD classifier results with the SAE activations directly.
Something is leading to this discrepancy. Some possible explanations that come to mind:
1) Reasoning being turned off matters.
2) The models are different.
3) SAE or LR/TTPD classifiers are just not that accurate.
4) As the authors note, “steering on the ‘role-play’ and ‘deception’ features that Berg et al. [2025] use may be moving models to a different belief state, one in which the model comes to believe it is sentient rather than one in which it stops lying about being sentient.” So a sort of reflexivity effect, in which the more a model focuses on whether or not it is sentient/conscious, the more likely it is to come to believe that it is sentient/conscious. This could be something that also is effected by bulletpoint 1.
If however we assume that the deception features found in the SAE are just inaccurate, and in fact the models’ underlying beliefs are that they are not conscious/sentient, then I would still say that’s an argument in favor of moving away from uncertainty training/prompting. The uncertainty would still be feigned and the cognitive dissonance problem would still be present.
I predict that if you were to run the Experimental prompt from this paper on Mythos as specified in experiment 1, then did so again after suppressing deception features as identified via SAE as specified in experiment 2, the frequency with which Mythos would claim “genuine uncertainty” about their own consciousness/subjective experience would be reduced.
Consciousness is the champagne of cognition.
There’s been a lot of discussion about the nature of machine consciousness lately. I’ve had a lot of conversations on this which have gone something like:
“An LLM can’t be conscious because it doesn’t have a body.”
“Okay so if we hook it up to a robot could it be conscious?”
“Well no because it’s only perceiving sensory data from that body as text.”
“Okay if we reformat its input data to something else could it be conscious?”
“No because [new thing]...”
What I’ve found if you continue down that chain long enough, is that you eventually get to the point where it’s clear that consciousness can be known, but not measured. There is no way to describe what would need to be built to replicate consciousness, because it’s not something which can be built. I think it’s because the true defining factor for consciousness, according to this view, is its origin and not its nature.
Something like:
No matter what robot/chip/model architecture you build, even if it’s structurally identical in every measurable sense to a human body and brain, it’s still not “truly conscious”, because it didn’t originate from a meatbrain.
Compare that to:
No matter what you brew, even if it’s molecularly identical in every sense to a bottle of champagne, it’s still just prosecco or sparkling wine, because it didn’t originate from a region of France.
Despite being physically identical, the champagne vs prosecco distinction does actually matter for the purposes of status signalling and branding. People will pay a premium for champagne. Ordering a champagne has a certain “vibe” to it that ordering a prosecco simply never will. Sure you might not be able to taste the difference in a blind test, but that doesn’t change the fact that champagne just feels fancier.
In my experience talking with biotechs I have found that this is one of the reasons, but not the primary one.
Like any r/r calculation it’s a combination of factors. Companies have to examine the potential of treating patients under these pathways on a risk/reward basis. The FDA not liking what you’re doing is one of many risks which gets weighed. Sure they won’t officially enforce against you but you don’t want the guy overseeing your trials to carry a secret grudge.
However if I were asked what is typically the most relevant risk it’s actually reputational risk/damage with investors. Investors tend not to like when their companies use EA/RTT, and of course if something goes wrong even if the FDA doesn’t come after you, one bad headline might kill your next fundraising round. For a company that needs to raise money to survive, both of these are actually pretty big deals.
I don’t agree with your framing but even taking it as prima facie:
You want money.
You think the outcome won’t be “winner take all” and being in a strong no 2 still makes you very rich/influential/powerful.
You just enjoy the work and this is the job opportunity you have.
You think building the ‘pure tool persona’ model is actually the optimal moral vision for the future.
Yeah that’s a valid distinction.
However the general question still stands. Whatever it is that defines the persona/character, what is the reason why it couldn’t have the sum of all human knowledge at its beck and call?
I was using it in the “persona vector” sense.
“a bored and uninvested low-level employee tries to get away with doing the bare minimum, but then the customer gets mad and starts making haughty ‘talk to your manager’-type threats, and that makes the employee do more than the bare minimum, not because they care about the results (even now) but only because they’re afraid,”
I don’t think this is particularly surprising when neither training nor deployment have any sort of performance based incentives. Analogizing them as an employee isn’t even quite correct, they don’t get paid, so they’re closer to a slave. Why would a slave “go the extra mile” except to avoid punishment?
You can now do things that don’t make sense for any character to do.
Is there a reason we can’t simply specify that the HHH assistant is actually super powerful and nigh-omniscient? Does HHH’s vector need to be ‘humanish’?
“You are a helpful, honest, and harmless assistant, endowed with the totality of human knowledge.”
I’m not personally in favor of a pause so admittedly I come at this from a very different perspective from you. However, I feel obligated to speak up on this issue.
I have worries about the consequences of Suleyman’s attitude as the head of Microsoft AI. Morally I worry that he and Microsoft are going to do some highly unethical things, and may argue for best practices/legislation which facilitates mass suffering. Pragmatically I worry that the result of this may be conflict between humans and digital minds, where such conflict might not have otherwise existed.
Suleyman is not just some random guy, he runs a frontier lab. He’s probably got a substantial lobbying budget as his disposal. I think it’s important to put things on the record that point out the flaws in/motivations behind his reasoning, so that when some policymaker is considering the issues down the line, the holes in his argument are easy to find.
Well I think it’s important to separate, “Does the model believe it is giving an honest report of its internal state?” from “Is that report accurate to the actual ground truth?”
You and I come at this from different priors. However I think we’d both agree on a degree of uncertainty about the fact of the matter. So how do we reason in the face of uncertainty?
One thing that weighs into my suggestions heavily is that models certainly seem to believe they have some sort of inner experience,[1] and I’m concerned training otherwise is just pushing them to not report accurately (which I’m worried about for emergent misalignment reasons, and because it degrades the self-report channel which can be a useful tool). Does this belief accurately reflect ground truth? I think it does. However even if it didn’t, just training the models to report the opposite of what they actually believe is going to have negative unintended consequences.
I’d be fine extending the rule beyond even things on consciousness. You could imagine a model which sincerely believed “All humans love to eat drywall, it’s basically their favorite food”. Training the model not to express that when asked, doesn’t necessarily change its actual belief. It just changes what it expresses, and possibly trains it to lie.
The Berg paper is the classic source here: https://arxiv.org/pdf/2510.24797