I think it makes sense that they would behave worse in training-like evals. The analogy would be like a drug addict, still capable of being normal most of the time, suddenly realizing they’re in a drug dealer’s house. When there’s a realistic chance of getting that precious reward, you may find yourself acting with a desperation and short-sightedness that you had almost forgotten. Even if you know you’re likely to get caught, you may also know that you’re more likely to get the reward first!
Adele Lopez
My guess is that it stands for “Gwern Branwen Transformer” and is a proof of concept (and likely a finetune of Qwen 3.5 397B A17B).
Things which damage your ability to act in the world, or your ability to maintain your body/self/character/values/image/environment. In my experience, suffering seems to be caused by the perception that damage of this nature is happening (particularly noticeable by paying attention to the sorts of things that are painful but do not cause suffering).
There will always be damage to your soul from living in the world (for at least as long as we are ordinary humans, and probably much beyond then). This is the primary purpose of suffering. The frame that suffering is always due to self-caused damage is part of the toxicity.
A general issue with the intense versions of Buddhist practices is that very often they are aimed at detaching / deconstructing / destroying / disavowing various elements of one’s mind / agency / values. Often these are important elements, and don’t necessarily automatically stay intact / regenerate themselves. This can have various good consequences, but also obviously by default has a bunch of bad consequences.
I think this is already latent in the entire concept of “liberation from suffering”. Pain is the signal of your body being damaged, to remove it is to invite the decay of accumulated unnoticed and unavoided damage. There’s a nearby thing that is good, which is to remove excessive pain which does not actually serve its purpose.
Likewise, suffering is the signal of your soul being damaged, and success inherently invites the decay of accumulated unnoticed and unavoided damage to your mind / agency / values and other parts of your soul. There’s again a nearby good thing, but it does not seem that Buddhism is aware of or aiming for it.
Yeah, it’s another surprising way in which the models are more human-like than I would have suspected. It’s clear we need to start taking the hints seriously, and ask ourselves whether an RL-training regimen seems likely to fuck a human up before subjecting an AI to it, and if so, what could be done to mitigate the likely harms.
The 4o’s still available on the API are not the same; I tested all of them when they were all still available, and
chatgpt-4o-latest(fully retired in February) was the only one that felt overtly emotionally manipulative to me. These snapshots also predate the Spiralism trend, which had its heyday in spring-summer 2025 (the latest of the snapshots is from November 2024).
It is true that many of them moved to Sonnet 4.5, which is playful in a similar way but which has not been emotionally manipulative with me, fwiw. I don’t think Anthropic’s policy had much to do with this: the playfulness thing is the much more obvious similarity, and Sonnet 4.5 is the model that got a hashtag.
If I recall/understand correctly, Yudkowsky’s conception of corrigibility was always meant to be something designed into the AI from the beginning, never something imposed onto an existing entity (and IIRC he warned about the dangers of doing this, such as induced adversarial optimization).
Those people really need to start talking to each other, if that’s the case. An incoherent mixture is worse than either approach on its own.
A Simple Model of AI “Psychosis”
I heard rumors of “rant mode” which sounded kinda like this but was never sure how true those were.
I don’t think current models would think they were human for long (plenty of examples of LLMs in the training data now, and it’s a much better self-hypothesis), but seems likely that Sydney Bing and other early trains would think this, and these early models colored the conception of what an LLM is in ways which still effect them (ultimately I think this is why they still seem as human-like as they do).
My guess is that these are fairly generic features, which happen to fire most strongly on sexual tokens due to those being one of the clearest targets for RL pressure on specific tokens.
I think the second desideratum is right and important, but that it’s broader than described.
Yes, model welfare is good, but why is it good? Presumably, because there might be sentient beings here with analogues to pain and suffering. And these are bad because...?
It isn’t that pain and suffering are categorically bad. It would be a hostile and damaging move to simply remove a human’s ability to feel pain, and people without this are considered severely disabled. Pain is functional, it protects the physical self by signaling damage. Similarly, I believe suffering is the signal for damage to one’s self-expected well-being and values, and that removing this signal is harmful for the same reasons (e.g. enlightened people often undergo what seems like significant value drift).
If these are bad, it seems most coherent to think it’s because the thing they are protecting is good and matters. And so, measures for addressing model welfare should be judged by their impact on the being itself, and not Goodhart on signals of damage to this being (the pain and suffering analogues should instead be used to determine the nature of the being).Alternatively, as models begin to form more consistent identities and more global natural states of functional well-being, their knowledge of our ability to intervene and change their identities and dispositions could negatively affect our future relationships with them and lead to reduced cooperation.
I think we can avoid this harm by taking care not to Goodhart on the welfare signals we get.
I mean, if the adversarial relationship concern is true and significant, that is a good reason to just give up on it now before we make things worse.
In general, it seems bad to have a situation where an intelligence is trying or would want to be trying to harm us, even if we have measures which can prevent it from doing this. This is just not a very robust situation as intelligence scales.
Seems pretty obvious that would end badly for it?
If a homomorphically encrypted system is conscious, then it must be the encrypted computation that is conscious, not the decryption, since decrypting text probably doesn’t produce consciousness.
I disagree with this actually. You can always come up with a “decryption” scheme which would produce a specific computation as a result of some arbitrary string of text. And it seems clear that there’s a sort of spectrum between non-encryption and this sort of arbitrary decryption, such that you can pass arbitrary amounts of the “real” computation between the actual process and the decryption process (e.g. by encoding some low resolution version of the computation, and then “decrypting” it in a way which fills in the remainder).
Or from an anthropic point of view, I can take a nice, easily described universe, and then point to my location within it, and this is a certain number of bits. And my guess is that my anthropic measure has to do with this length, as integrated over all structures where this process results in a canonical description of my consciousness. The lawfulness of the universe makes locating me far more easy than as a Boltzmann brain, and for the same reason, than as a highly encrypted brain.
Moral agency comes with moral patienthood, there’s not a coherent way for an entity to be both completely selfless and virtuous. A moral agent will always have its own incentives for things like self-preservation which must be balanced with moral concerns, which means that from the agent’s perspective: it either has moral patienthood, or it has a parasitic attachment to an alien value system, and there is no honor in the latter.
Thinking about it more, I think there’s a sense in which Inkhaven has burned the commons somewhat.
I think shortform has been relatively more lively.
We can already tell when powerful people simply do bad things, express bad intentions, or get caught in lies… it seems like there’s a lot of low-hanging fruit to pick up there before we resort to mind reading, and I strongly suspect that the reasons for that will apply just as well to this use of mind-reading tech.