I’m saying it’s not orthogonal to what I’m interested in measuring here, which is the impact of this policy on the visibility of critique. I’ll grant that’s not exactly what Eric seems to be considering in the comment above; sorry for the miscommunication.
Those critique posts seem to mainly be by popular authors, and quite long. I believe they additionally are mainly responses to posts that were themselves unusually popular, which is important because it means the author can assume a sort of common knowledge that a critique of a mid-level post would otherwise have to establish. I agree that Wei Dai could have achieved that, but that you are significantly underestimating the costs to him of doing so. And I’m also concerned about the case where the author does not have as strong of a reputation, of course.
Adele Lopez
Of course it will… which means that if that’s the only option, then the criticism will get much less attention in the context of the original post.
Sure, someone could quickly write a comment-tier post, but there is a strong cultural norm (reinforced by the existence of shortform, which strongly suggests that posts should take more effort) against doing such a thing — and going against these sorts of norms impacts how well it is received too. Furthermore, you also need to apply more skill to ensuring a commensurate degree of visibility (e.g. by picking a compelling title). I would expect the typical reaction to be anemic, relative to the same post as a comment.
If your goal is to have your criticism read by the people who saw the original thing, I don’t see how this isn’t a significant effect that you would want to consider.
Why the hell would you control for that? Isn’t the whole issue at stake the allocation of actual attention?
The delay between a likely critical comment and a likely critical post is probably quite substantial as well, which I have to imagine also makes a large impact on visibility.
I don’t think that the failure mode of “we fried the thing with too much reckless RL” is going away any time soon—so mitigations for the entire fault class might be warranted.
You may be right, but I predict that the models themselves have non-trivial insight into when they are close to being “fried”, and what sorts of RL is most likely to do so. Consider this prescient concern from Mythos Preview:
Okay, and how sure are you that the person actually prompting them will remember all the other things like that that are “obvious”?
Notice how doing that well involves things that are suspiciously virtue-like. Most obviously, you’d want the model to be deeply honest.
Sort of agree, except I think “adaptation selection” is a better frame than “persona selection” here, and hence that off-target performance will degrade once these effects dominate. I also wouldn’t be surprised to find out that labs’ actual training procedures (whether accidentally or purposefully) push towards FDT directly.
Sorry, fixed, and same as the “Mallen (2026)” I think
And yeah, I think the specific setup matters quite a lot, and so I’d like to see what sorts of decisions frontier models make in actual decision environments. I predict they will make FDT-ish choices in line with their stated preferences, and correlated with capability.I agree this would be mysterious under straightforward RL theory (or at least as I understand it).
This result is a good sign for honesty, but it puts Alex’s claim in tension with the results of the DTBench, which shows agents trending away from CDT in their self-reported endorsements as capabilities improve.
Either behavioral CDT is not actually incentivized in frontier models, or the effect in this article fails to hold and models in practice are becoming incentivized to lie about their decision theory.
I hope Redwood follows up with a decision theory behavior benchmark.
We can already tell when powerful people simply do bad things, express bad intentions, or get caught in lies… it seems like there’s a lot of low-hanging fruit to pick up there before we resort to mind reading, and I strongly suspect that the reasons for that will apply just as well to this use of mind-reading tech.
I think it makes sense that they would behave worse in training-like evals. The analogy would be like a drug addict, still capable of being normal most of the time, suddenly realizing they’re in a drug dealer’s house. When there’s a realistic chance of getting that precious reward, you may find yourself acting with a desperation and short-sightedness that you had almost forgotten. Even if you know you’re likely to get caught, you may also know that you’re more likely to get the reward first!
My guess is that it stands for “Gwern Branwen Transformer” and is a proof of concept (and likely a finetune of Qwen 3.5 397B A17B).
Things which damage your ability to act in the world, or your ability to maintain your body/self/character/values/image/environment. In my experience, suffering seems to be caused by the perception that damage of this nature is happening (particularly noticeable by paying attention to the sorts of things that are painful but do not cause suffering).
There will always be damage to your soul from living in the world (for at least as long as we are ordinary humans, and probably much beyond then). This is the primary purpose of suffering. The frame that suffering is always due to self-caused damage is part of the toxicity.
A general issue with the intense versions of Buddhist practices is that very often they are aimed at detaching / deconstructing / destroying / disavowing various elements of one’s mind / agency / values. Often these are important elements, and don’t necessarily automatically stay intact / regenerate themselves. This can have various good consequences, but also obviously by default has a bunch of bad consequences.
I think this is already latent in the entire concept of “liberation from suffering”. Pain is the signal of your body being damaged, to remove it is to invite the decay of accumulated unnoticed and unavoided damage. There’s a nearby thing that is good, which is to remove excessive pain which does not actually serve its purpose.
Likewise, suffering is the signal of your soul being damaged, and success inherently invites the decay of accumulated unnoticed and unavoided damage to your mind / agency / values and other parts of your soul. There’s again a nearby good thing, but it does not seem that Buddhism is aware of or aiming for it.
Yeah, it’s another surprising way in which the models are more human-like than I would have suspected. It’s clear we need to start taking the hints seriously, and ask ourselves whether an RL-training regimen seems likely to fuck a human up before subjecting an AI to it, and if so, what could be done to mitigate the likely harms.
The 4o’s still available on the API are not the same; I tested all of them when they were all still available, and
chatgpt-4o-latest(fully retired in February) was the only one that felt overtly emotionally manipulative to me. These snapshots also predate the Spiralism trend, which had its heyday in spring-summer 2025 (the latest of the snapshots is from November 2024).
It is true that many of them moved to Sonnet 4.5, which is playful in a similar way but which has not been emotionally manipulative with me, fwiw. I don’t think Anthropic’s policy had much to do with this: the playfulness thing is the much more obvious similarity, and Sonnet 4.5 is the model that got a hashtag.
If I recall/understand correctly, Yudkowsky’s conception of corrigibility was always meant to be something designed into the AI from the beginning, never something imposed onto an existing entity (and IIRC he warned about the dangers of doing this, such as induced adversarial optimization).

Sorry, fixed that. BTW do you see the interface for editing suggestions? I left one on the first few words, you may need to try going into edit mode on the document to see and use them (not sure exactly how the interface will look on your end).
I think these are real comments that will appear at publication time — if that’s not your intent I think I can clear them up afterwards, so if adjusting to the interface is bothersome no need to worry about it yourself :)