I’d like to see more engagement between Gwern’s position that catastrophic forgetting is solved and the Dwarkesh position that it remains one of the major barriers to useful AI and might require a breakthrough.
Matthew Khoriaty
I note that “correction” also refers to stocks going down. That’s not a great connotation to have, especially in the unfortunate futures where it becomes political.
This is surprising to me! I would have expected negligible differences.
What does this mean for how we should do research in practice? Should we avoid OpenRouter entirely?
At the least, we should have a norm of recording which provider was used.
There are sometimes subtle but significant drops in Claude and Codex model quality. This is something that could impact research involving those models, and it would be good if Anthropic and OpenAI were to have a log of what happened, why, and when the model is fully operational. I agree that first parties are generally fairly trustworthy, but I would still say they don’t reach the ideal for research. https://marginlab.ai/trackers/claude-code-historical-performance/
Though I admit I did not check all of Claude’s claims in the artifacts, I did spot check them and verified the specific claims (though not aggregate numbers) that made it to the post.
You’re right, sorry.
When I wrote “locally, yourself” I was thinking of either having or renting GPUs to run your models. I should have written “on hardware you control.”
For my use case, I sadly care a lot about “interactive debugging”. I’m iterating on prompts, control protocols, and models, so if I was running on rented GPUs they would be idling for a long time between experiments.
I don’t know too much about the details of running models on rented GPUs, but after looking into OpenRouter for the post, I am inclined to learn. Thanks for sharing!
Congratulations! Two years later, and we have our first major warning shot in Claude Mythos, which scared policymakers with its hacking abilities. You get bragging rights.
Do people really believe this?
“emergent misalignment” literature suggests that “good according to the human value system” and “evil according to the human value system” are salient enough vectors that pushing on them in some ways can “drag along” all of the rest of their content
I would expect that if you finetune on malicious data (eg insecure code) in non-English languages, you get a version of Emergent Misalignment for that specific culture. Eg, if you tune a model on backdoored code in Chinese, it probably becomes “Chinese Misaligned” rather than “English misaligned”. I don’t know enough about other cultures to say what this would look like.
I’d love to see someone do this experiment! It would demonstrate that Emergent Misalignment does not reveal a “human value system” in the model which it is easy and simple to move along.
I would say that “RLHF makes AI’s aligned in many ways, and Emergent Misalignment results in AIs trained on bad code to also be racist” is very weak evidence in favor of the Natural Abstraction Hypothesis since it can also arise from human statistical patters. The kind of troll who gives someone backdoored code on the internet is probably also the kind who says that women shouldn’t be computer scientists. (Some prompting-only experiments I ran confirm this).
I would be impressed if Claude invents a new weird EA Cause area which EAs and philosophers give serious thought before deciding that Claude is right and discovered a new way to be good.
Part of the problem with Personas is that it is blocking us from testing/evaluating the Natural Abstraction Hypothesis because the personas, mimicking humans, have their own bundles of beliefs and abstractions. Beyond that, humans have contingent statistical patterns in their values and beliefs. Studying the Natural Abstraction Hypothesis at frontier LLM scale requires we find a way to suppress/avoid the personas who already believe and use the abstractions in question.
As a piece of evidence that Claude isn’t going beyond mimicry to some transcendent fundamental goodness, when asked for “fix everything easily switch” policies, all of the policies it suggests are ones that are already popular with rationalist/technocrats. That isn’t to say that they are bad, but if Claude was really inferring some transcendent goodness, it probably would have suggested something that (eg) Zvi hadn’t already heard about, thought about, and liked. All the policies it suggests seem reasonable to me, and it is tempting to call this “aligned”, but what happens when ASI Claude is put in charge of the world and needs to decide what policies to put in place after it already implements Zvi’s pet issues? If there was an underlying goodness that underlies the personas, that would be amazing. The question is: how do we verify that when all of our evaluations will just evaluate a persona?
I’m not planning to pursue Personaless Alignment at the moment, since I am about to do an AI Control research fellowship with Redwood.
I don’t think the right path is to try starting from Zero pretraining. Pretraining is and will continue to be a huge benefit to capabilities, and aligning the most capable models is the goal. I’m saying that we should separate pretraining for capabilities from the alignment that can come from eliciting aligned personas from pretraining. I’m not sure exactly how to do this, but it would likely involve starting with carefully filtered large models.
I chose Beowulf because it is more alien and removed from the present day than the Bible. The Bible has had significant influence on our present-day values and culture, while Beowulf is still a human artifact containing human values, but extrapolating our current values from Beowulf would be very difficult. According to Claude, “Beowulf has nothing that reads as advocacy for animal rights or welfare, and the concept itself is anachronistic by roughly a millennium.”
Your second point isn’t really a nitpick. Rather, it is the alignment problem itself. Nobody really knows how to solve it, but techniques such as inverse reinforcement learning or Building AIs that do human-like philosophy don’t run into the persona problem the same way as techniques like RLHF.
Thank you for the clarification on Talkie behavior.
“We need Personaless Alignment” and “We don’t have time for Personaless Alignment“ can be true at the same time.
I agree with you that we look to be on track for getting ASI with short term techniques. I agree that because of that, persona alignment is a good thing for people to work on.
I’m drawing attention to how neglected Personaless Alignment seems to be and why it would be valuable.
Hopefully researchers can figure out a way to study this quickly.
I’m unsure what you mean. I consider the pretraining data to be the issue, not the particular language or output that the AI generates. If you pretrain on moral humans then convert your model to a neuralese model, then the model still has representations of good human personas.
I do agree that Personaless Alignment is a dual-use research direction in that it may benefit capabilities (you’re doing crazy things to a misaligned AI to try to get it to be good), but I consider its capability risks likely smaller than the benefit to alignment. Once the details of the experimental approaches are pinned down, it would be easier to say.
Agreed. Part of the problem is that it is hard to avoid alignment via personas. The LLMs already understand human goodness and you can elicit it with a few low-rank matrices. Personas are blocking researchers from doing real alignment work.
Thanks for the empirical check! Nodding along to whatever someone says in a positive light is pretty easy and doesn’t reflect much about the values of the model. And, as you say, its a small model.
I believe Personaless Alignment would be a good research direction if it could be made tractable. I’m just struggling how it could be made tractable. I’d appreciate your thoughts on why you think it won’t be a good research direction.
Yes, you described my argument correctly. I’m long-term bearish on trying to extrapolate the values and reasoning processes of good people in the training set out-of-distribution. Putting the goodness into the optimizer (or other targeting system) seems more promising. I don’t think that Claude is mimicing Atticus Finch’s specific actions, I think Claude learned about Atticus Finch, gained an understanding of what guided his actions, then was trained to act in that way. This still requires extrapolation.
Imagine if your training set is a group of 30 pre-schoolers. And your alignment technique is to try extrapolating the values of the most kind and conscientious of the pre-schoolers (according to the pre-schoolers themselves) to adult level. I don’t think that the actions of pre-schoolers have sufficient range for this extrapolation to work properly.
As with the other aspects of Personaless Alignment, our disagreement would be hard to test empirically.
The kind of thing which would impress me is if you can train a model on Beowulf-type values, scale the model, and the model correctly extrapolates “When I extrapolate the values of the wise King Hrothgar, I infer that animal welfare is an important thing to care about.” Aligning BeowulfLLM to modern values without imposing our values on it directly is a smaller, easier problem than aligning superintelligence to ‘true’ human values.
When you’re superintelligent, role models only cover a small fraction of the action space, and you have to genuinely be better/more moral than all of them to be aligned.
Tired of scrolling up to the top of Claude Code or Codex’s responses? Solution: Command-f the character that starts each prompt: “❯”. You can copy it right from your terminal, or you can make a keyboard shortcut to type it.
In that case it isn’t a scientist AI but rather a knowledge indexing AI. It can pull together causal concepts that humans have written about, but without RL-style interaction with the environment, it’s not a scientist.
Causes preceding effects is generally true but it’s not going to help you solve much.
I may have put extra focus on the upsides in an attempt to be more even-handed. The things that I say about the upsides are true and I do believe them but they aren’t really the meat of the paper.
I’m realizing that sharing writing and thoughts is largely appreciated. The EA forum has explicit events to try getting people to publish their drafts. In that spirit I decided to post this here, and I’m glad that some people are finding it valuable.
What process results in the highest quality CEV labels? That would be good to know, even if that process is expensive. Consider writing it up.
Thanks for writing this!
I graduated college June 2026, so I have some insight into what’s going on in universities. dvd’s has a better POV because he gathered information more systematically and I self-selected CS majors AI-risk-concerned students, but I figure I can share some of what I saw.
I can confirm that #2, #4, and #8 are common among my classmates. I haven’t seen the others.
I’m surprised by #1. People have been using AI for more and more difficult things over time. You can’t productively have ChatGPT 3.5 help you with anything difficult, and until very recently AIs couldn’t consistently read text in images.
#5 is a good point. Unfortunately, not all AI leaders see the issues, and they all think that it would be best for them selves to win, even though they are putting everyone at risk. I can understand why students would believe #5.
I’m guilty of #9. Even though I’m young and clueless and lack a lot of skills because of AI assistance, I feel worried for those younger than me who have relied on AI for an even longer portion of their lives. Coding by hand 3 years ago and coding with light AI assistance 1 year ago taught me a lot, and I’m prouder of the things I made without AI.
I have hoping since 2024 that my university group, Northwestern University AI Safety and Governance Group, would find it easier to recruit as AI risks became more obvious, and I had hoped that younger students would have an easier time understanding the power of AI (if they are “AI-native”, which you suggest they are not). In the following two years, I have been confused when DeepSeek-R1, o3, Claude Code, and Mythos came out, and I saw no influx of interested new members. Thank you for casting light on what has been going on with the students I have been trying to recruit.