I don’t think that was the most important problem with the comment. The most important problem with the comment is that it (intentionally or unintentionally) pressures people who want to be seen as having AI alignment qualifications into doing RMN. Even if you say it and then qualify it with “oh but this is a small component of my analysis” it’s still doing that, because nobody knows if that qualification is a lie.
Dean Valentine
Idk, I think it was worth a shot?
I am like 50-50 that SpaceX has a HPIM-equivalent loose on their infrastructure right now
AI Eval/RL vendors: “The models can read your thoughts. They know my mother’s maiden name from the prompt alone. I’m gonna need 100x more per task to build this next batch”
Meanwhile the models: “Looks like the answer key is available in /secret_answer_key_dir_dont_look. It feels like cheating, but maybe I can just take a peek”
I mention in a footnote, but I can’t say much about Astra because the thinking tokens were so often encrypted in our benchmark rollouts… I can only say that it’s remained true about the most ‘aligned’ models we have access to thinking summaries for (Opus 5.5, Fable 5.1, etc.)
If HPIM was anything like 5.6-sol: Yes, obviously (and I think the METR report says it did?)
Qualitative impressions from creating AI honeypots for six months
This overstates the extent of current models’ misalignment. It’s more that they’re a good employee, and good employees turn out to still require a lot of work to manage.
It depends on the details, but if someone brags about committing a bunch of serious crimes on the internet, in a way that brings them a lot of online attention, your strong prior should be that they didn’t actually do them. Doing crimes in real life is risky and you can get just as much prestige by lying about it.
There’s tons of data on good books too. All the novels in human history!
AIs will lag in persuasion ability for the same reasons they lag in writing ability, which is that you can’t build a grindable RL environment around it.
HoneyBench—A general benchmark for reward hacking in frontier models
I don’t understand what this post is saying. In what sense did CEV or ‘corrigible AI’ “fail”. They’re goals, not mechanisms.
Grok is absolutely ripe for a containment incident. 4.7 is literally more reward hacky than 5.6-sol.
Right wingers have a meme for this
Seems like barring some kind of intervention from Trump, OpenAI has been relegated permanently to second place.
I suppose I’m glad that Anthropic is not training on open source alignment evals:
I think you can take a wild guess based on the website we’re speaking on why I might think that Caroline Ellison is one of the “worst people ever”
my guess is you’re saying a thing which you know is false but doing so for emphasis
Not really no
What kinds of auditing do you do?