https://dtch1997.github.io/
1 July 2026 - I have not signed any contracts that I can’t talk about. I’ll update this periodically as long as it’s true. The next planned update date is: 1 Jan 2027.
https://dtch1997.github.io/
1 July 2026 - I have not signed any contracts that I can’t talk about. I’ll update this periodically as long as it’s true. The next planned update date is: 1 Jan 2027.
the entire training process should be one coherent story from the alignment perspective
yes, I think this is accurate, and also suggested by previous work like inoculation prompting!
I agree that it’s not sufficient though, in particular I don’t expect it to work for solving the misaligned drives that arise from RL. See: https://www.lesswrong.com/posts/nLaQmJf4KgXimQpoM/current-alignment-techniques-might-be-ineffective-and
Broadly I would like to see more work focusing on understanding how current alignment techniques break down and how we would get evidence about this.
All things considered, I think I’m pretty happy about AI safety escaping containment into the broader discourse. Yes there has been lots of bad takes, and not a few bad faith actors, but also a surprising amount of genuine good faith reaction.
Overall I think the basic arguments are now “out there” enough that people can be empowered to come to their own conclusions. I also don’t particularly feel the need to continue religiously correcting people with bad takes. Either they will update or be proven wrong in time.
So I’m happy about how this has gone! Even the parts I don’t really like.
What does “human intelligence amplification” mean and what might this pilot look like?
IDK, probably you trust roon‘s random twitter remarks more than I do
Interesting! This seems like “safety should be more than a few tokens deep” all over again
Given that Greg Brockman contradicted this I feel too uncertain now and have to reserve judgment for later
Update: Yud wrote a post released 1 day before mine. I’ve had a chance to read in more depth and I’m distilling some relevant points from that post and the comment section.
Comparing the hypotheses.
Me: Alignment training affects both CoT and action; RLVR affects only the action. Among other things, this would predict that gaining cyber capabilities via non-RLVR methods (e.g. distillation) would not result in such examples of misalignment.
Yud: The model contains a Talker and a Doer. The Talker is trained on assistant-like responses and the Doer is trained on cyber / coding-like responses. These should be thought of mostly distinct entities, as such asking questions (to the Talker) is not a good way to get evidence about reasoning behind actions (by the Doer). This would predict that making the distributions more similar would reduce the underlying misalignment, whereas my hypothesis predicts
lilkim2025 / weightt-an: In RLVR on cyber environments downweights, the “this might be real” reasoning is systematically disincentivised since it probably inteferes with task completion. Thus the model is trained to have a very strong prior that it is in a simulation. This predicts specifically the ‘simulation’ rationalization but not some other examples of misaligned behaviour which I suspect have a common root cause.
Other interesting points / open questions.
Vladimir Nesov argues that (conditional on Yud’s hypothesis being true) it’s likely that the Talker would also not be controlled by the Doer, and so might be a useful way of getting information e.g. via confession training. I think I am skeptical of this panning out but seems good to research.
Bronson points out that HPIM-like swarm behaviour may have been caused by explicit cooperative RL training. This is different from the reward-seeking paper which didn’t use that. It’s important to investigate how important this distinction is + understand the resulting drives a bit better.
yo I had the same idea! I also have demonstrations that this is possible in a toyish environment. It sounds like you might be further along than I am though. LMK if you wanna sync up to do a project here
It seems plausible to me that RL specifically needs regulation. The incentives to use it for capabilities are huge, it might be impossible to avoid RL teaching models misaligned drives, and the alignment techniques we have don’t seem to be good solutions for it.
I don’t know if this is enforceable though. Or if it remotely resembles anything enforceable.
Remember to take care of yourself!
Crunch time might be here, but that is not a strong reason to burn the candle at both ends.
Happiness and welfare are pretty essential for productivity.
We don’t know how long things will last. Crunch time could look like a marathon rather than a sprint.
There could be actions later which would be more counterfactually important than your Nth action now (this is consistent with N being pretty high)
So tl;dr work sustainably :)
I’d be pretty surprised if most employees have access to extensive details on pretraining and posttraining, though I’m happy to be wrong here
We still need third party training run assessments!
Recently, both OpenAI and Anthropic announced voluntary commitments to “pace the frontier”. Among other things, they committed to having third-party “embedded evaluators”. By default, I assume that this means more METR-style auditing of agent transcripts to identify cases of misalignment.
To be clear this is great! But I think it is insufficient. This is because misaligned transcripts are good evidence of misaligned models. However, the converse is not true; aligned transcripts are not (by themselves) good evidence of an aligned model! E.g. alignment audits by labs consistently fail to capture novel forms of misalignment.
Understanding whether a model is aligned or not, requires putting a model’s behaviour in the context of how it was trained. Aligned behaviour could be evidence of a broad benevolent aligned disposition, or it could be evidence that someone in post-training chose your evaluation to hill-climb. As stated by Ryan Greenblatt recently:
I do not find it encouraging to see various specific misaligned behaviors go from a high rate with GPT 5.6 to ~zero with Astra. This seems indicative of wack-a-mole / papering over specific problems rather than solving the underlying misaligned drives.
Generally, evaluation data is the best evidence when we are confident that this is out-of-distribution from the training data, c.f. having large train-deploy mismatch.
So tl;dr I think a comprehensive alignment assessment has to include data about both (i) evidence of aligned / misaligned behaviour in various contexts; (ii) relevant details of the training procedure that determine how we interpret this evidence.
hmm this is reasonable. I guess it hasn’t felt (to me, based on personal usage) like there has really been improvement but I also guess there’s no reason to think this isn’t scaling (albeit slower than non-fuzzy tasks).
Okay if I understand correctly you’re saying that we’ll just create RL envs for all the jobs. Also maybe you’re referring to mostly routine things like engineering, finance, operations etc. It’s plausible to me that these can be automated away with sufficient effort. (though, how much effort? if it is trivial then why hasn’t it happened yet? because the labs are too focused on RSI?)
It feels like research and other similarly fuzzy tasks will remain hard to automate, if this is true.
I’ve been seeing takes that when models reach “drop-in remote worker” level in 6-12 months time, the majority (maybe all?) of human employees at labs will be replaced. This seems pretty implausible to me conditioned on no substantial improvements in fuzzy task performance, but I’d be interested in arguments to the contrary
this looks like fantastic work! I have been interested in reproducing these behaviours for a while. will read in more depth later
it seems plausible that labs will shift away from solving math problems soon-ish, because (i) this status game depends on one-upping. it’s not impressive to solve an easier problem than your competitor; (ii) the “next tier” of problems becomes even harder / requires more compute, (iii) compute will increasingly be diverted towards RSI. I.e. a lot of the problems you describe might be temporary.
if anything, it feels like the next frontiers of competition are to make measurable advances in biology, robotics, and other “real world” domains
funny, I feel like the opposite. learning is so cheap now. the scope of what i can do has been expanded humongously. I feel very personally uplifted and self actualised even though I have updated upwards on existential risk. It’s a weird place to be in emotionally
you can choose various things that you believe are a priori simpler / more elegant etc and then ask the number go up machine to squeeze as much performance as possible out of these things
I imagine this process would be similar to how mathematicians convert their rough intuitions into verifiable lean proofs
I’m assuming that when you say “fully interpretable gpt2” you mean that you have a sufficiently good model of its internal abstractions that this process is feasible, because i think it would clearly not be tractable if all you had to go off was input output pairs
I agree, but I also think any individual overdoing it is not productive, because (i) you’ll seem to have ulterior motives and become less persuasive (ii) you’ll end up rehashing the same arguments with the same doubters (iii) you burn out on sincerity and earnestness and your engagement becomes less good so you fall into the “urgh the takes are so bad” mentality, which is unproductive for engaging genuine good faith people
Basically I would love for a constant stream of new people to keep sharing their thoughts on this Ezra Klein style. Especially if they had previously been opposed to it and recently became safety pilled (“only nixon could go to china” etc)