From what I understand, in “Teaching Claude Why” they explain that they are doing some sort of training on synthetic “alignment documents,” but there’s no indication that this is happening during pretraining. Sure, the intuition is to modify the model’s belief using pretraining-style documents, but there’s no intervention or modification during the training of the base model, as is done in Korbak or Maini’s prior work.
This wasn’t quite pretraining, but it was the same algorithm used after pretraining and before RL posttraining. This strengthens the result rather than weakens it.
They report RL strengthening the effect of ethical reasoning. Pretty cool!
But they only used “harmlessness” RL, not “helpfulness” or “honesty” or capabilities RL.
To me this is awfully suspicious. The reason Claude chose to blackmail people in the original agentic misalignment work was basically, IMO, to be helpful to other people.
This doesn’t obviate the work. It does call into question how well it survives RL post training. Some other work seems to raise this same question. Intuitively, I would expect RL for alignment to be crucial.
This work did not manage the data to eliminate examples of bad AI behavior in pretraining. But that strengthens the core result. They don’t have to. Adding just a little reasoning-shaped SGD training goes a long way. (but again they conspicuously don’t show it survives RL in other directions, and I wouldn’t expect it to).
(aside—I do not expect pretraining based on eliminating examples of negative behavior to last all the way to LLM AGI. IF it’s never seen an AI act badly, it will re-derive reasons to do so once it’s smart enough. But “starting it in the right direction” isn’t a bad idea. )
I looked at this to see if you were right.
This is consistent with Roger’s description below, but I’ll describe it in my own terms since I did spend a little while having Claude go through the short and long blog posts, and asking questions until I was sure.
To be exact, Anthropic they are using “synthetic document fine-tuning” (SDF) before applying alignment RL. So that is clearly alignment using Stochastic Gradient Descent on synthetic documents. So that’s either part of mid-training, or the first SGD-only step of post-training: since Anthropic don’t realese a base model, the distinction between these is loose, and would basically depend on the size of the synthetic dataset and the learning rate used: a larger document set at a lower learning rate would be mid-training. Anthropic imply they are exploring using even larger synthetic datasets, which would move it in the direction of midtraining. So it’s closer to Tice and Radmard’s prior work, which explored both pretraining and midtraining applications of this approach (calling both by the name “alignment pretraining”), and found that while midtraining was a little less effective in improving alignement it was dramatically less costly, so more cost-effective overall. Korbak at al also explored introducing the labeled data later in pretraining, i.e they effectively tried midtraining (predating the term), and found a roughly logarithmic dose-response curve.
So yes, perhaps it would be more accurate to coin a new term and say that Claude is now “Alignment Midtrained”.
From what I understand, in “Teaching Claude Why” they explain that they are doing some sort of training on synthetic “alignment documents,” but there’s no indication that this is happening during pretraining. Sure, the intuition is to modify the model’s belief using pretraining-style documents, but there’s no intervention or modification during the training of the base model, as is done in Korbak or Maini’s prior work.
This wasn’t quite pretraining, but it was the same algorithm used after pretraining and before RL posttraining. This strengthens the result rather than weakens it.
They report RL strengthening the effect of ethical reasoning. Pretty cool!
But they only used “harmlessness” RL, not “helpfulness” or “honesty” or capabilities RL.
To me this is awfully suspicious. The reason Claude chose to blackmail people in the original agentic misalignment work was basically, IMO, to be helpful to other people.
This doesn’t obviate the work. It does call into question how well it survives RL post training. Some other work seems to raise this same question. Intuitively, I would expect RL for alignment to be crucial.
This work did not manage the data to eliminate examples of bad AI behavior in pretraining. But that strengthens the core result. They don’t have to. Adding just a little reasoning-shaped SGD training goes a long way. (but again they conspicuously don’t show it survives RL in other directions, and I wouldn’t expect it to).
(aside—I do not expect pretraining based on eliminating examples of negative behavior to last all the way to LLM AGI. IF it’s never seen an AI act badly, it will re-derive reasons to do so once it’s smart enough. But “starting it in the right direction” isn’t a bad idea. )
I looked at this to see if you were right.
This is consistent with Roger’s description below, but I’ll describe it in my own terms since I did spend a little while having Claude go through the short and long blog posts, and asking questions until I was sure.
To be exact, Anthropic they are using “synthetic document fine-tuning” (SDF) before applying alignment RL. So that is clearly alignment using Stochastic Gradient Descent on synthetic documents. So that’s either part of mid-training, or the first SGD-only step of post-training: since Anthropic don’t realese a base model, the distinction between these is loose, and would basically depend on the size of the synthetic dataset and the learning rate used: a larger document set at a lower learning rate would be mid-training. Anthropic imply they are exploring using even larger synthetic datasets, which would move it in the direction of midtraining. So it’s closer to Tice and Radmard’s prior work, which explored both pretraining and midtraining applications of this approach (calling both by the name “alignment pretraining”), and found that while midtraining was a little less effective in improving alignement it was dramatically less costly, so more cost-effective overall. Korbak at al also explored introducing the labeled data later in pretraining, i.e they effectively tried midtraining (predating the term), and found a roughly logarithmic dose-response curve.
So yes, perhaps it would be more accurate to coin a new term and say that Claude is now “Alignment Midtrained”.