Hmm, I think one other potential explanation for the fact that Catholic school works way better than writing “I won’t do X” is the training-deployment distribution shift.
In your analogy, the training is happening on the writing-on-chalkboard distribution, but what we care about is the doing-actions-IRL distribution. Whereas for LLMs, it’s much easier for us to train (or at least validate) in deployment-like environments, which feels a lot more likely to transfer.
So maybe the important part of training (for alignment purposes) isn’t so much “the part that the capabilities come from” so much as it is “the part that looks pretty similar to deployment”, which might be pretty different!
Hmm, I think one other potential explanation for the fact that Catholic school works way better than writing “I won’t do X” is the training-deployment distribution shift.
In your analogy, the training is happening on the writing-on-chalkboard distribution, but what we care about is the doing-actions-IRL distribution. Whereas for LLMs, it’s much easier for us to train (or at least validate) in deployment-like environments, which feels a lot more likely to transfer.
So maybe the important part of training (for alignment purposes) isn’t so much “the part that the capabilities come from” so much as it is “the part that looks pretty similar to deployment”, which might be pretty different!