By “existing techniques” I do include “fiddle with the RL environments” or “midtrain on some documents about the intended behavior” or “run a prompted monitor over traffic in prod” or etc.
When I say “understand and shape generalization” I think the central example is (i) adjust parameters of the training process that don’t affect competitiveness, (ii) build some understanding of how those parameters affect generalization so that you can adjust them in a helpful way. I think you’re saying that’s not a “technique;” I don’t care much about the semantics. I’m sure I put a higher probability on those changes helping than you do but it doesn’t seem worth arguing about here.
Yep, makes sense. Don’t need to get into it here, but just to avoid a misunderstanding, I was not making a semantic point, I was making a point about what the majority of the field is working on.
I agree that if you count myopically fiddling with the RL environments as an example of the first one, then yeah, almost all alignment research is that, because that plus pretraining is what most of all ML research is. I think the case for “myopically fiddling with the RL environments scales to aligning superintelligence” is very weak, but I agree it’s not impossible!
By “existing techniques” I do include “fiddle with the RL environments” or “midtrain on some documents about the intended behavior” or “run a prompted monitor over traffic in prod” or etc.
When I say “understand and shape generalization” I think the central example is (i) adjust parameters of the training process that don’t affect competitiveness, (ii) build some understanding of how those parameters affect generalization so that you can adjust them in a helpful way. I think you’re saying that’s not a “technique;” I don’t care much about the semantics. I’m sure I put a higher probability on those changes helping than you do but it doesn’t seem worth arguing about here.
Yep, makes sense. Don’t need to get into it here, but just to avoid a misunderstanding, I was not making a semantic point, I was making a point about what the majority of the field is working on.
I agree that if you count myopically fiddling with the RL environments as an example of the first one, then yeah, almost all alignment research is that, because that plus pretraining is what most of all ML research is. I think the case for “myopically fiddling with the RL environments scales to aligning superintelligence” is very weak, but I agree it’s not impossible!