For example, in domains where people have the choice about whether to use a control or alignment lens, I think the control lens is better: I think the most fruitful interp work has been done by evaluating things control-style, evaluating techniques against model organisms.
I think I agree for interp, but don’t think I have this intuition in general. More generally, I am a bit confused on how this point fits in. If a technique is better fit for a control lens over an alignment lens, this seems to say more about the technique than the lens, to me. I am making the claim that we should focus our attention more toward the techniques that are better fit for an alignment lens. To the extent your point reduces down to something like “the control lens is more straightforward and easier to work with,” I agree.
Similarly, I think unsupervised elicitation work is best done using adversarially crafted initializations and datasets.
Clarification question: the idea is that when you are trying to elicit behavior we don’t have labels for, we should try to make sure this generalizes by testing it in situations where it is most likely to fail?
If this is right, this still seems alignment-y to me and not control-y (at risk of arguing over definitions). It just seems like you are using an adversarial / robustness lens, which seems valid for both alignment and control. I think I agree that this is good, though.
How cursed is your methodology for knowing whether you are making progress at all or not. I think the situation is better in control-land than in alignment-land.
(Moving back up to the top of your bullets) It seems to me like there are a bunch of cursed things about trying to know if you are making progress on control, such as: - A scheming AI may come up with strategies very dissimilar to your red-team - The environment may look super different with weird changes in scaffolding, coding language, etc. - The models themselves might be very different and this might change things (maybe continual learning could make collusion easier, etc.) (This is along with smaller, more basic disanalogies like maybe adding coding backdoors is a bit different than trying to have a rogue internal deployment or something.)
It seems to me that for both alignment and control, empirical experiments act largely as an excuse to reason in more detail about what reasonable things to do appear to be, and perhaps more value comes from this than numerical tests that indicate any progress (although I’ll avoid going far down this tangent).
How good are the marginal ideas and experiments to be explored. Here I think it’s a bit more ambiguous, and depends on what are the skill and information profile of the people “in the field”.
I think I basically agree with everything you are saying here.
(Moving back up to the top of your bullets) It seems to me like there are a bunch of cursed things about trying to know if you are making progress on control, such as: - A scheming AI may come up with strategies very dissimilar to your red-team - The environment may look super different with weird changes in scaffolding, coding language, etc. - The models themselves might be very different and this might change things (maybe continual learning could make collusion easier, etc.) (This is along with smaller, more basic disanalogies like maybe adding coding backdoors is a bit different than trying to have a rogue internal deployment or something.)
A lot of what excites me about control is that you can, at crunch time, run those experiments in the actual setting you care about. This means that a lot of our work now should be focused on prepping to run those experiments.
Thank you for your comment!
I think I agree for interp, but don’t think I have this intuition in general. More generally, I am a bit confused on how this point fits in. If a technique is better fit for a control lens over an alignment lens, this seems to say more about the technique than the lens, to me. I am making the claim that we should focus our attention more toward the techniques that are better fit for an alignment lens. To the extent your point reduces down to something like “the control lens is more straightforward and easier to work with,” I agree.
Clarification question: the idea is that when you are trying to elicit behavior we don’t have labels for, we should try to make sure this generalizes by testing it in situations where it is most likely to fail?
If this is right, this still seems alignment-y to me and not control-y (at risk of arguing over definitions). It just seems like you are using an adversarial / robustness lens, which seems valid for both alignment and control. I think I agree that this is good, though.
(Moving back up to the top of your bullets) It seems to me like there are a bunch of cursed things about trying to know if you are making progress on control, such as:
- A scheming AI may come up with strategies very dissimilar to your red-team
- The environment may look super different with weird changes in scaffolding, coding language, etc.
- The models themselves might be very different and this might change things (maybe continual learning could make collusion easier, etc.)
(This is along with smaller, more basic disanalogies like maybe adding coding backdoors is a bit different than trying to have a rogue internal deployment or something.)
It seems to me that for both alignment and control, empirical experiments act largely as an excuse to reason in more detail about what reasonable things to do appear to be, and perhaps more value comes from this than numerical tests that indicate any progress (although I’ll avoid going far down this tangent).
I think I basically agree with everything you are saying here.
A lot of what excites me about control is that you can, at crunch time, run those experiments in the actual setting you care about. This means that a lot of our work now should be focused on prepping to run those experiments.
If it’s crunch time, can’t a misaligned AI just sandbag (or do this kind of thing) on your control experiments?
Sandbagging is maybe really hard though! See here. In general, the possibility of running control evals at crunch time is a lot of why I like control.