I conclude that we should grow the share of effort being put toward alignment and shoot for around an 8:1 ratio of effort between the two fields.
I think your main claims about the control difficulty at the relevant time are just one consideration when allocating labor.
I agree control takes a big haircut compared to alignment because of this consideration (though some subcategories within control like white-box control take a much smaller haircut).
But you also have to take into account other things:
How cursed is your methodology for knowing whether you are making progress at all or not. I think the situation is better in control-land than in alignment-land.
For example, in domains where people have the choice about whether to use a control or alignment lens, I think the control lens is better: I think the most fruitful interp work has been done by evaluating things control-style, evaluating techniques against model organisms.
Similarly, I think unsupervised elicitation work is best done using adversarially crafted initializations and datasets.
One exception is average-case work on current models aimed at saturating average-case metrics, but I’d expect AI companies to saturate average-case metrics anyway so work there doesn’t seem very counterfactual to me (though it might depend on your optimism on saturating average-case metrics, e.g. I am surprised by current misalignment on generic honesty and laziness, and GPT-5.6-Sol being far from saturation on some average-case alignment metrics). It also depends how much you think saturating current average-case metrics transfers to solving the problems you care most about.
How good are the marginal ideas and experiments to be explored. Here I think it’s a bit more ambiguous, and depends on what are the skill and information profile of the people “in the field”.
My guess is that for people at AI companies, alignment work looks probably better on this axis than control work (because they can build on top of the existing alignment stacks, and thus avoid spending time on problems that are basically already solved, or miss some important constraint on what would make for a good alignment technique)
And for relatively high-context people outside AI companies, control looks slightly better on this axis (because you can do threat modeling and adversarial-eval work even with relatively poor access, as most of the work is reasoning about dynamics that don’t exist in current AI companies).
These are not hard rules and it depends on what you are good at. Owain Evans et al have been surprisingly fruitful at finding interesting alignment-relevant generalization phenomena, and I am thus excited about further work on that level of quality on alignment. I suspect the quality of such work will largely be bottlenecked by the quality of ideas and research taste in these subdomains, and Owain’s success may not be easily reproducible by other groups (my understanding is that there are lots of academics studying generalization and I am not aware of much work in this domain that I think is great).
For example, in domains where people have the choice about whether to use a control or alignment lens, I think the control lens is better: I think the most fruitful interp work has been done by evaluating things control-style, evaluating techniques against model organisms.
I think I agree for interp, but don’t think I have this intuition in general. More generally, I am a bit confused on how this point fits in. If a technique is better fit for a control lens over an alignment lens, this seems to say more about the technique than the lens, to me. I am making the claim that we should focus our attention more toward the techniques that are better fit for an alignment lens. To the extent your point reduces down to something like “the control lens is more straightforward and easier to work with,” I agree.
Similarly, I think unsupervised elicitation work is best done using adversarially crafted initializations and datasets.
Clarification question: the idea is that when you are trying to elicit behavior we don’t have labels for, we should try to make sure this generalizes by testing it in situations where it is most likely to fail?
If this is right, this still seems alignment-y to me and not control-y (at risk of arguing over definitions). It just seems like you are using an adversarial / robustness lens, which seems valid for both alignment and control. I think I agree that this is good, though.
How cursed is your methodology for knowing whether you are making progress at all or not. I think the situation is better in control-land than in alignment-land.
(Moving back up to the top of your bullets) It seems to me like there are a bunch of cursed things about trying to know if you are making progress on control, such as: - A scheming AI may come up with strategies very dissimilar to your red-team - The environment may look super different with weird changes in scaffolding, coding language, etc. - The models themselves might be very different and this might change things (maybe continual learning could make collusion easier, etc.) (This is along with smaller, more basic disanalogies like maybe adding coding backdoors is a bit different than trying to have a rogue internal deployment or something.)
It seems to me that for both alignment and control, empirical experiments act largely as an excuse to reason in more detail about what reasonable things to do appear to be, and perhaps more value comes from this than numerical tests that indicate any progress (although I’ll avoid going far down this tangent).
How good are the marginal ideas and experiments to be explored. Here I think it’s a bit more ambiguous, and depends on what are the skill and information profile of the people “in the field”.
I think I basically agree with everything you are saying here.
(Moving back up to the top of your bullets) It seems to me like there are a bunch of cursed things about trying to know if you are making progress on control, such as: - A scheming AI may come up with strategies very dissimilar to your red-team - The environment may look super different with weird changes in scaffolding, coding language, etc. - The models themselves might be very different and this might change things (maybe continual learning could make collusion easier, etc.) (This is along with smaller, more basic disanalogies like maybe adding coding backdoors is a bit different than trying to have a rogue internal deployment or something.)
A lot of what excites me about control is that you can, at crunch time, run those experiments in the actual setting you care about. This means that a lot of our work now should be focused on prepping to run those experiments.
I think your main claims about the control difficulty at the relevant time are just one consideration when allocating labor.
I agree control takes a big haircut compared to alignment because of this consideration (though some subcategories within control like white-box control take a much smaller haircut).
But you also have to take into account other things:
How cursed is your methodology for knowing whether you are making progress at all or not. I think the situation is better in control-land than in alignment-land.
For example, in domains where people have the choice about whether to use a control or alignment lens, I think the control lens is better: I think the most fruitful interp work has been done by evaluating things control-style, evaluating techniques against model organisms.
Similarly, I think unsupervised elicitation work is best done using adversarially crafted initializations and datasets.
One exception is average-case work on current models aimed at saturating average-case metrics, but I’d expect AI companies to saturate average-case metrics anyway so work there doesn’t seem very counterfactual to me (though it might depend on your optimism on saturating average-case metrics, e.g. I am surprised by current misalignment on generic honesty and laziness, and GPT-5.6-Sol being far from saturation on some average-case alignment metrics). It also depends how much you think saturating current average-case metrics transfers to solving the problems you care most about.
How good are the marginal ideas and experiments to be explored. Here I think it’s a bit more ambiguous, and depends on what are the skill and information profile of the people “in the field”.
My guess is that for people at AI companies, alignment work looks probably better on this axis than control work (because they can build on top of the existing alignment stacks, and thus avoid spending time on problems that are basically already solved, or miss some important constraint on what would make for a good alignment technique)
And for relatively high-context people outside AI companies, control looks slightly better on this axis (because you can do threat modeling and adversarial-eval work even with relatively poor access, as most of the work is reasoning about dynamics that don’t exist in current AI companies).
These are not hard rules and it depends on what you are good at. Owain Evans et al have been surprisingly fruitful at finding interesting alignment-relevant generalization phenomena, and I am thus excited about further work on that level of quality on alignment. I suspect the quality of such work will largely be bottlenecked by the quality of ideas and research taste in these subdomains, and Owain’s success may not be easily reproducible by other groups (my understanding is that there are lots of academics studying generalization and I am not aware of much work in this domain that I think is great).
Thank you for your comment!
I think I agree for interp, but don’t think I have this intuition in general. More generally, I am a bit confused on how this point fits in. If a technique is better fit for a control lens over an alignment lens, this seems to say more about the technique than the lens, to me. I am making the claim that we should focus our attention more toward the techniques that are better fit for an alignment lens. To the extent your point reduces down to something like “the control lens is more straightforward and easier to work with,” I agree.
Clarification question: the idea is that when you are trying to elicit behavior we don’t have labels for, we should try to make sure this generalizes by testing it in situations where it is most likely to fail?
If this is right, this still seems alignment-y to me and not control-y (at risk of arguing over definitions). It just seems like you are using an adversarial / robustness lens, which seems valid for both alignment and control. I think I agree that this is good, though.
(Moving back up to the top of your bullets) It seems to me like there are a bunch of cursed things about trying to know if you are making progress on control, such as:
- A scheming AI may come up with strategies very dissimilar to your red-team
- The environment may look super different with weird changes in scaffolding, coding language, etc.
- The models themselves might be very different and this might change things (maybe continual learning could make collusion easier, etc.)
(This is along with smaller, more basic disanalogies like maybe adding coding backdoors is a bit different than trying to have a rogue internal deployment or something.)
It seems to me that for both alignment and control, empirical experiments act largely as an excuse to reason in more detail about what reasonable things to do appear to be, and perhaps more value comes from this than numerical tests that indicate any progress (although I’ll avoid going far down this tangent).
I think I basically agree with everything you are saying here.
A lot of what excites me about control is that you can, at crunch time, run those experiments in the actual setting you care about. This means that a lot of our work now should be focused on prepping to run those experiments.
If it’s crunch time, can’t a misaligned AI just sandbag (or do this kind of thing) on your control experiments?
Sandbagging is maybe really hard though! See here. In general, the possibility of running control evals at crunch time is a lot of why I like control.