(Moving back up to the top of your bullets) It seems to me like there are a bunch of cursed things about trying to know if you are making progress on control, such as: - A scheming AI may come up with strategies very dissimilar to your red-team - The environment may look super different with weird changes in scaffolding, coding language, etc. - The models themselves might be very different and this might change things (maybe continual learning could make collusion easier, etc.) (This is along with smaller, more basic disanalogies like maybe adding coding backdoors is a bit different than trying to have a rogue internal deployment or something.)
A lot of what excites me about control is that you can, at crunch time, run those experiments in the actual setting you care about. This means that a lot of our work now should be focused on prepping to run those experiments.
A lot of what excites me about control is that you can, at crunch time, run those experiments in the actual setting you care about. This means that a lot of our work now should be focused on prepping to run those experiments.
If it’s crunch time, can’t a misaligned AI just sandbag (or do this kind of thing) on your control experiments?
Sandbagging is maybe really hard though! See here. In general, the possibility of running control evals at crunch time is a lot of why I like control.