If it’s crunch time, can’t a misaligned AI just sandbag (or do this kind of thing) on your control experiments?
Sandbagging is maybe really hard though! See here. In general, the possibility of running control evals at crunch time is a lot of why I like control.
If it’s crunch time, can’t a misaligned AI just sandbag (or do this kind of thing) on your control experiments?
Sandbagging is maybe really hard though! See here. In general, the possibility of running control evals at crunch time is a lot of why I like control.