There’s a tension between this and getting useful safety work out of models. We might need to use shallowly aligned models to accelerate safety, ban using them for capabilities, and use non-shallowly-aligned models to test alignment techniques.
Hopefully you can do this without getting rid of evidence. E.g. grader awareness should still be visible in CoT on these crucial tasks, and we probably shouldn’t just directly train against the concentrated failures (eg HF incident).
There’s a tension between this and getting useful safety work out of models. We might need to use shallowly aligned models to accelerate safety, ban using them for capabilities, and use non-shallowly-aligned models to test alignment techniques.
Hopefully you can do this without getting rid of evidence. E.g. grader awareness should still be visible in CoT on these crucial tasks, and we probably shouldn’t just directly train against the concentrated failures (eg HF incident).