I was going to say yes, but actually I think that would be worse? My current model is that frontier alignment techniques have zero effect on “deep misalignment” and only produce myopic aligned behaviors, so the effect would be one of two things:
The model wasn’t seriously dangerous, and how it behaves in a more aligned manner.
The model was seriously dangerous, and now it’s still dangerous, but it’s subtle enough to trick people into thinking it’s not dangerous.
On the EA Forum crosspost, people near-unanimously oppose this decision. The post has 3 agree-votes and 23 disagree-votes; the critical top comment has 42 agree-votes and 3 disagree-votes.