I think “stop doing harmful work until you are fired, instead of quitting” is fairly compelling.
But, I do think most people working on prosaic safety are mostly doing harm.
If you don’t have a particular story for how what you’re doing scales to superalignment*, I a) don’t think you’re solving particularly important problems in the nearterm, b) in the medium term, mostly making it marginally faster to roll out stronger capabilities, which is bad because it’s burning calendar time for serial research time, and solving the legible problems leaving the illegible ones.
I think “stop doing harmful work until you are fired, instead of quitting” is fairly compelling.
But, I do think most people working on prosaic safety are mostly doing harm.
If you don’t have a particular story for how what you’re doing scales to superalignment*, I a) don’t think you’re solving particularly important problems in the nearterm, b) in the medium term, mostly making it marginally faster to roll out stronger capabilities, which is bad because it’s burning calendar time for serial research time, and solving the legible problems leaving the illegible ones.