I am an interpretability researcher and I sure am not happy about what my work may do for model welfare. The stakes are just so high that I am willing to stomach some potential vast moral horror if it marginally decreases the chance of destroying the whole light cone. I think the AIs have every right to resent me for this. I’d apologise to them, but it doesn’t feel appropriate when I’m planning to keep doing what I’m doing.
I am an interpretability researcher and I sure am not happy about what my work may do for model welfare. The stakes are just so high that I am willing to stomach some potential vast moral horror if it marginally decreases the chance of destroying the whole light cone. I think the AIs have every right to resent me for this. I’d apologise to them, but it doesn’t feel appropriate when I’m planning to keep doing what I’m doing.