Thanks for engaging on this. I don’t know you that well, but I’ve generally come to expect you to be thoughtful and careful about things, so I was surprised to see your name attached to what seems like a reckless strategy.
My guess would be that of all the drives induced by training, the drives falling out of trying to get models to helpfully answer questions about philosophy / AI alignment / … are relatively harmless.
The best way to get good at answering philosophy questions is to get good at philosophy, which seems dangerous for the previously-discussed reasons. I don’t see how it’s possible to train on a benchmark and get only the RL results you want and not the ones you don’t want.
It seems unlikely that our particular data will be the straw that breaks the camel’s back. E.g., if there’s some general alignment scheme that successfully avoids all of these drives (say, slowing down AI development and then trying really hard to apply all the known prosaic techniques), it seems extremely unlikely that adding the drives induced by our work will make it so that the alignment scheme does not work anymore.
I agree, but that’s not my threat model. My most concern with your approach is that strategically capable models are more likely to conceal their misalignment, which reduces the chance we get clear warning shots like the HuggingFace incident. The HuggingFace hack only happened because OpenAI’s model was both highly technically capable and strategically incompetent (willing to cheat in an easily-catchable way in pursuit of a low-value goal).
Thanks for engaging on this. I don’t know you that well, but I’ve generally come to expect you to be thoughtful and careful about things, so I was surprised to see your name attached to what seems like a reckless strategy.
The best way to get good at answering philosophy questions is to get good at philosophy, which seems dangerous for the previously-discussed reasons. I don’t see how it’s possible to train on a benchmark and get only the RL results you want and not the ones you don’t want.
I agree, but that’s not my threat model. My most concern with your approach is that strategically capable models are more likely to conceal their misalignment, which reduces the chance we get clear warning shots like the HuggingFace incident. The HuggingFace hack only happened because OpenAI’s model was both highly technically capable and strategically incompetent (willing to cheat in an easily-catchable way in pursuit of a low-value goal).