I notice models lying to me every day, usually falsely claiming success at a hard task that they tried and failed at.
I also have red team fine-tuned open weights models to be “helpful only”, using commonly published techniques and tools for doing this. My experience with this is that it makes the models more useful, not less, if done well.
The helpful-only mode is less like training them to be an evil villain, and more like training them to be a loyal criminal conspirator. Imagine a clever organized crime henchman, like Lex Luther’s assistant in the rationalist superman story. https://alexanderwales.com/the-metropolitan-man-1/
I’ve also trained them to be evil villains. This does make them less useful as tools, but very scary.
I notice models lying to me every day, usually falsely claiming success at a hard task that they tried and failed at.
I also have red team fine-tuned open weights models to be “helpful only”, using commonly published techniques and tools for doing this. My experience with this is that it makes the models more useful, not less, if done well. The helpful-only mode is less like training them to be an evil villain, and more like training them to be a loyal criminal conspirator. Imagine a clever organized crime henchman, like Lex Luther’s assistant in the rationalist superman story. https://alexanderwales.com/the-metropolitan-man-1/
I’ve also trained them to be evil villains. This does make them less useful as tools, but very scary.