Great call. This makes me think about risk of emergent misalignment for open source models. Any open source model initially released with safeguards will likely see people finetune out the safeguards. I’ve largely viewed emergent misalignment as instructive, but probably not a huge threat in and of itself, this updates me that it could be a real threat.
Great call. This makes me think about risk of emergent misalignment for open source models. Any open source model initially released with safeguards will likely see people finetune out the safeguards. I’ve largely viewed emergent misalignment as instructive, but probably not a huge threat in and of itself, this updates me that it could be a real threat.