I think it does matter for emergent misalignment purposes. I’d think the emergent misalignment is more likely when you teach the model “these things are bad, don’t do them” then fine tune them to do one of those previously prohibited behaviors.
This is probably true, but emergent misalignment has been found in base models too (https://arxiv.org/pdf/2502.17424)
I think it does matter for emergent misalignment purposes. I’d think the emergent misalignment is more likely when you teach the model “these things are bad, don’t do them” then fine tune them to do one of those previously prohibited behaviors.
This is probably true, but emergent misalignment has been found in base models too (https://arxiv.org/pdf/2502.17424)