I was planning on doing my own fine-tuning to inject preferences, but this is a way easier way to de-risk!
It’s fairly in line with results in the post: the probes consistently transfer better than the utility correlation would predict, but this works less well on personas that are more different from the assistant.
I’m still pending approval for the misalignment one...
I was planning on doing my own fine-tuning to inject preferences, but this is a way easier way to de-risk!
It’s fairly in line with results in the post: the probes consistently transfer better than the utility correlation would predict, but this works less well on personas that are more different from the assistant.
I’m still pending approval for the misalignment one...
You should have access now, Sharan accepted a bunch yesterday