Very cool work. For certain forms of unlearning, e.g. dangerous Chemical/Biological/Radiological/Nuclear knowledge, this seems like a great idea.
However, for something like deception, which you quoted as a motivating example, there’s a challenge. We don’t want a model that has no idea what deception is:
a) it would be incredibly easy to jailbreak: just tell it you’re trustworthy and have a legitimate need to the information b) it would have a lot of trouble understanding human behavior, or writing stories, or doing computer security, etc. etc. c) it would generally have a big, glaring hole in its world model of people d) if and when it did figure out that humans were often deceptive, its reactions might be hard to predict
What we need is a model that understands deception, but doesn’t use it on us. Ideally, one that provably doesn’t use it, because we are sure we could detect and measure if it did.
Knowing exactly where the model’s representation of “deception” is might well still be a great start. But I suspect we wouldn’t just abliterate it, but would instead investigate, understand, and monitor it.
I’ve edited the article to use “The Golden Gate Bridge” instead of “deception” as the motivating example. It’s less motivating for sure, but I think it makes the technique easier to discuss. But you’ve really got me thinking about the complexities of capability removal.
True, concrete knowledge would have been a better motivating example. Intervening usefully on deception is a moonshot.
I wonder if it would be possible to intervene on a concept only for continuation tokens, while leaving prompt tokens uninhibited? I don’t know what might carry over via attention, but I’d like to try it.
Very cool work. For certain forms of unlearning, e.g. dangerous Chemical/Biological/Radiological/Nuclear knowledge, this seems like a great idea.
However, for something like deception, which you quoted as a motivating example, there’s a challenge. We don’t want a model that has no idea what deception is:
a) it would be incredibly easy to jailbreak: just tell it you’re trustworthy and have a legitimate need to the information
b) it would have a lot of trouble understanding human behavior, or writing stories, or doing computer security, etc. etc.
c) it would generally have a big, glaring hole in its world model of people
d) if and when it did figure out that humans were often deceptive, its reactions might be hard to predict
What we need is a model that understands deception, but doesn’t use it on us. Ideally, one that provably doesn’t use it, because we are sure we could detect and measure if it did.
Knowing exactly where the model’s representation of “deception” is might well still be a great start. But I suspect we wouldn’t just abliterate it, but would instead investigate, understand, and monitor it.
I’ve edited the article to use “The Golden Gate Bridge” instead of “deception” as the motivating example. It’s less motivating for sure, but I think it makes the technique easier to discuss. But you’ve really got me thinking about the complexities of capability removal.
True, concrete knowledge would have been a better motivating example. Intervening usefully on deception is a moonshot.
I wonder if it would be possible to intervene on a concept only for continuation tokens, while leaving prompt tokens uninhibited? I don’t know what might carry over via attention, but I’d like to try it.