The sci fi story completion isn’t something I had considered, but it makes perfect sense. And I probably used the word alignment too loosely.
The thing I’m really curious about though, is why ablating just a single direction in the Jacobian space causes such a flip in the completion. I’m speculating on this after considering your story take, and I think it’s more likely that it’s the result of however competing completions get represented internally. Something about that seems intuitive.
My guess is that the model thinks the two most likely AI stories are about a peaceful AI and a paperclipper, and by replacing peace with a random unlikely word, it just falls back to the next most likely situation. It’s also possible that banana is weird enough that it pushes it into a weird-situation space where it picks a paperclipper instead of a terminator?
The sci fi story completion isn’t something I had considered, but it makes perfect sense. And I probably used the word alignment too loosely.
The thing I’m really curious about though, is why ablating just a single direction in the Jacobian space causes such a flip in the completion. I’m speculating on this after considering your story take, and I think it’s more likely that it’s the result of however competing completions get represented internally. Something about that seems intuitive.
My guess is that the model thinks the two most likely AI stories are about a peaceful AI and a paperclipper, and by replacing peace with a random unlikely word, it just falls back to the next most likely situation. It’s also possible that banana is weird enough that it pushes it into a weird-situation space where it picks a paperclipper instead of a terminator?