To me it looks like a mix of cities, sex, food, politics and measurements. It doesn’t look random but does look polysemantic.
I looked at “Full”, which showed a lot more activations and variety of activations than “Snippet”, including strong green ones. I don’t know why this is, since I don’t know how neuronpedia works.
Tokens which occurred earlier in context window. With distance from the average position of previous occurrences determining the strength (so tends to fire on when a token was repeated a while back).
The way you find this neuron is by looking at the neurons which duplicate token heads feed into.
Tokens which occurred earlier in context window. With distance from the average position of previous occurrences determining the strength (so tends to fire on when a token was repeated a while back).
Could you explain that in more detail? I am looking at full example texts and failing to see a pattern like this. What is the relationship between the firing magnitude and the distance from the average position of previous occurrences?
What i’d say is that I think the neuron is probably not exactly monosemantic (it fires on “cons” always for instance). It’s just a good example of where the neuron’s behaviour is mostly explained by the duplicate pattern, but where this isn’t clear immediately.
What the draft analyses is the contribution just from positional information to the neurons activation, ignoring token specific contributions.
What would you say is the pattern?
To me it looks like a mix of cities, sex, food, politics and measurements. It doesn’t look random but does look polysemantic.
I looked at “Full”, which showed a lot more activations and variety of activations than “Snippet”, including strong green ones. I don’t know why this is, since I don’t know how neuronpedia works.
Tokens which occurred earlier in context window. With distance from the average position of previous occurrences determining the strength (so tends to fire on when a token was repeated a while back).
The way you find this neuron is by looking at the neurons which duplicate token heads feed into.
Could you explain that in more detail? I am looking at full example texts and failing to see a pattern like this. What is the relationship between the firing magnitude and the distance from the average position of previous occurrences?
It might be easiest if we just discuss tokens that you take issue with here?
If it helps I have a (very rough) draft from 1.5 years ago—Duplicate token neurons in the first layer of GPT-2 analysing the mechanism.
What i’d say is that I think the neuron is probably not exactly monosemantic (it fires on “cons” always for instance). It’s just a good example of where the neuron’s behaviour is mostly explained by the duplicate pattern, but where this isn’t clear immediately.
What the draft analyses is the contribution just from positional information to the neurons activation, ignoring token specific contributions.
I started lookin at what it doesn’t react to, too.
It reacted to “evolution”, “mating”, “evolved”, “genetics”, “condoms”, “erect penis” but not to “sex” or “sexuality”.
It reacts to “intermittent fasting”, “body fat”, “eating” but not consistently or strongly to “food”.