Ah yes, that makes perfect sense, though I must say that the reasoning behind my not considering it significant is because the non-linearity of attention happens along the sequence dimension and not the feature dimension of the MLP blocks. In my view, attention might just amplify which token but not the underlying features within that particular token’s feature space. Again, I am a fresh graduate, so might be wrong. But I this type of topic invigorates me.
I think we were fairly confident it was going to be the MLP blocks, but attention also has a non-linearity via the softmax.
Ah yes, that makes perfect sense, though I must say that the reasoning behind my not considering it significant is because the non-linearity of attention happens along the sequence dimension and not the feature dimension of the MLP blocks. In my view, attention might just amplify which token but not the underlying features within that particular token’s feature space. Again, I am a fresh graduate, so might be wrong. But I this type of topic invigorates me.