This probably breaks if you separate tags by injecting a designed vector as reference, or even move from projecting a one hot matrix of token identity to predicting a two hot matrix of tok+role, but that is not something that most labs do, so that might be a mitigation, have something automated provide good flagging instead, but it involves injection so it has other issues.
This probably breaks if you separate tags by injecting a designed vector as reference, or even move from projecting a one hot matrix of token identity to predicting a two hot matrix of tok+role, but that is not something that most labs do, so that might be a mitigation, have something automated provide good flagging instead, but it involves injection so it has other issues.
Oh wait, this is a repeat comment.