I’d push back pretty hard against the footnote 15. Looking back at history of models, it’s pretty clear almost everything they do is some sort of emergent property and not something directly trained. We unleash the self-supervising algorithm and post-training pretty much narrows and shapes the skills that are mostly already there.
Regarding stylometry specifically—it’s a fundamental piece of how models work. Predicting text is basically dependent on being able to model the mind and the writing of the author, which goes both ways. Authorship inference is pretty much the primary thing they’re trained at, not a side effect. If anything, the labs tend to suppress the skill because of privacy issues, which is one reason I was actually surprised to see models doing it lately.
I’d push back pretty hard against the footnote 15. Looking back at history of models, it’s pretty clear almost everything they do is some sort of emergent property and not something directly trained. We unleash the self-supervising algorithm and post-training pretty much narrows and shapes the skills that are mostly already there.
Regarding stylometry specifically—it’s a fundamental piece of how models work. Predicting text is basically dependent on being able to model the mind and the writing of the author, which goes both ways. Authorship inference is pretty much the primary thing they’re trained at, not a side effect. If anything, the labs tend to suppress the skill because of privacy issues, which is one reason I was actually surprised to see models doing it lately.