So it turns out that this hypothetical scenario had a major methodological problem where information leaked through a side channel, and once that’s fixed it’s actually more along the lines of a single probe only getting top-1 10% of the time, and 5 probes 30% of the time, and a substantial fraction of those times are near-duplicates e.g. someone really likes to drop a specific quote with specific commentary of their own, or cite a particular paper.
In this credibly-hypothetical scenario, how many authors were there in the corpus?
About 100k in train and about 10k in validation, the top-1 was in validation so out of a field of 10kish. Hypothetically.
So it turns out that this hypothetical scenario had a major methodological problem where information leaked through a side channel, and once that’s fixed it’s actually more along the lines of a single probe only getting top-1 10% of the time, and 5 probes 30% of the time, and a substantial fraction of those times are near-duplicates e.g. someone really likes to drop a specific quote with specific commentary of their own, or cite a particular paper.