I actually think your paper is great to start the discussion on EM, and it’s certainly not required for one of the first works in any field to already be perfect! The problem we see more is that many later works inherit these imperfections without questioning them, and already try to build early conclusions about EM (as a broad phenomenon) despite these evaluation/dataset limitations. Hence, we want to point them out in this position paper, but we certainly appreciate your effort in showing the phenomenon in the first place.
So, I’m somewhat worried that recommendations like these might lead to, let’s call it, “performative” science—“ok we have 4 pages on judges in the appendix, use 3 different models 7 prompts and 200 manually scored samples, check, good paper”—instead of focusing on what matters.
I think this is not contradictory to focusing on what matters. Usually (hopefully) people with ambitions to do great work start with what matters, and then be careful about how they evaluate model outputs. If a project itself isn’t interesting enough, I don’t think having multiple seeds/judges can significantly improve its significance.
I actually think your paper is great to start the discussion on EM, and it’s certainly not required for one of the first works in any field to already be perfect! The problem we see more is that many later works inherit these imperfections without questioning them, and already try to build early conclusions about EM (as a broad phenomenon) despite these evaluation/dataset limitations. Hence, we want to point them out in this position paper, but we certainly appreciate your effort in showing the phenomenon in the first place.
I think this is not contradictory to focusing on what matters. Usually (hopefully) people with ambitions to do great work start with what matters, and then be careful about how they evaluate model outputs. If a project itself isn’t interesting enough, I don’t think having multiple seeds/judges can significantly improve its significance.