This post makes me happy. It has comparison to other models, comparison to baselines, error bars, random control, no graph crimes, and its not in appendix.
This should be norm for all papers of this type.
J-Space paper from Anthropic broke some of these good practices of good science like this.
This post makes me happy. It has comparison to other models, comparison to baselines, error bars, random control, no graph crimes, and its not in appendix.
This should be norm for all papers of this type.
J-Space paper from Anthropic broke some of these good practices of good science like this.