Thanks for doing this! Do you have data on this stuff sorted by model and/or by date? I’d be curious to see if these numbers are trending up or down over time for example, or if they are better or worse for larger models.
One issue is that each model was run on a different subset of data so direct comparisons are confounded. We’re working on resampling (a subset of) all of the model turns under a fixed battery of models, so that we can make that sort of comparison. We do have (the probably very confounded) results by model in the “Results by model” header above, but I probably wouldn’t read into those comparisons much.
Thanks for doing this! Do you have data on this stuff sorted by model and/or by date? I’d be curious to see if these numbers are trending up or down over time for example, or if they are better or worse for larger models.
One issue is that each model was run on a different subset of data so direct comparisons are confounded. We’re working on resampling (a subset of) all of the model turns under a fixed battery of models, so that we can make that sort of comparison. We do have (the probably very confounded) results by model in the “Results by model” header above, but I probably wouldn’t read into those comparisons much.