I wouldn’t take one or two datapoints on a single benchmark too seriously, especially with a methodology as fiddly as time horizon and concerns like Ryan’s. Nevertheless seems like a good thought that you replicated using time estimates from commit data, as the original difficulty estimates seemed likely to be noisy. I’ll be interested to see if the trend continues and if the same is currently true with OSWorld (Looks like they had a big update so maybe it’s possible to get individual task data now.)
I wouldn’t take one or two datapoints on a single benchmark too seriously, especially with a methodology as fiddly as time horizon and concerns like Ryan’s. Nevertheless seems like a good thought that you replicated using time estimates from commit data, as the original difficulty estimates seemed likely to be noisy. I’ll be interested to see if the trend continues and if the same is currently true with OSWorld (Looks like they had a big update so maybe it’s possible to get individual task data now.)
Yeah, I am also pretty much on the fence right now. But time will tell.