Unfortunately, your post on EA doesn’t seem to account for the very enormous disparity in capabilities. For example, you wrote that “Elo is a transformation from the Bradley-Terry model, and although Elo has no true zero or multiple (Magnus Carlsen has double my Elo ≠ I’m half the chess player he is), Bradley-Terry strengths have both: 0 is an absolute zero of player strength—you always lose versus anyone else; doubling my strength means double my odds of winning, no matter who I am up against.”
The Elo rating and the BT scores are trivially connected by taking the logarithm: if two players have Bradley-Terry scores and , then it is equivalent to assigning them Elo ratings of and SOTA Elo ratings of humans span a wide range from ~1000 or less to ~2800. When translated to BT scores, this means a difference of at least tens of thousands of times. How are we supposed to land that onto a graph, if not by using logarithms?
Similarly, the METR time horizon should be treated like the BT scores varying tens of thousands of times. If we take the logarithm of the horizon, a wonder occurs and straight lines become obvious.
I’m afraid I don’t follow the objection. We can land 10k multiples onto a graph with a linear axis: the unlogged METR plot does exactly this, as do many pictures of exponential growth. I agree as a practical matter you often want to log the y axis to display this sort of data—you might get a nice straight line, and in any case your time series doesn’t spend a lot of its time within a pixel of the x-axis.
But that is distinct from (e.g.) whether “AI capabilities” are improving proportional to Time horizon or log(Time Horizon). Or, if you prefer deflationary/‘taboo X’ approaches, what is the transfer function between (these measurements) and (whatever you’re interested in).
Unfortunately, your post on EA doesn’t seem to account for the very enormous disparity in capabilities. For example, you wrote that “Elo is a transformation from the Bradley-Terry model, and although Elo has no true zero or multiple (Magnus Carlsen has double my Elo ≠ I’m half the chess player he is), Bradley-Terry strengths have both: 0 is an absolute zero of player strength—you always lose versus anyone else; doubling my strength means double my odds of winning, no matter who I am up against.”
The Elo rating and the BT scores are trivially connected by taking the logarithm: if two players have Bradley-Terry scores and , then it is equivalent to assigning them Elo ratings of and SOTA Elo ratings of humans span a wide range from ~1000 or less to ~2800. When translated to BT scores, this means a difference of at least tens of thousands of times. How are we supposed to land that onto a graph, if not by using logarithms?
Similarly, the METR time horizon should be treated like the BT scores varying tens of thousands of times. If we take the logarithm of the horizon, a wonder occurs and straight lines become obvious.
I’m afraid I don’t follow the objection. We can land 10k multiples onto a graph with a linear axis: the unlogged METR plot does exactly this, as do many pictures of exponential growth. I agree as a practical matter you often want to log the y axis to display this sort of data—you might get a nice straight line, and in any case your time series doesn’t spend a lot of its time within a pixel of the x-axis.
But that is distinct from (e.g.) whether “AI capabilities” are improving proportional to Time horizon or log(Time Horizon). Or, if you prefer deflationary/‘taboo X’ approaches, what is the transfer function between (these measurements) and (whatever you’re interested in).