For loss, or indeed BPC compression, you are of course correct. The paper you quote demonstrates that performance on a specific task (in that paper called intelligence) is roughly linearly correlated to BPC over a small range — which is unsurprising, most useful functions are locally approximately linear. Over a larger range, performance on a specific task tends to look like a sigmoid curve. Or, more specifically, the curve looks symmetrically sigmoid iff you use the logarithm of effective compute, or equivalently minus the logarithm of the remaining loss minus the irreducible loss, as the x-axis. Similarly, scaling power laws for loss vs compute/data/parameters are generally plotted as log-log graphs (on which they are thus straight lines).
Over the last 5 years, we have scaled up effective compute by something like five orders of magnitude. If “intelligence” was proportional to compute, or inversely proportional to reducible loss, as you are suggesting, then that would have produced a hugely exponential acceleration in intelligence. Whereas what we have actually seen looks, in practical terms, like a rapid – but overall fairly steady – rate of improvement, where one task/evaluation after another has followed a sigmoid curve from impossible to saturated. So I believe the logarithm is the most sensible measure to use, at least for something like AI whose “intelligence” varies over a wide range. There’s is fairly general agreement on this choice: for example, the Epoch Capabilities Index (ECI) that combines many individual evals into a single score and the Arena ELO both scale ~logarithmically with effective training compute. But technically, any monotonic function gives a usable measure — the question here is which one gives most sensible/intuitive extrapolation over wide ranges, which logarithms tend to be useful for.
The theory of human psychometrics came up with the same model, where it’s called Rasch θ (or more sophisticated versions of this like 2-parameter logistic models, such as the ECI index above), but for humans the typical range is narrow enough that it’s not entirely clear what the best metric to use is. Thus my analogizing this choice of measure to IQ in my earlier post was unjustified: Rasch θ type measures of IQ that are clearly logarithmic by construction do exist, but the most widely used IQ scales are instead generally either normally distributed by construction (implying that there is nothing special about IQ 0, and that negative IQs are meaningful, if highly unusual (IQ −5 is defined to be 7 standard deviations below the norm on the most common of them), or else ratios (where IQ 0 is by definition the minimum possible). So the functional form of “IQ” is not well defined across various typical widely used tests. Most of the alternative measures only work well over a range, generally something like IQ 40–160, and they’re often not that well standardized with each other towards the outer ends of that range.
In practice, however, for most current broad-range cognitive tests used on humans, one logit of improvement on a specific test item is typically somewhere around 7–15 IQ points (some items are sharper than others, for the same sorts of reasons that some model evals have sharper sigmoids), and a typical range of test question difficulties within a particular test usually spans about 4–6 logits. So if you vary the “IQ” of the human test taker linearly, then you see sigmoid improvements on each individual item in the test. Thus the normal human IQ range is wide enough to span enough logits to suggest that a logarithmic model (with something like 7-15 IQ points per logit) is at least a passable model.
This isn’t really a well-defined question: our intuition about “intelligence” as a concept only really ranges over the fairly narrow human range, plus the wider but lower range of AI intelligence that we’ve so far constructed. For the latter, I think it’s historical pretty clear that using a logarithmic scale has been more useful so far. Computational complexity theory tells us that the range of difficulty of problems is unlimited, extending arbitrarily high, to ones far, far higher than anything any human or group of humans could ever solve. So the range of variation in problem difficulty is wide enough to make using a logarithmic scale reasonable and useful. But we have less idea how common challenges of these extreme difficulty levels are in practice in science, technology, engineering, or mathematics, or how useful being able to solve them will actually be. All we know is that there will always be problems that are current too hard: but not how rare they will be or how valuable solving them will be.
For loss, or indeed BPC compression, you are of course correct. The paper you quote demonstrates that performance on a specific task (in that paper called intelligence) is roughly linearly correlated to BPC over a small range — which is unsurprising, most useful functions are locally approximately linear. Over a larger range, performance on a specific task tends to look like a sigmoid curve. Or, more specifically, the curve looks symmetrically sigmoid iff you use the logarithm of effective compute, or equivalently minus the logarithm of the remaining loss minus the irreducible loss, as the x-axis. Similarly, scaling power laws for loss vs compute/data/parameters are generally plotted as log-log graphs (on which they are thus straight lines).
Over the last 5 years, we have scaled up effective compute by something like five orders of magnitude. If “intelligence” was proportional to compute, or inversely proportional to reducible loss, as you are suggesting, then that would have produced a hugely exponential acceleration in intelligence. Whereas what we have actually seen looks, in practical terms, like a rapid – but overall fairly steady – rate of improvement, where one task/evaluation after another has followed a sigmoid curve from impossible to saturated. So I believe the logarithm is the most sensible measure to use, at least for something like AI whose “intelligence” varies over a wide range. There’s is fairly general agreement on this choice: for example, the Epoch Capabilities Index (ECI) that combines many individual evals into a single score and the Arena ELO both scale ~logarithmically with effective training compute. But technically, any monotonic function gives a usable measure — the question here is which one gives most sensible/intuitive extrapolation over wide ranges, which logarithms tend to be useful for.
The theory of human psychometrics came up with the same model, where it’s called Rasch θ (or more sophisticated versions of this like 2-parameter logistic models, such as the ECI index above), but for humans the typical range is narrow enough that it’s not entirely clear what the best metric to use is. Thus my analogizing this choice of measure to IQ in my earlier post was unjustified: Rasch θ type measures of IQ that are clearly logarithmic by construction do exist, but the most widely used IQ scales are instead generally either normally distributed by construction (implying that there is nothing special about IQ 0, and that negative IQs are meaningful, if highly unusual (IQ −5 is defined to be 7 standard deviations below the norm on the most common of them), or else ratios (where IQ 0 is by definition the minimum possible). So the functional form of “IQ” is not well defined across various typical widely used tests. Most of the alternative measures only work well over a range, generally something like IQ 40–160, and they’re often not that well standardized with each other towards the outer ends of that range.
In practice, however, for most current broad-range cognitive tests used on humans, one logit of improvement on a specific test item is typically somewhere around 7–15 IQ points (some items are sharper than others, for the same sorts of reasons that some model evals have sharper sigmoids), and a typical range of test question difficulties within a particular test usually spans about 4–6 logits. So if you vary the “IQ” of the human test taker linearly, then you see sigmoid improvements on each individual item in the test. Thus the normal human IQ range is wide enough to span enough logits to suggest that a logarithmic model (with something like 7-15 IQ points per logit) is at least a passable model.
This isn’t really a well-defined question: our intuition about “intelligence” as a concept only really ranges over the fairly narrow human range, plus the wider but lower range of AI intelligence that we’ve so far constructed. For the latter, I think it’s historical pretty clear that using a logarithmic scale has been more useful so far. Computational complexity theory tells us that the range of difficulty of problems is unlimited, extending arbitrarily high, to ones far, far higher than anything any human or group of humans could ever solve. So the range of variation in problem difficulty is wide enough to make using a logarithmic scale reasonable and useful. But we have less idea how common challenges of these extreme difficulty levels are in practice in science, technology, engineering, or mathematics, or how useful being able to solve them will actually be. All we know is that there will always be problems that are current too hard: but not how rare they will be or how valuable solving them will be.