One might equally object that Raven’s Progressive Matrices are just an arbitrary set of shapes. However, the partial validation is precisely the positive manifold: e.g., the ECI correlates with the METR time horizons at r=0.85 (https://arxiv.org/pdf/2512.00193, p. 5).
It’s like: “hey, all of these different cognitive tasks we throw at people point at the same thing”.
I agree, insofar as I’ve understood you correctly there, that external validity is lacking. We can’t say “at 4 hours it takes the junior’s job, at 8 hours it takes …” (akin to how IQ correlates with academic and job success, or our proxies thereof).
The fact that we can’t perfectly measure general intelligence is a blessing in disguise however. We can only say “smart people do this maths and coding thing, let’s throw the AI’s at that”, which then runs into Goodhart’s Law.
So we actually don’t want a good y-axis!
I agree you get partial validation from the positive manifold (this comes up at the top of the ECI section, although I appreciate this wasn’t in excerpt I originally posted which you replied to). Models tend to do better across the board of metrics we throw at them, so there’s some general sense of ‘more capable’.
As you say, external validity is tricky. In a sense, the ‘ordinal vs. interval/ratio’ is partly this problem: IQ/capabilities index proxy well enough for “If the person/model scores higher, they are expected to do better (ordinal)”, but struggle to resolve “x% in score = ?% of performance in something or other”.
(I’m less sure around ‘better measurements would be bad news because it could accelerate capabilities. But hopefully I could retreat to ‘probably better not to make mistaken claims based on overinterpreting measurements as if they were true y-axis’ - cf. “IQ 150 = 50% smarter than average”
One might equally object that Raven’s Progressive Matrices are just an arbitrary set of shapes. However, the partial validation is precisely the positive manifold: e.g., the ECI correlates with the METR time horizons at r=0.85 (https://arxiv.org/pdf/2512.00193, p. 5). It’s like: “hey, all of these different cognitive tasks we throw at people point at the same thing”.
I agree, insofar as I’ve understood you correctly there, that external validity is lacking. We can’t say “at 4 hours it takes the junior’s job, at 8 hours it takes …” (akin to how IQ correlates with academic and job success, or our proxies thereof).
The fact that we can’t perfectly measure general intelligence is a blessing in disguise however. We can only say “smart people do this maths and coding thing, let’s throw the AI’s at that”, which then runs into Goodhart’s Law. So we actually don’t want a good y-axis!
I agree you get partial validation from the positive manifold (this comes up at the top of the ECI section, although I appreciate this wasn’t in excerpt I originally posted which you replied to). Models tend to do better across the board of metrics we throw at them, so there’s some general sense of ‘more capable’.
As you say, external validity is tricky. In a sense, the ‘ordinal vs. interval/ratio’ is partly this problem: IQ/capabilities index proxy well enough for “If the person/model scores higher, they are expected to do better (ordinal)”, but struggle to resolve “x% in score = ?% of performance in something or other”.
(I’m less sure around ‘better measurements would be bad news because it could accelerate capabilities. But hopefully I could retreat to ‘probably better not to make mistaken claims based on overinterpreting measurements as if they were true y-axis’ - cf. “IQ 150 = 50% smarter than average”