I agree you get partial validation from the positive manifold (this comes up at the top of the ECI section, although I appreciate this wasn’t in excerpt I originally posted which you replied to). Models tend to do better across the board of metrics we throw at them, so there’s some general sense of ‘more capable’.
As you say, external validity is tricky. In a sense, the ‘ordinal vs. interval/ratio’ is partly this problem: IQ/capabilities index proxy well enough for “If the person/model scores higher, they are expected to do better (ordinal)”, but struggle to resolve “x% in score = ?% of performance in something or other”.
(I’m less sure around ‘better measurements would be bad news because it could accelerate capabilities. But hopefully I could retreat to ‘probably better not to make mistaken claims based on overinterpreting measurements as if they were true y-axis’ - cf. “IQ 150 = 50% smarter than average”
I agree you get partial validation from the positive manifold (this comes up at the top of the ECI section, although I appreciate this wasn’t in excerpt I originally posted which you replied to). Models tend to do better across the board of metrics we throw at them, so there’s some general sense of ‘more capable’.
As you say, external validity is tricky. In a sense, the ‘ordinal vs. interval/ratio’ is partly this problem: IQ/capabilities index proxy well enough for “If the person/model scores higher, they are expected to do better (ordinal)”, but struggle to resolve “x% in score = ?% of performance in something or other”.
(I’m less sure around ‘better measurements would be bad news because it could accelerate capabilities. But hopefully I could retreat to ‘probably better not to make mistaken claims based on overinterpreting measurements as if they were true y-axis’ - cf. “IQ 150 = 50% smarter than average”