I want to flag for local validity concerns that this specific eval is possibly contaimnated, and in particular this is only one eval (which is why we need more evals around this, now that OpenAI will probably scale this technique), as well as evals around tasks of fixed serial complexity with increasing information-free token production. And while I do find the supporting evidence strong enough to rule out claims that the eval is so contaminated that it isn’t a good measure of No-CoT capabilities whatsoever, and I think the eval is still good, I wouldn’t have gone too hard on OpenAI if the only evidence for that was the No-CoT evaluation.
To be clear, OpenAI comms are unacceptably misleading/non-informative, especially in light of new incidents that they knew happened and didn’t disclose (and plausibly did not disclose for legal reasons), but just chiming in about the issue here.
Relevant quote from Jai:
Testing could take the form of evaluating tasks of fixed serial complexity with increasing information-free token production. If capabilities predictably increase despite no additional useful CoT this would be strong evidence of effective depth chaining.
I want to flag for local validity concerns that this specific eval is possibly contaimnated, and in particular this is only one eval (which is why we need more evals around this, now that OpenAI will probably scale this technique), as well as evals around tasks of fixed serial complexity with increasing information-free token production. And while I do find the supporting evidence strong enough to rule out claims that the eval is so contaminated that it isn’t a good measure of No-CoT capabilities whatsoever, and I think the eval is still good, I wouldn’t have gone too hard on OpenAI if the only evidence for that was the No-CoT evaluation.
To be clear, OpenAI comms are unacceptably misleading/non-informative, especially in light of new incidents that they knew happened and didn’t disclose (and plausibly did not disclose for legal reasons), but just chiming in about the issue here.
Relevant quote from Jai: