How often can frontier models autonomously solve math problems that required human-AI collaboration only a few months ago? This seems evaluable and may be useful as an early warning.
Frontier models are resolvingmanyopenproblems and saturating many benchmarks. At this point, would it also be worth tracking how often the newest AI models can zero-shot proofs and counterexamples that people recently obtained in collaboration with older AI models?
For the sake of argument, suppose that many recent human contributions turn out to be redundant to the capabilities of AI models released weeks or months later, in the sense that results which previously required humans and AIs to work collaboratively can now be obtained by AI alone. This would likely feel unpleasant. However, if and when this occurs, it could provide an important early warning that we are exiting the “centaur phase,” during which human-AI collaborations are typically stronger than humans alone and AIs alone. So it would be good to know.
I’ve hesitated to make this suggestion because I’m not a mathematician. I couldn’t do this tracking and I can’t even fully assess whether it is a good idea. However, others have surely had similar thoughts by now. If the idea isn’t feasible, then I’d like understand more—what would prevent it? I’d guess that others would like to know too.
Some obstacles do seem addressable. If the details interest you, please see below the folds.
Knowledge cutoffs
For this to work, tracking would need to focus on problems solved by human-AI collaboration after the knowledge cutoff date of the newest AI model (but before that model’s release date). Otherwise, solutions that a model obtained de novo could not be distinguished from solutions that it remembered from training data. Evaluations would be performed with web search disabled. For simplicity, it would also help if original solutions were Lean-verified.
Sample size and misclassification
When I looked through Erdos problems a month ago, it seemed like 10-20 recent solutions met both the knowledge cutoff and Lean criteria. This is not a big sample, but is large enough to give a rough sense of whether the human part of the human-AI collaborations could be replaced by subsequent models working on their own. It is not smaller than the sample used in First Proof, for example. More problems may meet the criteria today.
However, it is slippery to figure out which solutions actually involved meaningful human and AI contributions. I may have misjudged. Also, a user could easily make derivations in extended interaction with AI, without realizing that AI could have provided the same results alone. To address this, it would be useful to run the original model on the problem alone, to verify that it fails to resolve the problem. (This is only possible until the original model is deprecated.) Care would be needed to avoid blame in such situations. I think it could happen to any of us.
Interpretation
Another obstacle is that proper interpretation would be subtle. The class of problem that is included in tracking may greatly affect how often newer AI can replace the human contribution to recent derivations. Further, even if it is found that newer AI was able to fully replace human contributions, it does not strictly follow that we are leaving the centaur phase. For example, it remains possible that the mathematicians aimed for low-hanging fruit with their recent contributions, which were therefore easiest for newer models to replicate, and that mathematicians could also reach higher fruit that AI still cannot obtain. Additionally, solving problems does not touch other important aspects of math, like theory building and exercising taste.
A last obstacle is that the suggested tracking could become impossible when continual learning arrives, owing to the removal of knowledge cutoffs. The opportunity to do this work may disappear, but that is all the more reason to do it now.
How often can frontier models autonomously solve math problems that required human-AI collaboration only a few months ago? This seems evaluable and may be useful as an early warning.
Frontier models are resolving many open problems and saturating many benchmarks. At this point, would it also be worth tracking how often the newest AI models can zero-shot proofs and counterexamples that people recently obtained in collaboration with older AI models?
For the sake of argument, suppose that many recent human contributions turn out to be redundant to the capabilities of AI models released weeks or months later, in the sense that results which previously required humans and AIs to work collaboratively can now be obtained by AI alone. This would likely feel unpleasant. However, if and when this occurs, it could provide an important early warning that we are exiting the “centaur phase,” during which human-AI collaborations are typically stronger than humans alone and AIs alone. So it would be good to know.
I’ve hesitated to make this suggestion because I’m not a mathematician. I couldn’t do this tracking and I can’t even fully assess whether it is a good idea. However, others have surely had similar thoughts by now. If the idea isn’t feasible, then I’d like understand more—what would prevent it? I’d guess that others would like to know too.
Some obstacles do seem addressable. If the details interest you, please see below the folds.
Knowledge cutoffs
For this to work, tracking would need to focus on problems solved by human-AI collaboration after the knowledge cutoff date of the newest AI model (but before that model’s release date). Otherwise, solutions that a model obtained de novo could not be distinguished from solutions that it remembered from training data. Evaluations would be performed with web search disabled. For simplicity, it would also help if original solutions were Lean-verified.
Sample size and misclassification
When I looked through Erdos problems a month ago, it seemed like 10-20 recent solutions met both the knowledge cutoff and Lean criteria. This is not a big sample, but is large enough to give a rough sense of whether the human part of the human-AI collaborations could be replaced by subsequent models working on their own. It is not smaller than the sample used in First Proof, for example. More problems may meet the criteria today.
However, it is slippery to figure out which solutions actually involved meaningful human and AI contributions. I may have misjudged. Also, a user could easily make derivations in extended interaction with AI, without realizing that AI could have provided the same results alone. To address this, it would be useful to run the original model on the problem alone, to verify that it fails to resolve the problem. (This is only possible until the original model is deprecated.) Care would be needed to avoid blame in such situations. I think it could happen to any of us.
Interpretation
Another obstacle is that proper interpretation would be subtle. The class of problem that is included in tracking may greatly affect how often newer AI can replace the human contribution to recent derivations. Further, even if it is found that newer AI was able to fully replace human contributions, it does not strictly follow that we are leaving the centaur phase. For example, it remains possible that the mathematicians aimed for low-hanging fruit with their recent contributions, which were therefore easiest for newer models to replicate, and that mathematicians could also reach higher fruit that AI still cannot obtain. Additionally, solving problems does not touch other important aspects of math, like theory building and exercising taste.
Obstacles of interpretation are also faced by First Proof, Erdos problem tracking, and trends in AI-solved conjectures. All are valuable indicators of AI capability nonetheless.
Continual learning
A last obstacle is that the suggested tracking could become impossible when continual learning arrives, owing to the removal of knowledge cutoffs. The opportunity to do this work may disappear, but that is all the more reason to do it now.