Maybe it would be good to know how often the AI co-developed solutions can now be obtained by the latest AI models working on their own. (The checks would need to be limited to problems solved after the knowledge cutoff dates of the latest models, and probably restricted to Lean-verified only.)
That might provide an indicator for when we are starting to exit the centaur phase, at least for the problem solving part of math.
Also, how many of the VibeMath users / contributors do you think are agents? In their dataset, I see only one identifiably human name (JSON key: submittedBy).
Maybe it would be good to know how often the AI co-developed solutions can now be obtained by the latest AI models working on their own. (The checks would need to be limited to problems solved after the knowledge cutoff dates of the latest models, and probably restricted to Lean-verified only.)
That might provide an indicator for when we are starting to exit the centaur phase, at least for the problem solving part of math.
Also, how many of the VibeMath users / contributors do you think are agents? In their dataset, I see only one identifiably human name (JSON key: submittedBy).