Does anyone know how systematically OpenAI/Anthropic are throwing the models at open conjectures? Are they selecting conjectures that seem tractable? Or are they just throwing thousands at the wall and seeing which small subset end up getting solved?
This seems relevant! If the models are solving a high proportion of the conjectures given to them, that means something quite different than if they are solving a tiny proportion (in particular, it might update you towards these conjectures being outliers in “easiness”).
My expectation is that the labs are probably (?) doing the firehose of conjectures thing, based on the fact that the recent 10 Astra results apparently only cost 2k in total.
My guess is that it did start off as one off employee projects, but now that they’ve gotten so much news it will be rapidly institutionalised and a pipeline will be drawn up (similar to the “GPT for science”/Claude Science/GPT-Rosalind style initiatives)
It seemed to me that OpenAI have been somewhat intentionally fostering this (starting with providing early access to famous mathematicians and shopping any results around on twitter back when the models weren’t good enough to one shot significant stuff)
Seems plausible! Perhaps the Astra results are the first outputs of this pipeline? I imagine it wouldn’t be super hard to put together a functional pipeline here (maybe this is naive, but apparently Fable solved 5⁄10 of these with a bare prompt)
Surely with Lean this is just part and parcel of RLVR?
My best guess is “a huge number, and now the tail of the band between reliably solves and cannot solve is newsworthy,” but I don’t have particular knowledge of what they’re actually doing. It’s surprising to me that the Lean was apparently a post-hoc artifact from Astra, I would’ve expected all of these results to have been Lean spat out by a RLVR trace.
I doubt seemingly-curiosity-driven results like the Jacobian conjecture will remain typical for long.
Does anyone know how systematically OpenAI/Anthropic are throwing the models at open conjectures? Are they selecting conjectures that seem tractable? Or are they just throwing thousands at the wall and seeing which small subset end up getting solved?
This seems relevant! If the models are solving a high proportion of the conjectures given to them, that means something quite different than if they are solving a tiny proportion (in particular, it might update you towards these conjectures being outliers in “easiness”).
My expectation is that the labs are probably (?) doing the firehose of conjectures thing, based on the fact that the recent 10 Astra results apparently only cost 2k in total.
My guess is that it did start off as one off employee projects, but now that they’ve gotten so much news it will be rapidly institutionalised and a pipeline will be drawn up (similar to the “GPT for science”/Claude Science/GPT-Rosalind style initiatives)
It seemed to me that OpenAI have been somewhat intentionally fostering this (starting with providing early access to famous mathematicians and shopping any results around on twitter back when the models weren’t good enough to one shot significant stuff)
Seems plausible! Perhaps the Astra results are the first outputs of this pipeline? I imagine it wouldn’t be super hard to put together a functional pipeline here (maybe this is naive, but apparently Fable solved 5⁄10 of these with a bare prompt)
Surely with Lean this is just part and parcel of RLVR?
My best guess is “a huge number, and now the tail of the band between reliably solves and cannot solve is newsworthy,” but I don’t have particular knowledge of what they’re actually doing. It’s surprising to me that the Lean was apparently a post-hoc artifact from Astra, I would’ve expected all of these results to have been Lean spat out by a RLVR trace.
I doubt seemingly-curiosity-driven results like the Jacobian conjecture will remain typical for long.