From personal experience, ChatGPT still likes to spin in circles when it’s unable to completely one- or two-shot the task, all the while sounding like it’s making continual progress.
I spent a week trying to vibe-prove an optimization bound, and it almost always reached the same conclusion using different notation, presented as a “new useful reduction” and an almost completed proof with a minor “missing lemma”. Prompting it to attempt to prove this lemma only resulted in another restatement of the problem in new notation and a demand for an (essentially) equivalent lemma.
To be clear: ChatGPT’s work was neither trivial nor useless. I had it create a comprehensive pdf write-up and could then prove the result by actually steering the model and doing some work myself. I’m also like 80% sure that running an actual multi-agent workflow with a higher subscription tier would have succeeded.
multi-agent workflow with a higher subscription tier would have succeeded
you need Mythos/Fable as the research director. GPT 5.6 Sol High for literature search and review, and adversarial result reviewer and critic. Opus 5 (or Sol—it is comparatively cheap) as the agentic mathematician that operates within reviewable rounds or blocks.
From personal experience, ChatGPT still likes to spin in circles when it’s unable to completely one- or two-shot the task, all the while sounding like it’s making continual progress.
I spent a week trying to vibe-prove an optimization bound, and it almost always reached the same conclusion using different notation, presented as a “new useful reduction” and an almost completed proof with a minor “missing lemma”. Prompting it to attempt to prove this lemma only resulted in another restatement of the problem in new notation and a demand for an (essentially) equivalent lemma.
To be clear: ChatGPT’s work was neither trivial nor useless. I had it create a comprehensive pdf write-up and could then prove the result by actually steering the model and doing some work myself. I’m also like 80% sure that running an actual multi-agent workflow with a higher subscription tier would have succeeded.
you need Mythos/Fable as the research director. GPT 5.6 Sol High for literature search and review, and adversarial result reviewer and critic. Opus 5 (or Sol—it is comparatively cheap) as the agentic mathematician that operates within reviewable rounds or blocks.