AI is eating Mathematics. What are we waiting for to scale up automated safety work?
Epistemic status: I spent 5 hours thinking about this in March 2026. This is a Claude-written summary of a longer document, then improved manually. This post has been on my to-do list for a while; I did the 80⁄20 to get this out.
I buy the core case for AI labor. There may be a brief window between early AGIs and superintelligence in which AI could do orders of magnitude more safety-relevant work than humans alone, and we are already behind in preparing for it.
Tldr: What are we waiting for? Claude just formalized Fermat’s last theorem. OpenAI released PaperBench. People are already automating mathematics. Why don’t we do the same in AI Safety? Nobody is pointing compute at replicating safety papers, rerunning the fine-tunes, testing the empirical predictions we keep writing down and never checking. A lot of slop is coming, and investing in evaluating it and identifying the potential nuggets it contains is probably very important.
P.S. Since writing the initial version of the post, Resolution has been announced, but they seem to be heavily focused on alignment. This might be a good call, but this post is talking about broader safety work, not just alignment. Also, it’s not nearly enough to have a single initiative working on this.
Scheming is not what I expect to break first. Writing alignment papers is much lower-stakes than automating an entire lab’s codebase, which companies are already doing. The time horizon is measured in months; papers are largely independent of each other, so cumulative sabotage is unlikely, and any sabotage we catch would be useful evidence.
What breaks first is slop/evaluation. A scaled-up program produces a flood of automated research, with no reliable way to tell which parts are any good. I have run into this directly. I co-directed GRASP 2 years ago, an AI-assisted “Wikipedia of AI safety for policymakers”, and we produced a large volume of AI-written content: it was very sloppy. Every single output needed a human to check it. Also, even if factually correct, readers discount anything that looks AI-generated (for good reasons: generally, something written by AI hasn’t been thoroughly endorsed. The writing can be tactically brilliant and strategically very bad, with quite egregious mistakes). CURL canceled its bug bounty under a flood of AI-generated spam, even though genuine CVEs were arriving through the same channel. Scaling AI-assisted research before solving quality evaluation buys expensive slop.
I’d be interested in seeing those two benchmarks study quality evaluation:
First: take a funder’s own historical grant decisions (for small grants), give a model the proposal, web search and the relevant context, and have it produce a structured evaluation and a fund-or-reject with reasoning. Measure agreement with the human decision at each stage of the pipeline. Then measure how much human oversight, at what level of expertise, is needed to close the gap. That is a scalable oversight experiment on a task we already understand and have labeled data for. It is also immediately useful, because grantmaker capacity is a real constraint. Note: Manifund has started doing this.
Second: reproduction of AI safety technical research. Take published empirical alignment papers, mask the quantitative results, hand the model the research question and the methodology, and see whether it reproduces the numbers. The harder version gives it only the abstract (Contamination is the obvious objection. But a model that cannot reproduce a result it has already read is not ready for the novel case, so the failure is informative either way.) Nobody (I think) has run this for safety research specifically, despite it being clearly tractable, and people are already doing this.[1] As far as I know, the only thing preventing this from happening is the fact that we are just not starting.
2 other experiments to study the current uplift and see how far it is possible to automate today?
A pilot measuring uplift of a team of researchers when paying per result: 10 safety researchers, 3 weeks, frontier access with the best scaffolding from the benchmarks, paid per result. Measure against what they would have produced under their normal incentives, and rerun every few months to track the trend. The quality bar should be “useful on the Alignment Forum or at a workshop”, not NeurIPS.
Why could this be cool? Inkhaven shows that throughput shifts when incentives shift. I think tight deadlines are good. Part of what this would buy here is an exit from the academic treadmill, where researchers spend months over-polishing a useless paper because the reward is an ICML acceptance rather than impact.
Why not? To some extent, the karma on LW or Twitter is already a pay-per-result incentive.
Another option could be a bounty system: an expert panel with AI-assisted screening, paying per result that clears the bar, over a public list of concrete safety engineering problems. The bounty only works if the evaluation works; we probably need to scale up evaluation infrastructure for this type of production. This enables the separation of ideas from engineering.
Why not? Maybe this is already how organizations work implicitly: the director proposes the idea, and the team implements it? This is potentially already similar to Requests for Proposals from big funders.
My main concerns:
A lot of good safety research already exists and is not being absorbed, and as Buck Shlegeris put it, labs do not have the time or appetite for the 40 things that would probably solve most of the problem. Ten-xing research production without fixing the absorption and enforcement is filling a leaky bucket. But I guess we can do this in parallel.
Solving slop is a dangerous capability externality?
On your first point, I think an important reason that we’re making so much progress in mathematics as opposed to AI safety is that AI safety is not well formalized yet. It is impressive to see how much rigorous and intelligent Fable and Astra becomes when they have access to a Lean MCP. If you want to scale up alignment with automatization this way, I believe the most important thing is to formalize alignment directly, as in, allow the agent to have a test to immediately and empirically test its idea. The full blown version of this is Davidad’s infrabayesianism and hypersimulations, but we can also do this on a shorter scale, with e.g. a framework that allows Claude/Codex to specify the behavior they want to observe, and run many simulations of it.
I already have a good meta-plugin to write such MCP, so I think I could do this in the span of a week, if anyone is interested to work on this, let me know.
I agree with the sentiment of this post, but I think it is factually incorrect about Coordinal Research activities and as such I strongly downvoted.
I don’t believe I or Coordinal has anywhere stated or claimed that we have “curated 400+ open safety research questions”. I think that there are likely many groups/organizations/frontier-company-projects that include datasets like this; I am aware of a few personally. Coordinal never did this.
I worry that this failure was the result of poor agent/LLM use and oversight. It’s possible I am wrong and Coordinal has made this claim in the past, could you cite your source for this claim?
Alignment progress could be seen as contributing to capabilities work (see here and here) which is why I’m personally more excited about progress on broader questions (better information tech for coordination, robust institution design, etc) that target the meta-problem of humanity failing to effectively coordinate towards shared goals—though this is of course far more difficult (nearly impossible?) to formalize and thus harder for present day AIs to make progress on.
Not disagreeing; I just think the community needs extensive discussion/thinking on the net-impact of alignment work before deciding to throw more resources at it (whereas something like discovering/demonstrating misalignment seems more robustly beneficial).
AI is eating Mathematics. What are we waiting for to scale up automated safety work?
Epistemic status: I spent 5 hours thinking about this in March 2026. This is a Claude-written summary of a longer document, then improved manually. This post has been on my to-do list for a while; I did the 80⁄20 to get this out.
I buy the core case for AI labor. There may be a brief window between early AGIs and superintelligence in which AI could do orders of magnitude more safety-relevant work than humans alone, and we are already behind in preparing for it.
Tldr: What are we waiting for? Claude just formalized Fermat’s last theorem. OpenAI released PaperBench. People are already automating mathematics. Why don’t we do the same in AI Safety? Nobody is pointing compute at replicating safety papers, rerunning the fine-tunes, testing the empirical predictions we keep writing down and never checking. A lot of slop is coming, and investing in evaluating it and identifying the potential nuggets it contains is probably very important.
P.S. Since writing the initial version of the post, Resolution has been announced, but they seem to be heavily focused on alignment. This might be a good call, but this post is talking about broader safety work, not just alignment. Also, it’s not nearly enough to have a single initiative working on this.
Scheming is not what I expect to break first. Writing alignment papers is much lower-stakes than automating an entire lab’s codebase, which companies are already doing. The time horizon is measured in months; papers are largely independent of each other, so cumulative sabotage is unlikely, and any sabotage we catch would be useful evidence.
What breaks first is slop/evaluation. A scaled-up program produces a flood of automated research, with no reliable way to tell which parts are any good. I have run into this directly. I co-directed GRASP 2 years ago, an AI-assisted “Wikipedia of AI safety for policymakers”, and we produced a large volume of AI-written content: it was very sloppy. Every single output needed a human to check it. Also, even if factually correct, readers discount anything that looks AI-generated (for good reasons: generally, something written by AI hasn’t been thoroughly endorsed. The writing can be tactically brilliant and strategically very bad, with quite egregious mistakes). CURL canceled its bug bounty under a flood of AI-generated spam, even though genuine CVEs were arriving through the same channel. Scaling AI-assisted research before solving quality evaluation buys expensive slop.
I’d be interested in seeing those two benchmarks study quality evaluation:
First: take a funder’s own historical grant decisions (for small grants), give a model the proposal, web search and the relevant context, and have it produce a structured evaluation and a fund-or-reject with reasoning. Measure agreement with the human decision at each stage of the pipeline. Then measure how much human oversight, at what level of expertise, is needed to close the gap. That is a scalable oversight experiment on a task we already understand and have labeled data for. It is also immediately useful, because grantmaker capacity is a real constraint. Note: Manifund has started doing this.
Second: reproduction of AI safety technical research. Take published empirical alignment papers, mask the quantitative results, hand the model the research question and the methodology, and see whether it reproduces the numbers. The harder version gives it only the abstract (Contamination is the obvious objection. But a model that cannot reproduce a result it has already read is not ready for the novel case, so the failure is informative either way.) Nobody (I think) has run this for safety research specifically, despite it being clearly tractable, and people are already doing this.[1] As far as I know, the only thing preventing this from happening is the fact that we are just not starting.
2 other experiments to study the current uplift and see how far it is possible to automate today?
A pilot measuring uplift of a team of researchers when paying per result: 10 safety researchers, 3 weeks, frontier access with the best scaffolding from the benchmarks, paid per result. Measure against what they would have produced under their normal incentives, and rerun every few months to track the trend. The quality bar should be “useful on the Alignment Forum or at a workshop”, not NeurIPS.
Why could this be cool? Inkhaven shows that throughput shifts when incentives shift. I think tight deadlines are good. Part of what this would buy here is an exit from the academic treadmill, where researchers spend months over-polishing a useless paper because the reward is an ICML acceptance rather than impact.
Why not? To some extent, the karma on LW or Twitter is already a pay-per-result incentive.
Another option could be a bounty system: an expert panel with AI-assisted screening, paying per result that clears the bar, over a public list of concrete safety engineering problems. The bounty only works if the evaluation works; we probably need to scale up evaluation infrastructure for this type of production. This enables the separation of ideas from engineering.
Why not? Maybe this is already how organizations work implicitly: the director proposes the idea, and the team implements it? This is potentially already similar to Requests for Proposals from big funders.
My main concerns:
A lot of good safety research already exists and is not being absorbed, and as Buck Shlegeris put it, labs do not have the time or appetite for the 40 things that would probably solve most of the problem. Ten-xing research production without fixing the absorption and enforcement is filling a leaky bucket. But I guess we can do this in parallel.
Solving slop is a dangerous capability externality?
Coordinal Research was not ultimately funded.
On your first point, I think an important reason that we’re making so much progress in mathematics as opposed to AI safety is that AI safety is not well formalized yet. It is impressive to see how much rigorous and intelligent Fable and Astra becomes when they have access to a Lean MCP. If you want to scale up alignment with automatization this way, I believe the most important thing is to formalize alignment directly, as in, allow the agent to have a test to immediately and empirically test its idea. The full blown version of this is Davidad’s infrabayesianism and hypersimulations, but we can also do this on a shorter scale, with e.g. a framework that allows Claude/Codex to specify the behavior they want to observe, and run many simulations of it.
I already have a good meta-plugin to write such MCP, so I think I could do this in the span of a week, if anyone is interested to work on this, let me know.
I agree with the sentiment of this post, but I think it is factually incorrect about Coordinal Research activities and as such I strongly downvoted.
I don’t believe I or Coordinal has anywhere stated or claimed that we have “curated 400+ open safety research questions”. I think that there are likely many groups/organizations/frontier-company-projects that include datasets like this; I am aware of a few personally. Coordinal never did this.
I worry that this failure was the result of poor agent/LLM use and oversight. It’s possible I am wrong and Coordinal has made this claim in the past, could you cite your source for this claim?
Thanks for commenting. I edited the post.
Alignment progress could be seen as contributing to capabilities work (see here and here) which is why I’m personally more excited about progress on broader questions (better information tech for coordination, robust institution design, etc) that target the meta-problem of humanity failing to effectively coordinate towards shared goals—though this is of course far more difficult (nearly impossible?) to formalize and thus harder for present day AIs to make progress on.
Not disagreeing; I just think the community needs extensive discussion/thinking on the net-impact of alignment work before deciding to throw more resources at it (whereas something like discovering/demonstrating misalignment seems more robustly beneficial).