Plan A scales to top-expert-level AI in ~5 years, then to superintelligence once very high confidence in alignment is achieved; in the scenario this is 5 more years later, for a total of 10 years from the deal beginning until superintelligence. (More on the scaling strategy in https://ai-2040.com/supplements/capability-scaling-strategy.)
The version of Plan S that we’re most sympathetic to involves a halt on AI capabilities for a minimum of ~3-5 years with intent to eventually scale to top-expert-level AI and then superintelligence, but much slower than in Plan A.
(The below is all my own view, other authors may disagree on the details.)
I think Plan S is a big improvement over Plans D, C, or B but worse than Plan A.
The main reason that I prefer Plan A to Plan S is that I think the risk of the international slowdown/pause deal declining is very significant: i.e. either dissolving or having its effectiveness becoming highly impaired. I think that the risk is roughly 35-40% within 5 years, with high error bars (and with deal dissolution and deal impairment contributing roughly equally to that tootal). More on this in https://ai-2040.com/supplements/deal-decline.
Given a fixed amount of time bought, it’s better to scale capabilities as long as this can be done with high confidence in safety, in order to get useful work out of AIs and study AIs that are closer to being able to take over. If the deal dissolves 3 years into Plan A, you’ve used improve AIs to achieve lots of useful alignment research, decision-making/epistemics improvements, etc. If it dissolves 3 years into Plan S, you’re better off than in Plan D because of the increased time you’ve bought, but you’ve gotten much less useful work out of your AIs.
Another key questions is whether the chance of deal decline is higher in Plan A or S: my guess is that it’s higher in Plan S because better AIs allow you to develop technologies that stabilize the deal, though I’m not confident as the AI progress could be destabilizing in other ways; if we went all out Plan-D-style that would likely be more destabilizing to the deal than Plan A even pre-TED-AI, so faster isn’t always better.
The main upside of Plan S is that the initial pause phase is simpler than Plan A and thus harder to mess up; in particular, in Plan A the risk of catastrophe due to a misjudgment that led to scaling too fast is higher probability. If you want to eventually resume scaling, you will need to transition to a regime that can handle this. But the extra time you’ve bought and the slower pace at which you might scale should help with managing this risk.
If the risk of deal decline were very low, this would provide more reason to do Plan S instead of A and I’d think that they were pretty close in value with Plan S potentially being better. Even then, I’d think that Plan S should aim to scale to superintelligence eventually and probably within around 30-100 years, because there are other reasons to scale at some point besides deal decline: background risks such as pandemics and nuclear war, and covert projects.
I think it’s much easier to achieve Plan S than Plan A.
Plan S is simple “Don’t train new AI until there’s very strong consensus it’s a good idea.”
Plan A requires continuously making nuanced judgment calls about what counts as safe, while generally maintaining momentum on training more powerful AIs (and leaving lots of dry tinder around). It’s essentially an unsolved problem to make regulations careful enough to distinguish good vs bad safety cases, and I don’t think the Plan A documents had particularly good ideas for now to do so.
I think Plan A basically only makes sense if we get pleasantly surprisingly good governance (which would be a marked departure from the governance we currently seem on track to have).
We’ve seen examples of Plan S for cloning/eugenics and (sorta) for nuclear power, so it’s not like it’s obviously intractable.
I realize Plan A / Plan S is a spectrum. But one of the main things I’d want to see to feel safer is interrupting the momentum of the AI labs and transitioning the world to “training a new AI is treated as a dangerous, careful endeavor.” I think this requires multiple years (vs the approximately 1 year pause in Plan A).
Some things that’d update me include seeing a) a significant reduction in US gov corruption after the 2028 election, one way or another, b) seeing “AI for epistemics” actually begin to play out.
I don’t think AI for epistemics is actually that bottlenecked on better AIs (community notes didn’t require LLMs at all). I think it’s more just “actually bothering to design social media and other infrastructure for epistemics at all.”)
Plan S is better at averting permanent disempowerment (it enables more chances to take this specific problem seriously), even as perhaps it’s modestly worse than Plan A at averting extinction. In the futures where Plan A doesn’t lead to extinction, it still almost certainly ends in permanent disempowerment (the future of humanity only gets breadcrumbs of the reachable universe).
(My impression is that Yudkowsky/Soares expect extinction where I expect permanent disempowerment. And I expect permanent disempowerment to be likely averted in the futures where they expect extinction to be averted.)
I think Plan A is significantly worse than Plan S.
I am uncertain that even the best versions of the sorts of control/superficial alignment techniques portrayed in the plan will be sufficient to make sure nothing catastrophic happens, when “Top-Expert-Dominating AI” is on the table.
And it does not seem to me that every time, across years and dozens of companies, the best possible versions of said alignment techniques will be implemented. It seems very plausible that something of the flavor of Anthropic and OpenAI training against the CoT will happen, except much more dangerous, since “Top-Expert-Dominating AI”.
I do not think that scaling to what Plan A portrays, will speed up the time to “alignment is solved” sufficiently to outweigh the risk; in my experience, the sorts of things current AIs are good at, or are on track to become good at, are not the limiting factor in alignment research, especially the sorts of exceptional alignment research that bring us substantially closer to “alignment is solved”.
I agree that deal breakdown is a big problem, one that people should be working on post-pause. I do not think getting closer to the edge of existentially dangerous capabilities helps with that.[1]
On Plan A vs. Plan S
Plan A scales to top-expert-level AI in ~5 years, then to superintelligence once very high confidence in alignment is achieved; in the scenario this is 5 more years later, for a total of 10 years from the deal beginning until superintelligence. (More on the scaling strategy in https://ai-2040.com/supplements/capability-scaling-strategy.)
The version of Plan S that we’re most sympathetic to involves a halt on AI capabilities for a minimum of ~3-5 years with intent to eventually scale to top-expert-level AI and then superintelligence, but much slower than in Plan A.
(The below is all my own view, other authors may disagree on the details.)
I think Plan S is a big improvement over Plans D, C, or B but worse than Plan A.
The main reason that I prefer Plan A to Plan S is that I think the risk of the international slowdown/pause deal declining is very significant: i.e. either dissolving or having its effectiveness becoming highly impaired. I think that the risk is roughly 35-40% within 5 years, with high error bars (and with deal dissolution and deal impairment contributing roughly equally to that tootal). More on this in https://ai-2040.com/supplements/deal-decline.
Given a fixed amount of time bought, it’s better to scale capabilities as long as this can be done with high confidence in safety, in order to get useful work out of AIs and study AIs that are closer to being able to take over. If the deal dissolves 3 years into Plan A, you’ve used improve AIs to achieve lots of useful alignment research, decision-making/epistemics improvements, etc. If it dissolves 3 years into Plan S, you’re better off than in Plan D because of the increased time you’ve bought, but you’ve gotten much less useful work out of your AIs.
Another key questions is whether the chance of deal decline is higher in Plan A or S: my guess is that it’s higher in Plan S because better AIs allow you to develop technologies that stabilize the deal, though I’m not confident as the AI progress could be destabilizing in other ways; if we went all out Plan-D-style that would likely be more destabilizing to the deal than Plan A even pre-TED-AI, so faster isn’t always better.
The main upside of Plan S is that the initial pause phase is simpler than Plan A and thus harder to mess up; in particular, in Plan A the risk of catastrophe due to a misjudgment that led to scaling too fast is higher probability. If you want to eventually resume scaling, you will need to transition to a regime that can handle this. But the extra time you’ve bought and the slower pace at which you might scale should help with managing this risk.
If the risk of deal decline were very low, this would provide more reason to do Plan S instead of A and I’d think that they were pretty close in value with Plan S potentially being better. Even then, I’d think that Plan S should aim to scale to superintelligence eventually and probably within around 30-100 years, because there are other reasons to scale at some point besides deal decline: background risks such as pandemics and nuclear war, and covert projects.
(crossposted from https://x.com/eli_lifland/status/2075734827832401959 with minor edits)
I think it’s much easier to achieve Plan S than Plan A.
Plan S is simple “Don’t train new AI until there’s very strong consensus it’s a good idea.”
Plan A requires continuously making nuanced judgment calls about what counts as safe, while generally maintaining momentum on training more powerful AIs (and leaving lots of dry tinder around). It’s essentially an unsolved problem to make regulations careful enough to distinguish good vs bad safety cases, and I don’t think the Plan A documents had particularly good ideas for now to do so.
I think Plan A basically only makes sense if we get pleasantly surprisingly good governance (which would be a marked departure from the governance we currently seem on track to have).
We’ve seen examples of Plan S for cloning/eugenics and (sorta) for nuclear power, so it’s not like it’s obviously intractable.
I realize Plan A / Plan S is a spectrum. But one of the main things I’d want to see to feel safer is interrupting the momentum of the AI labs and transitioning the world to “training a new AI is treated as a dangerous, careful endeavor.” I think this requires multiple years (vs the approximately 1 year pause in Plan A).
Some things that’d update me include seeing a) a significant reduction in US gov corruption after the 2028 election, one way or another, b) seeing “AI for epistemics” actually begin to play out.
I don’t think AI for epistemics is actually that bottlenecked on better AIs (community notes didn’t require LLMs at all). I think it’s more just “actually bothering to design social media and other infrastructure for epistemics at all.”)
Plan S is better at averting permanent disempowerment (it enables more chances to take this specific problem seriously), even as perhaps it’s modestly worse than Plan A at averting extinction. In the futures where Plan A doesn’t lead to extinction, it still almost certainly ends in permanent disempowerment (the future of humanity only gets breadcrumbs of the reachable universe).
(My impression is that Yudkowsky/Soares expect extinction where I expect permanent disempowerment. And I expect permanent disempowerment to be likely averted in the futures where they expect extinction to be averted.)
I think Plan A is significantly worse than Plan S.
I am uncertain that even the best versions of the sorts of control/superficial alignment techniques portrayed in the plan will be sufficient to make sure nothing catastrophic happens, when “Top-Expert-Dominating AI” is on the table.
And it does not seem to me that every time, across years and dozens of companies, the best possible versions of said alignment techniques will be implemented. It seems very plausible that something of the flavor of Anthropic and OpenAI training against the CoT will happen, except much more dangerous, since “Top-Expert-Dominating AI”.
I do not think that scaling to what Plan A portrays, will speed up the time to “alignment is solved” sufficiently to outweigh the risk; in my experience, the sorts of things current AIs are good at, or are on track to become good at, are not the limiting factor in alignment research, especially the sorts of exceptional alignment research that bring us substantially closer to “alignment is solved”.
I agree that deal breakdown is a big problem, one that people should be working on post-pause. I do not think getting closer to the edge of existentially dangerous capabilities helps with that.[1]
If anything, it might make a deal weaker, since it goes past a natural Schelling fence.