The key argument against the superalignment/automated alignment agenda is that while AIs will excel in verifiable domains, such as code, they will struggle with hard-to-verify tasks.
For example, science in domains we have little data (alignment of superintelligence) and techniques that work for weaker models will be poor proxies and break at superintelligence (i.e. harder to monitor, internal reasoning, models are no longer stateless and are continually learning, tangibly different reasoning than the weak reasoning that currently exists, etc).
Ultimately, you get convincing slop, and even though you might catch non-superintelligent AIs doing so-called “scheming”, it’s not that helpful because they are not capable enough to cause a catastrophe at this point.
The crux is whether AIs end up capable of +10x-ing actually useful superalignment research while you are in the valley of life, which is when you can quickly verify outputs are not slop (no longer severely bottlenecked on human talent; after the slop era), but before all your control techniques are basically doomed.
So, you hope to prevent AIs from sabotaging AI safety research AND that the resulting safety research isn’t just a poor proxy that works well at a specific model size/shape, but then completely fails when you have self-modifying superintelligence.
Ultimately, you’d better have a backup plan for superalignment that isn’t just, “we’ll stop if we catch the AIs being deceptively aligned.” There are worlds where everything seems plausibly safe, you have a very convincing, vetted safety plan, you implement it, and you die.
That said, I am a little bit confused by folks who both say, “current AI models have nothing to do with future powerful (real) AIs” yet also consistently use “bad” behaviour from current AIs as a reason to stop.
Often, the argument made is, “we don’t even understand the previous generations of AIs, how do we even hope to align future AIs?”
I guess the way I understand it is that given that we can’t even get current AIs to do exactly what we want, then we should expect the same for future AIs. However, this feels somewhat connected to the fact that current AIs are just sloppy and lack the capability, not only some thing about “we don’t know how to align current models perfectly to our intentions.”
The issue is that they are getting better at making the slop convincing, and in the predicted ways—ways that got reward in training due to under-specified goals. The canonical example is Claude Code’s tendency to delete tests, or make tests pass by mocking the part that we wanted to check.
Ultimately, you’d better have a backup plan for superalignment that isn’t just, “we’ll stop if we catch the AIs being deceptively aligned.” There are worlds where everything seems plausibly safe, you have a very convincing, vetted safety plan, you implement it, and you die.
An alignment plan that ignores “normalization of deviance” seems likely to fail in practice, in much the same way the OpenAI’s non-profit oversight has struggled to resist the temptations of success. Human systems fail in some very predictable ways.
The key argument against the superalignment/automated alignment agenda is that while AIs will excel in verifiable domains, such as code, they will struggle with hard-to-verify tasks.
For example, science in domains we have little data (alignment of superintelligence) and techniques that work for weaker models will be poor proxies and break at superintelligence (i.e. harder to monitor, internal reasoning, models are no longer stateless and are continually learning, tangibly different reasoning than the weak reasoning that currently exists, etc).
Ultimately, you get convincing slop, and even though you might catch non-superintelligent AIs doing so-called “scheming”, it’s not that helpful because they are not capable enough to cause a catastrophe at this point.
The crux is whether AIs end up capable of +10x-ing actually useful superalignment research while you are in the valley of life, which is when you can quickly verify outputs are not slop (no longer severely bottlenecked on human talent; after the slop era), but before all your control techniques are basically doomed.
So, you hope to prevent AIs from sabotaging AI safety research AND that the resulting safety research isn’t just a poor proxy that works well at a specific model size/shape, but then completely fails when you have self-modifying superintelligence.
Ultimately, you’d better have a backup plan for superalignment that isn’t just, “we’ll stop if we catch the AIs being deceptively aligned.” There are worlds where everything seems plausibly safe, you have a very convincing, vetted safety plan, you implement it, and you die.
That said, I am a little bit confused by folks who both say, “current AI models have nothing to do with future powerful (real) AIs” yet also consistently use “bad” behaviour from current AIs as a reason to stop.
Often, the argument made is, “we don’t even understand the previous generations of AIs, how do we even hope to align future AIs?”
I guess the way I understand it is that given that we can’t even get current AIs to do exactly what we want, then we should expect the same for future AIs. However, this feels somewhat connected to the fact that current AIs are just sloppy and lack the capability, not only some thing about “we don’t know how to align current models perfectly to our intentions.”
The issue is that they are getting better at making the slop convincing, and in the predicted ways—ways that got reward in training due to under-specified goals. The canonical example is Claude Code’s tendency to delete tests, or make tests pass by mocking the part that we wanted to check.
There are also worlds where we we get many, many warnings that we are not in control, and we make excuses, and we plunge ahead anyway. This is classic “normalization of deviance”, and it’s a notorious problem in aviation safety.
An alignment plan that ignores “normalization of deviance” seems likely to fail in practice, in much the same way the OpenAI’s non-profit oversight has struggled to resist the temptations of success. Human systems fail in some very predictable ways.