There’s an argument against slowing down AI capabilities in order to focus on alignment, which goes “we have historically made technologies safe by iterating on them; we need a system that’s representative of the system we’re trying to align or else we won’t know what we’re trying to do”.
One argument against this, which you can see in today’s ACX post goes something like “AI systems are sufficiently complex/agentic/capable of scheming/potentially destructive that we need a much more complete science in order to deal with them than we do for past technologies”.
This is true, but the original argument would fail even if this were false. There’s a good chance that significant harms come from systems that are extremely similar to present ones, and present ones are poorly enough understood that we could likely spend years doing research to understand them better, and have much of it generalize to dangerous systems.
(This has been true as long as it’s been true that “there are significant empirical things that we don’t understand about contemporary AI systems”, which is like 4-12 years depending on how you count. (Of course 6-12 years ago you’d have missed significant information about the current paradigm, leaving only very general and abstract research like agent foundations and theoretical ML stuff. However, it’s also fair to say that we would love to have made more progress on those things at this point in time. Regardless of your stance towards more theoretical AIS work, the same goes for a bunch of things which today would be called “empirical alignment”, so the statement is certainly true now.)
There’s an argument against slowing down AI capabilities in order to focus on alignment, which goes “we have historically made technologies safe by iterating on them; we need a system that’s representative of the system we’re trying to align or else we won’t know what we’re trying to do”.
One argument against this, which you can see in today’s ACX post goes something like “AI systems are sufficiently complex/agentic/capable of scheming/potentially destructive that we need a much more complete science in order to deal with them than we do for past technologies”.
This is true, but the original argument would fail even if this were false. There’s a good chance that significant harms come from systems that are extremely similar to present ones, and present ones are poorly enough understood that we could likely spend years doing research to understand them better, and have much of it generalize to dangerous systems.
(This has been true as long as it’s been true that “there are significant empirical things that we don’t understand about contemporary AI systems”, which is like 4-12 years depending on how you count. (Of course 6-12 years ago you’d have missed significant information about the current paradigm, leaving only very general and abstract research like agent foundations and theoretical ML stuff. However, it’s also fair to say that we would love to have made more progress on those things at this point in time. Regardless of your stance towards more theoretical AIS work, the same goes for a bunch of things which today would be called “empirical alignment”, so the statement is certainly true now.)