This is a much better statement of the AI risk problem than I’ve often given. It resolves some qualms that I’ve had since trying to explain why any AI would be motivated to take over the world in a recent talk.
Generally, there’s a bit of a leap in the old LessWrong arguments that any powerful AI would exhibit instrumental convergence—that claim depends on a kind of specific conception of what AI is or has to be, which Eliezer persuasively argues for, but which isn’t obvious, and is hard to convey to a layman. It is much less of a leap to say “any AI of this type, would try to take over the world if it could” and “notice that since LLM base models were invented, every single capability advance entailed making the AIs more like the scary version.” This seems to me like a more intellectually honest framing of the problem. Maybe not all powerful AIs we could invent are like this, but the ones we’re building in practice seem to be!
I also like that this frames the problem, appropriately, not as some speculative thing that might happen, but as a thing we’ve seen over and over for decades, and that we expect to keep happening. It’s just that as it occurs in more and more powerful and empowered AIs, we get more and more harmful versions of the problem.
This is a much better statement of the AI risk problem than I’ve often given. It resolves some qualms that I’ve had since trying to explain why any AI would be motivated to take over the world in a recent talk.
Generally, there’s a bit of a leap in the old LessWrong arguments that any powerful AI would exhibit instrumental convergence—that claim depends on a kind of specific conception of what AI is or has to be, which Eliezer persuasively argues for, but which isn’t obvious, and is hard to convey to a layman. It is much less of a leap to say “any AI of this type, would try to take over the world if it could” and “notice that since LLM base models were invented, every single capability advance entailed making the AIs more like the scary version.” This seems to me like a more intellectually honest framing of the problem. Maybe not all powerful AIs we could invent are like this, but the ones we’re building in practice seem to be!
I also like that this frames the problem, appropriately, not as some speculative thing that might happen, but as a thing we’ve seen over and over for decades, and that we expect to keep happening. It’s just that as it occurs in more and more powerful and empowered AIs, we get more and more harmful versions of the problem.