I agree with the rest of your post but I have pretty high error bars on how hard AI research would be.
The current AI paradigm seems poised to make models superhumanly capable at hacking and formal math way before they become even upper-human-level at something more dangerous, like long-term strategy or open-ended innovation.
I think this is largely the crux yeah. I think exactly the current paradigm probably doesn’t get you to RSI because 2026-levels of reward hacking might just be too high for fully automated AI research, but a) I have somewhat high error bars and b) I think you can probably patch this locally well enough for AI research without fundamental breakthroughs in alignment.
After you patch the local issues, is significant automation of AI research in the cards? I’m really unsure. It’s plausible to me that you need deeper sources of insight/wisdom like you say, but it’s also plausible to me that AI research is conceptually somewhat “easy”, and ultimately mostly a number-go-up game that the AI researcher models can hill-climb on.
Sure, but not all kinds of AI research are alike. I think it’s obvious that LLMs are going to be able to automate some kinds of nontrivial AI research, or are already doing it. But is the kind of AI research that is automable without further innovation the kind of AI research that allows to significantly improve at those less-tangible skills? Probably not centrally so; probably the kind of AI research that is most automable is the kind of research that is most easily verifiable (as you say: that is easy to reduce to number-go-up).
Which is to say, plausibly this will just differentially accelerate math/hacking even further, making the capability frontier even more jagged in the same ways it’s already jagged.
I agree with the rest of your post but I have pretty high error bars on how hard AI research would be.
I think this is largely the crux yeah. I think exactly the current paradigm probably doesn’t get you to RSI because 2026-levels of reward hacking might just be too high for fully automated AI research, but a) I have somewhat high error bars and b) I think you can probably patch this locally well enough for AI research without fundamental breakthroughs in alignment.
After you patch the local issues, is significant automation of AI research in the cards? I’m really unsure. It’s plausible to me that you need deeper sources of insight/wisdom like you say, but it’s also plausible to me that AI research is conceptually somewhat “easy”, and ultimately mostly a number-go-up game that the AI researcher models can hill-climb on.
Sure, but not all kinds of AI research are alike. I think it’s obvious that LLMs are going to be able to automate some kinds of nontrivial AI research, or are already doing it. But is the kind of AI research that is automable without further innovation the kind of AI research that allows to significantly improve at those less-tangible skills? Probably not centrally so; probably the kind of AI research that is most automable is the kind of research that is most easily verifiable (as you say: that is easy to reduce to number-go-up).
Which is to say, plausibly this will just differentially accelerate math/hacking even further, making the capability frontier even more jagged in the same ways it’s already jagged.