I’m worried that you’re missing something important because you mostly argue against large AI R&D multipliers, but you don’t spend much time directly referencing compute bottlenecks in your arguments that the forecast is too aggressive.
Consider the case of doing pure math research (which we’ll assume for simplicity doesn’t benefit from compute at all). If we made emulated versions of the 1000 best math researchers and then we made 1 billion copies of each of them them which all ran at 1000x speed, I expect we’d get >1000x faster progress. As far as I can tell, the words in your arguments don’t particularly apply less to this situation than the AI R&D situation.
Going through the object level response for each of these arguments in the case of pure math research and the correspondence to the AI R&D:
Simplified Model of AI R&D
Math: Yes, there are many tasks in math R&D, but the 1000 best math researchers could already do them or learn to do them.
AI R&D: By the time you have SAR (superhuman AI researcher), we’re assuming the AIs are better than the best human researchers(!), so heterogenous tasks don’t matter if you accept the premise of SAR: whatever the humans could have done, the AIs can do better. It does apply to the speed ups at superhuman coders, but I’m not sure this will make a huge difference to the bottom line (and you seem to mostly be referencing later speed ups).
Amdahl’s Law
Math: The speed up is near universal because we can do whatever the humans could do.
AI R&D: Again, the SAR is strictly better than humans, so hard-to-automate activities aren’t a problem. When we’re talking about ~1000x speed up, the authors are imagining AIs which are much smarter than humans at everything and which are running 100x faster than humans at immense scale. So, “hard to automate tasks” is also not relevant.
All this said, compute bottlenecks could be very important here! But the bottlenecking argument must directly reference these compute bottlenecks and there has to be no way to route around this. My sense is that much better research taste and perfect implementation could make experiments with some fixed amount of compute >100x more useful. To me, this feels like the important question: how much can labor results in routing around compute bottlenecks and utilizing compute much more effectively. The naive extrapolation out of the human range makes this look quite aggressive: the median AI company employee is probably 10x worse at using compute than the best, so an AI which as superhuman as 2x the gap between median and best would naively be 100x better at using compute than the best employee. (Is the research taste ceiling plausibly this high? I currently think extrapolating out another 100x is reasonable given that we don’t see things slowing down in the human range as far as we can tell.)
Dependence on Narrow Data Sets
This is only applicable to the timeline to the superhuman coder milestone, not to takeoff speeds once we have a superhuman coder. (Or maybe you think a similar argument applies to the time between superhuman coder and SAR.)
Hofstadter’s Law As Prior
Math: We’re talking about speed up relative to what the human researchers would have done by default, so this just divides both sides equally and cancels out.
AI R&D: The should also just divide both sides. That said, Hofstadter’s Law does apply to the human-only, software-only times between milestones. But note that these times are actually quite long! (Maybe you think they are still too short, in which case fair enough.)
I’ve (briefly) addressed the compute bottleneck question on a different comment branch, and “hard-to-automate activities aren’t a problem” on another (confusion regarding the definition of various milestones).
[Dependence on Narrow Data Sets] is only applicable to the timeline to the superhuman coder milestone, not to takeoff speeds once we have a superhuman coder. (Or maybe you think a similar argument applies to the time between superhuman coder and SAR.)
I do think it applies, if indirectly. Most data relating to progress in AI capabilities comes from benchmarks of crisply encapsulated tasks. I worry this may skew our collective intuitions regarding progress toward broader capabilities, especially as I haven’t seen much attention paid to exploring the delta between things we currently benchmark and “everything”.
Hofstadter’s Law As Prior
Math: We’re talking about speed up relative to what the human researchers would have done by default, so this just divides both sides equally and cancels out.
This feels like one of those “the difference between theory and practice is smaller in theory than in practice” situations… Hofstadter’s Law would imply that Hofstadter’s Law applies here. :-)
For one concrete example of how that could manifest, perhaps there is a delay between “AI models exist that are superhuman at all activities involved in developing better models” and “those models have been fully adopted across the organization”. Interior to a frontier lab, that specific delay might be immaterial, it’s just meant as an existence proof that there’s room for us to be missing things.
I’m worried that you’re missing something important because you mostly argue against large AI R&D multipliers, but you don’t spend much time directly referencing compute bottlenecks in your arguments that the forecast is too aggressive.
Consider the case of doing pure math research (which we’ll assume for simplicity doesn’t benefit from compute at all). If we made emulated versions of the 1000 best math researchers and then we made 1 billion copies of each of them them which all ran at 1000x speed, I expect we’d get >1000x faster progress. As far as I can tell, the words in your arguments don’t particularly apply less to this situation than the AI R&D situation.
Going through the object level response for each of these arguments in the case of pure math research and the correspondence to the AI R&D:
Math: Yes, there are many tasks in math R&D, but the 1000 best math researchers could already do them or learn to do them.
AI R&D: By the time you have SAR (superhuman AI researcher), we’re assuming the AIs are better than the best human researchers(!), so heterogenous tasks don’t matter if you accept the premise of SAR: whatever the humans could have done, the AIs can do better. It does apply to the speed ups at superhuman coders, but I’m not sure this will make a huge difference to the bottom line (and you seem to mostly be referencing later speed ups).
Math: The speed up is near universal because we can do whatever the humans could do.
AI R&D: Again, the SAR is strictly better than humans, so hard-to-automate activities aren’t a problem. When we’re talking about ~1000x speed up, the authors are imagining AIs which are much smarter than humans at everything and which are running 100x faster than humans at immense scale. So, “hard to automate tasks” is also not relevant.
All this said, compute bottlenecks could be very important here! But the bottlenecking argument must directly reference these compute bottlenecks and there has to be no way to route around this. My sense is that much better research taste and perfect implementation could make experiments with some fixed amount of compute >100x more useful. To me, this feels like the important question: how much can labor results in routing around compute bottlenecks and utilizing compute much more effectively. The naive extrapolation out of the human range makes this look quite aggressive: the median AI company employee is probably 10x worse at using compute than the best, so an AI which as superhuman as 2x the gap between median and best would naively be 100x better at using compute than the best employee. (Is the research taste ceiling plausibly this high? I currently think extrapolating out another 100x is reasonable given that we don’t see things slowing down in the human range as far as we can tell.)
This is only applicable to the timeline to the superhuman coder milestone, not to takeoff speeds once we have a superhuman coder. (Or maybe you think a similar argument applies to the time between superhuman coder and SAR.)
Math: We’re talking about speed up relative to what the human researchers would have done by default, so this just divides both sides equally and cancels out.
AI R&D: The should also just divide both sides. That said, Hofstadter’s Law does apply to the human-only, software-only times between milestones. But note that these times are actually quite long! (Maybe you think they are still too short, in which case fair enough.)
I’ve (briefly) addressed the compute bottleneck question on a different comment branch, and “hard-to-automate activities aren’t a problem” on another (confusion regarding the definition of various milestones).
I do think it applies, if indirectly. Most data relating to progress in AI capabilities comes from benchmarks of crisply encapsulated tasks. I worry this may skew our collective intuitions regarding progress toward broader capabilities, especially as I haven’t seen much attention paid to exploring the delta between things we currently benchmark and “everything”.
This feels like one of those “the difference between theory and practice is smaller in theory than in practice” situations… Hofstadter’s Law would imply that Hofstadter’s Law applies here. :-)
For one concrete example of how that could manifest, perhaps there is a delay between “AI models exist that are superhuman at all activities involved in developing better models” and “those models have been fully adopted across the organization”. Interior to a frontier lab, that specific delay might be immaterial, it’s just meant as an existence proof that there’s room for us to be missing things.