Arguments/ideas about what would make smaller models catch up to the frontier should confront what happens when you use the same improved methods for the biggest models. Usually, it doesn’t narrow the gap, because both things get stronger.
It’s not clear to me that large models getting better with CL means the adoption (not capability, as measured by centralized benchmarks) gap between them and small models will hold. First, small models will get more benefit than large models. But more importantly, if the extra capabilities of the larger model are not economically meaningful against the costs, then the market will select the smaller model. A historical analog is the application of Moore’s law to both mainframes and PCs. Both got better, but the market selected the lower-cost, right-sized, and personalized option.
First, small models will get more benefit than large models.
Why is that? Bigger models are better at making use of any given amount of data, so they might benefit more, if there’s enough compute to go around.
A historical analog is the application of Moore’s law to both mainframes and PCs. Both got better, but the market selected the lower-cost, right-sized, and personalized option.
In some ways, the modern 1 GW datacenters beg to differ.
Why is that? Bigger models are better at making use of any given amount of data, so they might benefit more, if there’s enough compute to go around.
The argument assumes that frontier models are already capable enough (“PhD level”) for most general whilte collar economic tasks that they’re being applied to, and that what’s missing is the tacit knowledge enabled by CL. Making bigger models bigger just won’t fix this blocker.
In some ways, the modern 1 GW datacenters beg to differ.
Yes, and to use the analogy, it’s because we’re in the mainframe era. Once CL is available, it would be akin to the introduction of the PC.
Mainframes never went away. They became data centers, which are even more centralized, and got bigger. The tasks that could be offloaded to small local computers were, and the ones that couldn’t, didn’t, and we kept inventing new jobs for both as both got stronger.
A mainframe is a single large very powerful machine. A datacenter is a bunch of servers (the PC model) that are conveniently colocated. I think the analogy is consistent with data centers.
GPU datacenters are different and closer to one giant machine, but that’s a recent phenomenon with centralized training of AI.
GPU datacenters are different and closer to one giant machine, but that’s a recent phenomenon with centralized training of AI.
It’s worth distinguishing big datacenters (especially for pretraining, where scale-out networks need to be unusually good) from big scale-up systems/pods. Big scale-up systems (currently mostly rack-level, though TPUs were multi-rack for many years now) are motivated by very big numbers of total params in MoE models, not by centralized training of AI. Centralized training of AI (in the sense of pretraining) instead motivates big datacenters with good scale-out networks, but the individual scale-up systems within these datacenters could be small (even for models with a lot of total params). So these are completely different desiderata, calling for technologically unrelated things.
A mainframe is a single large very powerful machine. A datacenter is a bunch of servers (the PC model) that are conveniently colocated.
That’s a good point, regarding the CPU datacenters (though I don’t see how the analogy could transfer something useful at this level, if the conclusion isn’t already accepted and the analogy just illustrates what it looks like based on a more familiar story).
Making bigger models bigger just won’t fix this blocker.
The question was what happens to bigger vs. smaller models when we do fix this blocker for both. You can’t just apply the improved method to the smaller models, and compare that to the bigger models without the improved method (whether the bigger models in fact get trained depends on the answer to this question about the capability consequences of hypothetically training the improved bigger models regardless of their presumed economic usefulness). So the relevant thing is whether the bigger models can still make better use of any given amount of data (continual learning or not), and thus whether the gap between the smaller and the bigger models persists.
The argument assumes that frontier models are already capable enough … economic tasks
They still aren’t capable enough at whatever they aren’t in fact capable enough for. And those further things can have enormous TAM, all the way to taking over the world and the reachable universe. This is only irrelevant for the question of big vs. small models if the gap in fact disappers.
If the gap remains, while the smaller models are capable enough for most general economic tasks, then the bigger models will be even more capable than that. The crux then shifts to whether this is even possible (but then it can’t affect the argument itself, that would be rationalization from the bottom line back to the argument), or whether it’s valuable to be significantly more capable than whatever most modern economic tasks require.
You can’t just apply the improved method to the smaller models, and compare that to the bigger models without the improved method
I tried to address this but might have been missed:
It’s not clear to me that large models getting better with CL means the adoption (not capability, as measured by centralized benchmarks) gap between them and small models will hold. First, small models will get more benefit than large models. But more importantly, if the extra capabilities of the larger model are not economically meaningful against the costs, then the market will select the smaller model.
So both get a benefit, but if what was blocking economic ROI was CL, and not the extra capabilities of the larger models, then smaller models will win.
They still aren’t capable enough at whatever they aren’t in fact capable enough for. And those further things can have enormous TAM,
Yes, and I think we’ve established the “further thing” is CL (at least in the context of this argument), so then the enourmous TAM is tied to CL, not how big the model is. You would obviously want to own the weights, and you’d want to serve the smallest possible model that delivered the returns of the CL to your org.
If the gap remains, while the smaller models are capable enough for most general economic tasks, then the bigger models will be even more capable than that.
That extra capability only matters if it’s economically useful. Like I said before (and analogized with the mainframe vs. PC argument), the model being able to solve math conjectures is not useful for the white collar work that has repeatedly resisted automation with AI. More of that will not change the equilibrium.
That extra capability only matters if it’s economically useful
I’m not even insisting that the extra capabilities matter. I’m insisting that you didn’t argue that they aren’t there (capabilities of big over small models, given CL or whatever other improvements). Or that the capability (rather than adoption) gap gets smaller. Whether the bigger models are economically useful is downstream of whether they’re importantly more capable, the question of relative capability is a key input to the outcome of adoption, while the question of adoption doesn’t inform the question of relative capability at all.
That extra capability only matters if it’s economically useful. Like I said before (and analogized with the mainframe vs. PC argument), the model being able to solve math conjectures is not useful for the white collar work that has repeatedly resisted automation with AI. More of that will not change the equilibrium.
That’s why I mentioned the reachable universe. Some amount of ASI-pilledness is necessary for a reasonable discussion about what happens when the bigger models saturate the status quo level of capabilities of the modern humanity.
Yes, and I think we’ve established the “further thing” is CL
What I said was
They still aren’t capable enough at whatever they aren’t in fact capable enough for. And those further things can have enormous TAM
The “further things” in my intended meaning are capabilities (and the accomplishment of the more difficult tasks), not methods.
It’s not clear to me that large models getting better with CL means the adoption (not capability, as measured by centralized benchmarks) gap between them and small models will hold. First, small models will get more benefit than large models. But more importantly, if the extra capabilities of the larger model are not economically meaningful against the costs, then the market will select the smaller model. A historical analog is the application of Moore’s law to both mainframes and PCs. Both got better, but the market selected the lower-cost, right-sized, and personalized option.
Why is that? Bigger models are better at making use of any given amount of data, so they might benefit more, if there’s enough compute to go around.
In some ways, the modern 1 GW datacenters beg to differ.
The argument assumes that frontier models are already capable enough (“PhD level”) for most general whilte collar economic tasks that they’re being applied to, and that what’s missing is the tacit knowledge enabled by CL. Making bigger models bigger just won’t fix this blocker.
Yes, and to use the analogy, it’s because we’re in the mainframe era. Once CL is available, it would be akin to the introduction of the PC.
Mainframes never went away. They became data centers, which are even more centralized, and got bigger. The tasks that could be offloaded to small local computers were, and the ones that couldn’t, didn’t, and we kept inventing new jobs for both as both got stronger.
So far I see AI following the same path.
A mainframe is a single large very powerful machine. A datacenter is a bunch of servers (the PC model) that are conveniently colocated. I think the analogy is consistent with data centers.
GPU datacenters are different and closer to one giant machine, but that’s a recent phenomenon with centralized training of AI.
It’s worth distinguishing big datacenters (especially for pretraining, where scale-out networks need to be unusually good) from big scale-up systems/pods. Big scale-up systems (currently mostly rack-level, though TPUs were multi-rack for many years now) are motivated by very big numbers of total params in MoE models, not by centralized training of AI. Centralized training of AI (in the sense of pretraining) instead motivates big datacenters with good scale-out networks, but the individual scale-up systems within these datacenters could be small (even for models with a lot of total params). So these are completely different desiderata, calling for technologically unrelated things.
That’s a good point, regarding the CPU datacenters (though I don’t see how the analogy could transfer something useful at this level, if the conclusion isn’t already accepted and the analogy just illustrates what it looks like based on a more familiar story).
The question was what happens to bigger vs. smaller models when we do fix this blocker for both. You can’t just apply the improved method to the smaller models, and compare that to the bigger models without the improved method (whether the bigger models in fact get trained depends on the answer to this question about the capability consequences of hypothetically training the improved bigger models regardless of their presumed economic usefulness). So the relevant thing is whether the bigger models can still make better use of any given amount of data (continual learning or not), and thus whether the gap between the smaller and the bigger models persists.
They still aren’t capable enough at whatever they aren’t in fact capable enough for. And those further things can have enormous TAM, all the way to taking over the world and the reachable universe. This is only irrelevant for the question of big vs. small models if the gap in fact disappers.
If the gap remains, while the smaller models are capable enough for most general economic tasks, then the bigger models will be even more capable than that. The crux then shifts to whether this is even possible (but then it can’t affect the argument itself, that would be rationalization from the bottom line back to the argument), or whether it’s valuable to be significantly more capable than whatever most modern economic tasks require.
I tried to address this but might have been missed:
So both get a benefit, but if what was blocking economic ROI was CL, and not the extra capabilities of the larger models, then smaller models will win.
Yes, and I think we’ve established the “further thing” is CL (at least in the context of this argument), so then the enourmous TAM is tied to CL, not how big the model is. You would obviously want to own the weights, and you’d want to serve the smallest possible model that delivered the returns of the CL to your org.
That extra capability only matters if it’s economically useful. Like I said before (and analogized with the mainframe vs. PC argument), the model being able to solve math conjectures is not useful for the white collar work that has repeatedly resisted automation with AI. More of that will not change the equilibrium.
I’m not even insisting that the extra capabilities matter. I’m insisting that you didn’t argue that they aren’t there (capabilities of big over small models, given CL or whatever other improvements). Or that the capability (rather than adoption) gap gets smaller. Whether the bigger models are economically useful is downstream of whether they’re importantly more capable, the question of relative capability is a key input to the outcome of adoption, while the question of adoption doesn’t inform the question of relative capability at all.
That’s why I mentioned the reachable universe. Some amount of ASI-pilledness is necessary for a reasonable discussion about what happens when the bigger models saturate the status quo level of capabilities of the modern humanity.
What I said was
The “further things” in my intended meaning are capabilities (and the accomplishment of the more difficult tasks), not methods.