You can’t just apply the improved method to the smaller models, and compare that to the bigger models without the improved method
I tried to address this but might have been missed:
It’s not clear to me that large models getting better with CL means the adoption (not capability, as measured by centralized benchmarks) gap between them and small models will hold. First, small models will get more benefit than large models. But more importantly, if the extra capabilities of the larger model are not economically meaningful against the costs, then the market will select the smaller model.
So both get a benefit, but if what was blocking economic ROI was CL, and not the extra capabilities of the larger models, then smaller models will win.
They still aren’t capable enough at whatever they aren’t in fact capable enough for. And those further things can have enormous TAM,
Yes, and I think we’ve established the “further thing” is CL (at least in the context of this argument), so then the enourmous TAM is tied to CL, not how big the model is. You would obviously want to own the weights, and you’d want to serve the smallest possible model that delivered the returns of the CL to your org.
If the gap remains, while the smaller models are capable enough for most general economic tasks, then the bigger models will be even more capable than that.
That extra capability only matters if it’s economically useful. Like I said before (and analogized with the mainframe vs. PC argument), the model being able to solve math conjectures is not useful for the white collar work that has repeatedly resisted automation with AI. More of that will not change the equilibrium.
That extra capability only matters if it’s economically useful
I’m not even insisting that the extra capabilities matter. I’m insisting that you didn’t argue that they aren’t there (capabilities of big over small models, given CL or whatever other improvements). Or that the capability (rather than adoption) gap gets smaller. Whether the bigger models are economically useful is downstream of whether they’re importantly more capable, the question of relative capability is a key input to the outcome of adoption, while the question of adoption doesn’t inform the question of relative capability at all.
That extra capability only matters if it’s economically useful. Like I said before (and analogized with the mainframe vs. PC argument), the model being able to solve math conjectures is not useful for the white collar work that has repeatedly resisted automation with AI. More of that will not change the equilibrium.
That’s why I mentioned the reachable universe. Some amount of ASI-pilledness is necessary for a reasonable discussion about what happens when the bigger models saturate the status quo level of capabilities of the modern humanity.
Yes, and I think we’ve established the “further thing” is CL
What I said was
They still aren’t capable enough at whatever they aren’t in fact capable enough for. And those further things can have enormous TAM
The “further things” in my intended meaning are capabilities (and the accomplishment of the more difficult tasks), not methods.
I tried to address this but might have been missed:
So both get a benefit, but if what was blocking economic ROI was CL, and not the extra capabilities of the larger models, then smaller models will win.
Yes, and I think we’ve established the “further thing” is CL (at least in the context of this argument), so then the enourmous TAM is tied to CL, not how big the model is. You would obviously want to own the weights, and you’d want to serve the smallest possible model that delivered the returns of the CL to your org.
That extra capability only matters if it’s economically useful. Like I said before (and analogized with the mainframe vs. PC argument), the model being able to solve math conjectures is not useful for the white collar work that has repeatedly resisted automation with AI. More of that will not change the equilibrium.
I’m not even insisting that the extra capabilities matter. I’m insisting that you didn’t argue that they aren’t there (capabilities of big over small models, given CL or whatever other improvements). Or that the capability (rather than adoption) gap gets smaller. Whether the bigger models are economically useful is downstream of whether they’re importantly more capable, the question of relative capability is a key input to the outcome of adoption, while the question of adoption doesn’t inform the question of relative capability at all.
That’s why I mentioned the reachable universe. Some amount of ASI-pilledness is necessary for a reasonable discussion about what happens when the bigger models saturate the status quo level of capabilities of the modern humanity.
What I said was
The “further things” in my intended meaning are capabilities (and the accomplishment of the more difficult tasks), not methods.