if you’re a SMB and need a slack AI to manage customer service, you don’t need that AI to also know organic chemistry. That it does today is an economic abberation simmilar to the adoption of mainframes before the PC came around.
Technical details of token costs say it’s active params that determine costs, not total params (when running within sufficiently big scale-up systems). Knowing more things (as opposed to being smarter about them) is more a total params thing (and currently there is no way to avoid knowing specific things; you can only avoid putting the effort into getting even better at them than you would by default). If there is some niche application that benefits from knowing organic chemistry, and the AI company puts in the work to get some model good at it, then there is no reason not to make all the models it makes good at it, to the extent they are able to stay coherent about sufficiently complicated questions like that (given the number of active params their cost permits them to have).
So there is no technical reason for there being an economic benefit from knowing less, if an AI company already knows how to teach AIs those things. Even cheap models with few active params can receive all the training in organic chemistry, they might just remain unable to get very good at it, but judging by the current capabilities even small cheap models will still be better than non-specialist humans. Future bigger scale-up systems/pods will enable more total params, to the point that in mid-2030s there might be “small” quadrillion param models that are very cheap (if very high levels of sparsity are technically feasible, which they might be with an absurd total number of reasonably sized routed experts).
Currently, models with fewer total params are useful because there are many legacy servers that can’t run the big models at all, not because they are inherently cheaper to run on the new bigger servers (which in 2026 are mostly rack-level scale-up systems like GB300 NVL72, but the future is multi-rack). This consideration will eventually go away, because bigger servers/pods are made out of the same stuff, the technology for connecting them into a single whole (scale-up system) just didn’t catch up to the demands of frontier LLMs yet, and therefore it’s not available for cheap models with few active params either, which is why cheap models with few active params also tend to have fewer total params.
So there is no technical reason for there being an economic benefit from knowing less
I think you’re still assuming there’s this one megacompany that’s training a single model distributed to the whole world. A small model can be trained by a specialized company that’s a good enough bootstrap based for a basket of skills such that CL can then specialize it for a user environment. If that small model can take market share for applications that require that basket of skills away from the mega AI corp, then you have the PC/mainframe dynamic again where the mainframe play gets completely disrupted.
My technical point is that “a small model” equivocates between active and total params, and a small-meaning-cheap model might have a quadrillion of total params in mid-2030s. Even with a few trillion total params, it can know all the things. Small-meaning-cheap 50B-100B active param models with trillions of total params are already happening; though such models are not yet the cheapest models, because GB300 NVL72 is not yet the low-end kind of hardware. So saying that a customer service model knowing organic chemistry is an abberation that will go away needs the premise that the company training the customer service model has difficulty obtaining the training data/process for organic chemistry, or vice versa, because there is no downside at the unit economics stage of the model’s lifecycle.
Because of the non-negative (and probably positive) transfer, meaning knowing more in one domain doesn’t hurt model performance in the other domains, and no effect on unit costs because of the active/total params distinction, it’s useful to make models that know all of the things simultaneously, for any target price tier. So there is a natural pressure for AI companies that specialize in different domains (to the extent such a thing happens at all) to merge, or to never seriously develop separately past a startup-acquisition phase (or beyond data-sourcing activities, selling the data to the companies actually training the models). This is a consequence of the unusual technological properties of MoE models, which don’t hold for many other kinds of products. If anything, CL makes it easier for a poorly focused AI megacorp to enter niche domains, which it doesn’t then need to know in subtle detail, letting its AIs get to know these domains on their own instead.
Technical details of token costs say it’s active params that determine costs, not total params (when running within sufficiently big scale-up systems). Knowing more things (as opposed to being smarter about them) is more a total params thing (and currently there is no way to avoid knowing specific things; you can only avoid putting the effort into getting even better at them than you would by default). If there is some niche application that benefits from knowing organic chemistry, and the AI company puts in the work to get some model good at it, then there is no reason not to make all the models it makes good at it, to the extent they are able to stay coherent about sufficiently complicated questions like that (given the number of active params their cost permits them to have).
So there is no technical reason for there being an economic benefit from knowing less, if an AI company already knows how to teach AIs those things. Even cheap models with few active params can receive all the training in organic chemistry, they might just remain unable to get very good at it, but judging by the current capabilities even small cheap models will still be better than non-specialist humans. Future bigger scale-up systems/pods will enable more total params, to the point that in mid-2030s there might be “small” quadrillion param models that are very cheap (if very high levels of sparsity are technically feasible, which they might be with an absurd total number of reasonably sized routed experts).
Currently, models with fewer total params are useful because there are many legacy servers that can’t run the big models at all, not because they are inherently cheaper to run on the new bigger servers (which in 2026 are mostly rack-level scale-up systems like GB300 NVL72, but the future is multi-rack). This consideration will eventually go away, because bigger servers/pods are made out of the same stuff, the technology for connecting them into a single whole (scale-up system) just didn’t catch up to the demands of frontier LLMs yet, and therefore it’s not available for cheap models with few active params either, which is why cheap models with few active params also tend to have fewer total params.
I think you’re still assuming there’s this one megacompany that’s training a single model distributed to the whole world. A small model can be trained by a specialized company that’s a good enough bootstrap based for a basket of skills such that CL can then specialize it for a user environment. If that small model can take market share for applications that require that basket of skills away from the mega AI corp, then you have the PC/mainframe dynamic again where the mainframe play gets completely disrupted.
My technical point is that “a small model” equivocates between active and total params, and a small-meaning-cheap model might have a quadrillion of total params in mid-2030s. Even with a few trillion total params, it can know all the things. Small-meaning-cheap 50B-100B active param models with trillions of total params are already happening; though such models are not yet the cheapest models, because GB300 NVL72 is not yet the low-end kind of hardware. So saying that a customer service model knowing organic chemistry is an abberation that will go away needs the premise that the company training the customer service model has difficulty obtaining the training data/process for organic chemistry, or vice versa, because there is no downside at the unit economics stage of the model’s lifecycle.
Because of the non-negative (and probably positive) transfer, meaning knowing more in one domain doesn’t hurt model performance in the other domains, and no effect on unit costs because of the active/total params distinction, it’s useful to make models that know all of the things simultaneously, for any target price tier. So there is a natural pressure for AI companies that specialize in different domains (to the extent such a thing happens at all) to merge, or to never seriously develop separately past a startup-acquisition phase (or beyond data-sourcing activities, selling the data to the companies actually training the models). This is a consequence of the unusual technological properties of MoE models, which don’t hold for many other kinds of products. If anything, CL makes it easier for a poorly focused AI megacorp to enter niche domains, which it doesn’t then need to know in subtle detail, letting its AIs get to know these domains on their own instead.