In addition to what others have said, I’m not sure it would be good from a safety perspective if smaller, personalized models became more viable. A big part of what makes current governance proposals anywhere close to workable is the fact that big models need big datacenters to train, so by restricting the datacenters you can restrict the big training runs. If it were possible to train AGI on consumer-grade hardware, it would be much harder to coordinate to slow down or pause development if that becomes necessary (I’d argue it’s already necessary but that’s beside the point).
Without CL, the main axis the model companies compete on is larger, more broadly capable models. The point I made in the other threads is that if you’re a SMB and need a slack AI to manage customer service, you don’t need that AI to also know organic chemistry. That it does today is an economic abberation simmilar to the adoption of mainframes before the PC came around. The PC decimated mainframes not because it was more powerful, but because it was cheaper, right-sized, and personalized.
If it’s true that CL is the missing capability for businesses to get ROI on AI, and not marginal improvement on group theory, or nuclear engineering, or biology, then its introduction will dampend incentives to ship ever larger models with “dangerous” knowledge very few businesses actually need, even if those larger models would similarly benefit from CL.
if you’re a SMB and need a slack AI to manage customer service, you don’t need that AI to also know organic chemistry. That it does today is an economic abberation simmilar to the adoption of mainframes before the PC came around.
Technical details of token costs say it’s active params that determine costs, not total params (when running within sufficiently big scale-up systems). Knowing more things (as opposed to being smarter about them) is more a total params thing (and currently there is no way to avoid knowing specific things; you can only avoid putting the effort into getting even better at them than you would by default). If there is some niche application that benefits from knowing organic chemistry, and the AI company puts in the work to get some model good at it, then there is no reason not to make all the models it makes good at it, to the extent they are able to stay coherent about sufficiently complicated questions like that (given the number of active params their cost permits them to have).
So there is no technical reason for there being an economic benefit from knowing less, if an AI company already knows how to teach AIs those things. Even cheap models with few active params can receive all the training in organic chemistry, they might just remain unable to get very good at it, but judging by the current capabilities even small cheap models will still be better than non-specialist humans. Future bigger scale-up systems/pods will enable more total params, to the point that in mid-2030s there might be “small” quadrillion param models that are very cheap (if very high levels of sparsity are technically feasible, which they might be with an absurd total number of reasonably sized routed experts).
Currently, models with fewer total params are useful because there are many legacy servers that can’t run the big models at all, not because they are inherently cheaper to run on the new bigger servers (which in 2026 are mostly rack-level scale-up systems like GB300 NVL72, but the future is multi-rack). This consideration will eventually go away, because bigger servers/pods are made out of the same stuff, the technology for connecting them into a single whole (scale-up system) just didn’t catch up to the demands of frontier LLMs yet, and therefore it’s not available for cheap models with few active params either, which is why cheap models with few active params also tend to have fewer total params.
So there is no technical reason for there being an economic benefit from knowing less
I think you’re still assuming there’s this one megacompany that’s training a single model distributed to the whole world. A small model can be trained by a specialized company that’s a good enough bootstrap based for a basket of skills such that CL can then specialize it for a user environment. If that small model can take market share for applications that require that basket of skills away from the mega AI corp, then you have the PC/mainframe dynamic again where the mainframe play gets completely disrupted.
My technical point is that “a small model” equivocates between active and total params, and a small-meaning-cheap model might have a quadrillion of total params in mid-2030s. Even with a few trillion total params, it can know all the things. Small-meaning-cheap 50B-100B active param models with trillions of total params are already happening; though such models are not yet the cheapest models, because GB300 NVL72 is not yet the low-end kind of hardware. So saying that a customer service model knowing organic chemistry is an abberation that will go away needs the premise that the company training the customer service model has difficulty obtaining the training data/process for organic chemistry, or vice versa, because there is no downside at the unit economics stage of the model’s lifecycle.
Because of the non-negative (and probably positive) transfer, meaning knowing more in one domain doesn’t hurt model performance in the other domains, and no effect on unit costs because of the active/total params distinction, it’s useful to make models that know all of the things simultaneously, for any target price tier. So there is a natural pressure for AI companies that specialize in different domains (to the extent such a thing happens at all) to merge, or to never seriously develop separately past a startup-acquisition phase (or beyond data-sourcing activities, selling the data to the companies actually training the models). This is a consequence of the unusual technological properties of MoE models, which don’t hold for many other kinds of products. If anything, CL makes it easier for a poorly focused AI megacorp to enter niche domains, which it doesn’t then need to know in subtle detail, letting its AIs get to know these domains on their own instead.
Hmm, I kind of suspect that there’s something deep about the fact that making models larger makes them smarter. Something like, I have a prior that all knowledge is connected, such that in principle getting better at nuclear engineering does make you better at, like, biology, even if only a little. Maybe you’ll say “okay, but with CL companies could make models that are good enough at what they need while still being non-general”. Sure, maybe, but the worry is that if you can make a model like that on consumer grade hardware, then eventually you could also make a general “AGI” on consumer-level hardware if you wanted to, and that would be dangerous for all the usual reasons. And if that’s possible, it’s much harder to stop than the current large training runs for the reasons I mentioned in my previous comment.
In addition to what others have said, I’m not sure it would be good from a safety perspective if smaller, personalized models became more viable. A big part of what makes current governance proposals anywhere close to workable is the fact that big models need big datacenters to train, so by restricting the datacenters you can restrict the big training runs. If it were possible to train AGI on consumer-grade hardware, it would be much harder to coordinate to slow down or pause development if that becomes necessary (I’d argue it’s already necessary but that’s beside the point).
Without CL, the main axis the model companies compete on is larger, more broadly capable models. The point I made in the other threads is that if you’re a SMB and need a slack AI to manage customer service, you don’t need that AI to also know organic chemistry. That it does today is an economic abberation simmilar to the adoption of mainframes before the PC came around. The PC decimated mainframes not because it was more powerful, but because it was cheaper, right-sized, and personalized.
If it’s true that CL is the missing capability for businesses to get ROI on AI, and not marginal improvement on group theory, or nuclear engineering, or biology, then its introduction will dampend incentives to ship ever larger models with “dangerous” knowledge very few businesses actually need, even if those larger models would similarly benefit from CL.
Technical details of token costs say it’s active params that determine costs, not total params (when running within sufficiently big scale-up systems). Knowing more things (as opposed to being smarter about them) is more a total params thing (and currently there is no way to avoid knowing specific things; you can only avoid putting the effort into getting even better at them than you would by default). If there is some niche application that benefits from knowing organic chemistry, and the AI company puts in the work to get some model good at it, then there is no reason not to make all the models it makes good at it, to the extent they are able to stay coherent about sufficiently complicated questions like that (given the number of active params their cost permits them to have).
So there is no technical reason for there being an economic benefit from knowing less, if an AI company already knows how to teach AIs those things. Even cheap models with few active params can receive all the training in organic chemistry, they might just remain unable to get very good at it, but judging by the current capabilities even small cheap models will still be better than non-specialist humans. Future bigger scale-up systems/pods will enable more total params, to the point that in mid-2030s there might be “small” quadrillion param models that are very cheap (if very high levels of sparsity are technically feasible, which they might be with an absurd total number of reasonably sized routed experts).
Currently, models with fewer total params are useful because there are many legacy servers that can’t run the big models at all, not because they are inherently cheaper to run on the new bigger servers (which in 2026 are mostly rack-level scale-up systems like GB300 NVL72, but the future is multi-rack). This consideration will eventually go away, because bigger servers/pods are made out of the same stuff, the technology for connecting them into a single whole (scale-up system) just didn’t catch up to the demands of frontier LLMs yet, and therefore it’s not available for cheap models with few active params either, which is why cheap models with few active params also tend to have fewer total params.
I think you’re still assuming there’s this one megacompany that’s training a single model distributed to the whole world. A small model can be trained by a specialized company that’s a good enough bootstrap based for a basket of skills such that CL can then specialize it for a user environment. If that small model can take market share for applications that require that basket of skills away from the mega AI corp, then you have the PC/mainframe dynamic again where the mainframe play gets completely disrupted.
My technical point is that “a small model” equivocates between active and total params, and a small-meaning-cheap model might have a quadrillion of total params in mid-2030s. Even with a few trillion total params, it can know all the things. Small-meaning-cheap 50B-100B active param models with trillions of total params are already happening; though such models are not yet the cheapest models, because GB300 NVL72 is not yet the low-end kind of hardware. So saying that a customer service model knowing organic chemistry is an abberation that will go away needs the premise that the company training the customer service model has difficulty obtaining the training data/process for organic chemistry, or vice versa, because there is no downside at the unit economics stage of the model’s lifecycle.
Because of the non-negative (and probably positive) transfer, meaning knowing more in one domain doesn’t hurt model performance in the other domains, and no effect on unit costs because of the active/total params distinction, it’s useful to make models that know all of the things simultaneously, for any target price tier. So there is a natural pressure for AI companies that specialize in different domains (to the extent such a thing happens at all) to merge, or to never seriously develop separately past a startup-acquisition phase (or beyond data-sourcing activities, selling the data to the companies actually training the models). This is a consequence of the unusual technological properties of MoE models, which don’t hold for many other kinds of products. If anything, CL makes it easier for a poorly focused AI megacorp to enter niche domains, which it doesn’t then need to know in subtle detail, letting its AIs get to know these domains on their own instead.
Hmm, I kind of suspect that there’s something deep about the fact that making models larger makes them smarter. Something like, I have a prior that all knowledge is connected, such that in principle getting better at nuclear engineering does make you better at, like, biology, even if only a little. Maybe you’ll say “okay, but with CL companies could make models that are good enough at what they need while still being non-general”. Sure, maybe, but the worry is that if you can make a model like that on consumer grade hardware, then eventually you could also make a general “AGI” on consumer-level hardware if you wanted to, and that would be dangerous for all the usual reasons. And if that’s possible, it’s much harder to stop than the current large training runs for the reasons I mentioned in my previous comment.