Because the marginal benefit from a smaller model learning your company’s tacit knowledge far outweights the marginal benefit from training a model with larger parameters. Like your marketing agent just doesn’t benefit from GPT6 solving yet another math conjecture. It does benefit from knowing your business and your customers and the subtleties of your particular industry.
Not without destroying the current prospects for AI company growth.
Yes, and I don’t think AI company growth prospects are a universal given. If we do end up with CL models, companies likely won’t want the learned knowledge to be distributed with other users of the model. So the incentive will be to own the weights, which directly threatens the business of the frontier model companies and their capacity to centrally train ever-larger models.
The crux is that bigger models are not very expensive. If bigger models are importantly more capable, if wouldn’t matter if they are moderately more expensive; and if they are only marginally more capable (even for the most difficult tasks), it doesn’t matter if they get trained (they might still in fact get trained). Prospects for AI company growth don’t obviously suffer even if bigger models aren’t important, because established AI companies can remain good at making smaller models as well.
And if AI companies do grow (for the biggest companies to individually command about 10-20% of global compute, out of hundreds of gigawatts), this then determines if the bigger models get trained. The fixed costs of training won’t be very high compared to the AI company scale, and they are still not very expensive so they’ll have their uses (even if they aren’t much more capable for many easier applications), and their dangers. Distillation motivates making very big models even when they are not directly useful, to make the smaller models marginally more capable, but the biggest models still not being very expensive makes it likely that they are available directly.
Because the marginal benefit from a smaller model learning your company’s tacit knowledge far outweights the marginal benefit from training a model with larger parameters. Like your marketing agent just doesn’t benefit from GPT6 solving yet another math conjecture. It does benefit from knowing your business and your customers and the subtleties of your particular industry.
Yes, and I don’t think AI company growth prospects are a universal given. If we do end up with CL models, companies likely won’t want the learned knowledge to be distributed with other users of the model. So the incentive will be to own the weights, which directly threatens the business of the frontier model companies and their capacity to centrally train ever-larger models.
The crux is that bigger models are not very expensive. If bigger models are importantly more capable, if wouldn’t matter if they are moderately more expensive; and if they are only marginally more capable (even for the most difficult tasks), it doesn’t matter if they get trained (they might still in fact get trained). Prospects for AI company growth don’t obviously suffer even if bigger models aren’t important, because established AI companies can remain good at making smaller models as well.
And if AI companies do grow (for the biggest companies to individually command about 10-20% of global compute, out of hundreds of gigawatts), this then determines if the bigger models get trained. The fixed costs of training won’t be very high compared to the AI company scale, and they are still not very expensive so they’ll have their uses (even if they aren’t much more capable for many easier applications), and their dangers. Distillation motivates making very big models even when they are not directly useful, to make the smaller models marginally more capable, but the biggest models still not being very expensive makes it likely that they are available directly.