The crux is that bigger models are not very expensive. If bigger models are importantly more capable, if wouldn’t matter if they are moderately more expensive; and if they are only marginally more capable (even for the most difficult tasks), it doesn’t matter if they get trained (they might still in fact get trained). Prospects for AI company growth don’t obviously suffer even if bigger models aren’t important, because established AI companies can remain good at making smaller models as well.
And if AI companies do grow (for the biggest companies to individually command about 10-20% of global compute, out of hundreds of gigawatts), this then determines if the bigger models get trained. The fixed costs of training won’t be very high compared to the AI company scale, and they are still not very expensive so they’ll have their uses (even if they aren’t much more capable for many easier applications), and their dangers. Distillation motivates making very big models even when they are not directly useful, to make the smaller models marginally more capable, but the biggest models still not being very expensive makes it likely that they are available directly.
The crux is that bigger models are not very expensive. If bigger models are importantly more capable, if wouldn’t matter if they are moderately more expensive; and if they are only marginally more capable (even for the most difficult tasks), it doesn’t matter if they get trained (they might still in fact get trained). Prospects for AI company growth don’t obviously suffer even if bigger models aren’t important, because established AI companies can remain good at making smaller models as well.
And if AI companies do grow (for the biggest companies to individually command about 10-20% of global compute, out of hundreds of gigawatts), this then determines if the bigger models get trained. The fixed costs of training won’t be very high compared to the AI company scale, and they are still not very expensive so they’ll have their uses (even if they aren’t much more capable for many easier applications), and their dangers. Distillation motivates making very big models even when they are not directly useful, to make the smaller models marginally more capable, but the biggest models still not being very expensive makes it likely that they are available directly.