The clearest source I know on compute multipliers of MoE is this Jan 2025 paper. The key data is in Appendix D.4, Figure 11 and Figure 12, left. My takeaway is that 8x sparsity gives a 3x compute multiplier, and 30x sparsity gives a 6x compute multiplier. Figure 11 suggests that the gains continue with more sparsity, and are linear in the logarithm of sparsity (but then we run out of data, since MoEs also increase the compute optimal D/N ratio, seemingly by exactly the same factor as the compute multiplier boon).
(I make a couple more points on why it’s hard to get a sense for these numbers from indirect estimates, experiments that are not specifically about this particular question, in a comment I left under the “Origin of Algorithmic Progress” linkpost 6 months ago.)
The clearest source I know on compute multipliers of MoE is this Jan 2025 paper. The key data is in Appendix D.4, Figure 11 and Figure 12, left. My takeaway is that 8x sparsity gives a 3x compute multiplier, and 30x sparsity gives a 6x compute multiplier. Figure 11 suggests that the gains continue with more sparsity, and are linear in the logarithm of sparsity (but then we run out of data, since MoEs also increase the compute optimal D/N ratio, seemingly by exactly the same factor as the compute multiplier boon).
(I make a couple more points on why it’s hard to get a sense for these numbers from indirect estimates, experiments that are not specifically about this particular question, in a comment I left under the “Origin of Algorithmic Progress” linkpost 6 months ago.)