It’s not clear that a FLOP is an appropriate functional entity to base a comparison on. In biological brains, cognitive tasks do not appear to be implemented from a composition of FLOPs, like they are in computers. Since we don’t know the granular units of computation in biological brains, or even if cognitive tasks neatly decompose into granular units in brains, we should instead use holistic cognitive tasks as a basis for apples-to-apples comparison. We could look to the AI benchmark space for such potential tasks, such as solving a math test, like the IMO. How much energy does a human brain need to complete an IMO (in competition humans are given 9 hours, so about 180 watt-hours based on the common 20-watt assumption). How about a flagship frontier LLM? How about a lower parameter fine-tuned-for-math distilled LLM? Harder to estimate given shifting sands and proprietary tech co’s, but it is not wildly orders of magnitude more than 180 watt-hours, supporting the sentiment of this article.
It’s not clear that a FLOP is an appropriate functional entity to base a comparison on. In biological brains, cognitive tasks do not appear to be implemented from a composition of FLOPs, like they are in computers. Since we don’t know the granular units of computation in biological brains, or even if cognitive tasks neatly decompose into granular units in brains, we should instead use holistic cognitive tasks as a basis for apples-to-apples comparison. We could look to the AI benchmark space for such potential tasks, such as solving a math test, like the IMO. How much energy does a human brain need to complete an IMO (in competition humans are given 9 hours, so about 180 watt-hours based on the common 20-watt assumption). How about a flagship frontier LLM? How about a lower parameter fine-tuned-for-math distilled LLM? Harder to estimate given shifting sands and proprietary tech co’s, but it is not wildly orders of magnitude more than 180 watt-hours, supporting the sentiment of this article.