I looked up EpochAI’s list of benchmarks. The very poor ECI of Grok 4.5 as opposed to GPT-5.5 seems to be a result of xAI not caring about math. Additionally, Grok 4.6 seems to be closer to Opus 4.8 with respect to ARC-AGI-3, if not outright outperforming Opus, as the actual ARC-AGI-3 score suggests. Does it mean that xAI is less behind than we think? I guess that Zvi will have to write something like “Grok 4.6 is three, not six, mounths behind. Stop xAI to hell!”
Also, SpaceX has caught up to Meta in credibly having enough compute in 2027-2028[1] to stay in the game, if either of them can assemble a functional model development team. As Google illustrates, it’s not easy to do that (even when you have some of the best people), but as OpenAI and Anthropic illustrate, it’s not so difficult that it can’t be replicated. Muse Sparks are probably small enough models that their non-frontier performance doesn’t count as evidence that they’re not well-made (and that a Mythos-sized Muse model won’t have Mythos-level capabilities). SpaceX is in a worse position in 2026 on priors, because the xAI team wasn’t doing that well in 2025 and then got disrupted in early 2026, but they aren’t in a worse situation than Meta was in 2025, so it remains plausible that in 2027 they catch up.
The public part of the text article leaves many gaps in the arguments (wildly gesturing to a somewhat unusual extent at the presumed arguments hidden behind the various paywalls), but some relevant things were discussed in the video version. Basically, it’s a feasibility and track record argument. I’m guessing there’s an assumption that others aren’t competing too strongly for the sites that SpaceX could in principle use for the 2027 buildout (since others may prefer greenfield developments, and SpaceX might be willing to outbid people who are planning for the more distant future), or that you can find enough additional sites when you have an unlimited budget for searching.
I looked up EpochAI’s list of benchmarks. The very poor ECI of Grok 4.5 as opposed to GPT-5.5 seems to be a result of xAI not caring about math. Additionally, Grok 4.6 seems to be closer to Opus 4.8 with respect to ARC-AGI-3, if not outright outperforming Opus, as the actual ARC-AGI-3 score suggests. Does it mean that xAI is less behind than we think? I guess that Zvi will have to write something like “Grok 4.6 is three, not six, mounths behind. Stop xAI to hell!”
P.S. The same issue seems to apply to Meta’s Muse Spark, except that it has even less evaluated benchmarks.
Also, SpaceX has caught up to Meta in credibly having enough compute in 2027-2028 [1] to stay in the game, if either of them can assemble a functional model development team. As Google illustrates, it’s not easy to do that (even when you have some of the best people), but as OpenAI and Anthropic illustrate, it’s not so difficult that it can’t be replicated. Muse Sparks are probably small enough models that their non-frontier performance doesn’t count as evidence that they’re not well-made (and that a Mythos-sized Muse model won’t have Mythos-level capabilities). SpaceX is in a worse position in 2026 on priors, because the xAI team wasn’t doing that well in 2025 and then got disrupted in early 2026, but they aren’t in a worse situation than Meta was in 2025, so it remains plausible that in 2027 they catch up.
The public part of the text article leaves many gaps in the arguments (wildly gesturing to a somewhat unusual extent at the presumed arguments hidden behind the various paywalls), but some relevant things were discussed in the video version. Basically, it’s a feasibility and track record argument. I’m guessing there’s an assumption that others aren’t competing too strongly for the sites that SpaceX could in principle use for the 2027 buildout (since others may prefer greenfield developments, and SpaceX might be willing to outbid people who are planning for the more distant future), or that you can find enough additional sites when you have an unlimited budget for searching.