The current writeup on Opus 4 and Mythos 5 (and their costs/prices) is in my last post (that’s more recent than the above comment).
The 104B figure was a bit surprising, the shared experts turn out to be very big. Maybe next year we’ll even see a model targeting the hardware available in China that’s genuinely Sonnet-class (in shape/size and thus capabilities that are downstream of pretraining and can’t be comprehensively RLVRed; something like 200B-300B active params). The way the current below-Sonnet class models are at the level of GPT-5.4 (probably itself Sonnet-class in the above sense), a Sonnet-class model made this well might be somewhere between Opus 4.5 and Opus 5 in agentic usage across the board (plus whatever more narrow things can be RLVRed at that time). But a Mythos/Astra-class model likely needs bigger scale-up systems or HBM4 (so that more pipelining becomes feasible), the current method of doing expert parallelism on the scale-out network (since the scale-up nodes are too small) doesn’t seem sustainable for that many active params.
The current writeup on Opus 4 and Mythos 5 (and their costs/prices) is in my last post (that’s more recent than the above comment).
The 104B figure was a bit surprising, the shared experts turn out to be very big. Maybe next year we’ll even see a model targeting the hardware available in China that’s genuinely Sonnet-class (in shape/size and thus capabilities that are downstream of pretraining and can’t be comprehensively RLVRed; something like 200B-300B active params). The way the current below-Sonnet class models are at the level of GPT-5.4 (probably itself Sonnet-class in the above sense), a Sonnet-class model made this well might be somewhere between Opus 4.5 and Opus 5 in agentic usage across the board (plus whatever more narrow things can be RLVRed at that time). But a Mythos/Astra-class model likely needs bigger scale-up systems or HBM4 (so that more pipelining becomes feasible), the current method of doing expert parallelism on the scale-out network (since the scale-up nodes are too small) doesn’t seem sustainable for that many active params.