Astra is Latin for “stars”, so it’s a plausible next step in the Luna-Terra-Sol sequence. As I argued in the second paragraph, the hypothesis of Sol being Opus-class in shape/size has merit (from earlier clues, before taking the mention of Astra into account). In which case there is room in 2026 for a bigger OpenAI model, and it’s already August.
yes, but that doesn’t mean the sequence to more power is based on”bigger models”, and in fact the metaphor from sol (single star) to astra (many stars), suggests some sort of multi-agent scaffolding mode
“Astra” being plural does suggest multi-agent mode, but they already have the “Pro” qualifier, and “Codex” for the harness. Also, Luna-Terra-Sol doesn’t naturally suggest good choices for the next step. Zvi chose “Galaxy”; I was thinking “Antares” because there are going to be at least 2 more models beyond Mythos-class deserving of their own weight class names (possibly 4+ if half-steps like Opus 4 to Mythos 5 should count), and “Galaxy” takes too large a step. So the suboptimal “Astra” makes sense in this framing, an uneasy fit that’s not much worse than the alternatives. “Astra” has the same issue as “Galaxy”, it’s not very future-proof, but then neither is “Mythos”.
The sequence of more capability is inevitably based on bigger models (at least in effective active params; but because of legacy hardware from multi-year rental contracts, using fewer total params for weaker models is also important). There is a training process that’s essentially the same for all the models, the main difference between the models is their size. If a smaller model is as capable as a bigger model, there is no use at all for the bigger model, and something still needs to distinguish the differently-capable and differently-priced models.
Smaller models (in active params and KV cache per token) naturally cost less, while being less capable when trained with as much compute (which gets worse when they’re trained with only as much data and thus less compute). A model smaller than the compute optimal frontier model essentially can’t become as capable through overtraining (while using the same training process and data), because it would take more compute than the frontier model did, and that greater amount of compute is either unavailable, or it could be used for an even more capable bigger frontier model instead. This only happens when the smaller model is trained at a different AI company with different methods and data, but not when it’s another model in a series from a single AI company.
Astra is Latin for “stars”, so it’s a plausible next step in the Luna-Terra-Sol sequence. As I argued in the second paragraph, the hypothesis of Sol being Opus-class in shape/size has merit (from earlier clues, before taking the mention of Astra into account). In which case there is room in 2026 for a bigger OpenAI model, and it’s already August.
yes, but that doesn’t mean the sequence to more power is based on”bigger models”, and in fact the metaphor from sol (single star) to astra (many stars), suggests some sort of multi-agent scaffolding mode
“Astra” being plural does suggest multi-agent mode, but they already have the “Pro” qualifier, and “Codex” for the harness. Also, Luna-Terra-Sol doesn’t naturally suggest good choices for the next step. Zvi chose “Galaxy”; I was thinking “Antares” because there are going to be at least 2 more models beyond Mythos-class deserving of their own weight class names (possibly 4+ if half-steps like Opus 4 to Mythos 5 should count), and “Galaxy” takes too large a step. So the suboptimal “Astra” makes sense in this framing, an uneasy fit that’s not much worse than the alternatives. “Astra” has the same issue as “Galaxy”, it’s not very future-proof, but then neither is “Mythos”.
The sequence of more capability is inevitably based on bigger models (at least in effective active params; but because of legacy hardware from multi-year rental contracts, using fewer total params for weaker models is also important). There is a training process that’s essentially the same for all the models, the main difference between the models is their size. If a smaller model is as capable as a bigger model, there is no use at all for the bigger model, and something still needs to distinguish the differently-capable and differently-priced models.
Smaller models (in active params and KV cache per token) naturally cost less, while being less capable when trained with as much compute (which gets worse when they’re trained with only as much data and thus less compute). A model smaller than the compute optimal frontier model essentially can’t become as capable through overtraining (while using the same training process and data), because it would take more compute than the frontier model did, and that greater amount of compute is either unavailable, or it could be used for an even more capable bigger frontier model instead. This only happens when the smaller model is trained at a different AI company with different methods and data, but not when it’s another model in a series from a single AI company.