Should we expect the PRC to support open weight models indefinitely? Or will incentives change such that Chinese frontier labs are forced to keep their weights secret?
It’s still too early to tell whether GLM 5.2 might have a “Deepseek” moment. Observers have become skeptical of benchmaxed open weight models that feel under performant in practice. However, the recent benchmarks are impressive, and we should have more “vibes” reports over the next couple days.
The releases comes on the back of Fable’s problematic safeguards and a subsequent executive order taking the model offline. Both events led to capabilities researchers emphasizing the importance of open weight models. So GLM 5.2 is well positioned to capture the attention of a wary research community.
Conversely, might this not also mean that the PRC is cued in to the risks of a model which claims to reach capabilities just behind Fable?
Open weight models seem to offer two advantages for the Chinese government. They allow Chinese labs to compete for consumers resistant to American proprietary models (e.g. anyone resistant to Anthropic API fees), and (more importantly?) they allow China to extend influence over countries which feel strategically disadvantaged by America proprietary models.
This is the problem of “mid-tier” powers, unable to compete with China and the US, uncertain of how to situate themselves in the AI race. The recent shutdown of Fable raises questions for how erstwhile American allies can ensure access to frontier intelligence. But if France can just use an open weight model from China, maybe they don’t need a fictional “le gros chaton” (assuming they are comfortable asking no questions about Tiananmen Square)?
However, if Chinese labs are entering into the same supposedly dangerous territory as Mythos, will China remain comfortable leaving these capabilities freely available to potential adversaries? They will draw their own conclusions from the US government’s erratic but wary approach to Mythos. And open weight models seem to have an inherently higher risk of jailbreaking and subterfuge. Is it consistent with the history of the PRC to assume they will want to give their populace greater access toward unregulated intelligence?
Allowing frontier labs to pursue open weight models has been advantageous to China until now, but I would anticipate incentives will change in the coming year, such that China implements industrial policy forbidding open weights for at least some subset of frontier intelligence.
I’m not sure if this would be good or bad for AI risk. Increased secrecy seems dangerous, but highly capable open weight models are perhaps more so?
Regards safe outcomes for superintelligence, your parenthetical remark is the one I believe most important. Far above any prosaic or theoretical safety work, our priority should be regulation preventing the development and release of superintelligence, at least until we have strong guarantees on its safety.
I don’t really disagree with any of the other points in your comment. Without a regulatory framework, it seems very likely that prosaic safety techniques will only contribute to bad outcomes. So it makes sense to me if one wants to focus on agent foundations and similar theoretic work. My post is not intended as a critique of agent foundations per se!
However, I do believe that one must be clearsighted on the risks of theoretic work, particularly when built upon abstractions. My critique is that agent foundations sometimes fails to make its assumptions explicit and works backwards from abstractions, effectively building a castle in the sky. A more robust approach would be to make these assumptions very explicit, ideally linking a theory to a set of axioms, so that we can better assess the defensibility of a theory. Some branches of continental philosophy are very bad at this (e.g. Lacan), starting from “metaphor” rather than an axiom, which is why I draw the parallel.
I will note that prosaic safety work could be relevant under a strong regulatory framework. For example, suppose we established an international treaty to freeze AI development at ChatGPT 5.5 Pro / Mythos. The treaty states that we can only advance to higher capabilities/intelligence when we are “sure” that the next model is aligned. With huge amounts of resources dedicated to verifying the next model if safe, it seems feasible to me that prosaic approaches could play a large role in building safe AI under such a regime.
Now, setting up sufficiently strong regulation is of course very hard, and one might critique that “proving” that the next generation of a model is aligned is akin to solving alignment itself! But I suspect that guaranteeing a single model is aligned is much easier than solving alignment for all possible models.
I would still guess it is better not to do prosaic safety work until a global regulatory framework exists, since it accelerates AI progress and thus reduces opportunities to implement said regulation. But there are enough counterarguments that I would be careful moralizing over it (not suggesting anyone in the comments is doing so!).