The two-layer design is right, and there is a physical reason. Ideas are information, and information leaks: the US tried to export-control cryptography in the 1990s and failed completely. Compute is physical, and physical things can be verified, the way fissile material was under the NPT. So open research plus restricted hardware is the right split. But it means the whole plan depends on one variable: does frontier AI stay compute-hungry? If algorithmic efficiency keeps letting last year’s frontier run on ordinary hardware, the world is licensed on paper but leaky in practice. That variable should be the plan’s primary tripwire, not a background assumption.
My main worry is that the plan defends against the wrong failure mode. MACD deters overt defection: leaving the treaty, building a secret datacenter. But history says the overt move almost never comes. Preventive war against the Soviet nuclear program was seriously advocated and never executed. What actually happens is covert friction: quiet, deniable interference kept below the retaliation threshold. So the real question is whether the verification layer can detect steady covert cheating, and the plan itself admits the gap. Footnote 11 of the transparency plan concedes a defector could do research steganographically through the monitored channels, and the misuse defense depends on refusal training, which jailbreak research shows is reliably bypassable.
The pace mismatch makes it worse. Nuclear arms control took about twenty years to reach real verification, for a technology that barely changed while being negotiated. AI capabilities turn over in months, so a treaty will always describe last year’s AI. That’s why AI-assisted verification, their own “AI Silver Bullet”, is not a side project: it is the only mechanism that can re-match the two speeds.
One concrete ask: make the security assumptions as measurable as the compute assumptions. The plan quantifies its compute verification (“99.99% confidence that 99.9% of the compute was verified”), but there are no comparable numbers for refusal robustness or for the steganographic capacity of the monitored channels. Quantified bypass rates and capacity bounds would make the security claims treaty-grade, the way the compute accounting already is. Bounding steganographic capacity in particular looks tractable, and nobody is doing it.
The two-layer design is right, and there is a physical reason. Ideas are information, and information leaks: the US tried to export-control cryptography in the 1990s and failed completely. Compute is physical, and physical things can be verified, the way fissile material was under the NPT. So open research plus restricted hardware is the right split. But it means the whole plan depends on one variable: does frontier AI stay compute-hungry? If algorithmic efficiency keeps letting last year’s frontier run on ordinary hardware, the world is licensed on paper but leaky in practice. That variable should be the plan’s primary tripwire, not a background assumption.
My main worry is that the plan defends against the wrong failure mode. MACD deters overt defection: leaving the treaty, building a secret datacenter. But history says the overt move almost never comes. Preventive war against the Soviet nuclear program was seriously advocated and never executed. What actually happens is covert friction: quiet, deniable interference kept below the retaliation threshold. So the real question is whether the verification layer can detect steady covert cheating, and the plan itself admits the gap. Footnote 11 of the transparency plan concedes a defector could do research steganographically through the monitored channels, and the misuse defense depends on refusal training, which jailbreak research shows is reliably bypassable.
The pace mismatch makes it worse. Nuclear arms control took about twenty years to reach real verification, for a technology that barely changed while being negotiated. AI capabilities turn over in months, so a treaty will always describe last year’s AI. That’s why AI-assisted verification, their own “AI Silver Bullet”, is not a side project: it is the only mechanism that can re-match the two speeds.
One concrete ask: make the security assumptions as measurable as the compute assumptions. The plan quantifies its compute verification (“99.99% confidence that 99.9% of the compute was verified”), but there are no comparable numbers for refusal robustness or for the steganographic capacity of the monitored channels. Quantified bypass rates and capacity bounds would make the security claims treaty-grade, the way the compute accounting already is. Bounding steganographic capacity in particular looks tractable, and nobody is doing it.