Which country in particular? In civilized countries local authorities will likely shut such operations down as soon as there is real damage from cybercriminals using it, failed states generally lack infrastructure to sustain them, and pariah states will control such GPUs as a national asset
Petropolitan
Hosted where?
Such activity requires feeding lots of cyber-related prompts (including offensive) to powerful models (at least Kimi K3-level). Such a possibility would be very useful to human cybercriminals, why would API providers allow this without KYC?
The reason consumer GPUs and even a single A100 might be scattered around without use is that one can’t inference any useful coding agents on them, making the scenario you suggest impossible.
If they were to appear (very improbable by the end of this year and unlikely even next year), this hardware would become much more valuable both for its legitimate owners and for cybercriminals.
It’s quite obvious that professional (human) cybercriminals with advanced agents inferenced on large clusters (say, 8xH200s) will exploit such hardware much earlier and more effectively than AIs, making such compute basically unavailable to the latter
Mind you, the AI capabilities are very jagged and will remain so. This is BTW entirely missing from the Zvi post, but I don’t have time to argue that, especially since the author doesn’t respond in comments anyway.
For me it looks more like LLMs around 4o cracked the theory of mind and figured out how to persuade people to some extent, likely better than humans writing text—even though humans actually achieve their (our) best level of persuasion in person. But the progress basically stagnated on that, and from the apparent lack of impact of the current AI persuasion we can infer that the current level of capabilities is not world-changing
Which kinds of misalignment might one get from the on-policy distillation with no direct RL on release candidates as practiced by DeepSeek on v4 (e. g., see https://youtu.be/AIRfT41A89s?t=1213 )? How likely would undesirable characteristics of the third and fourth kind be “smuggled” from the RL’d checkpoints via mechanisms similar to subliminal learning? Could the filtering mechanisms prevent that?
Looks like a rich and interesting empirical research direction
Let me clarify what I had in mind by “smart engineering”: sure, 8-high HBF is slower than 8-high HBM but you can supplement part of 12-high HBM with, say, 6-high HBF and get same or better bandwidth for a lower chip price but higher electricity consumption.[1]
As for the prefill, if you can allocate a certain share of hardware to prefill-only, you can minimize HBM on that hardware. However, there are usually ~3x less prefill than decode nodes and in practice even less than a quarter of hardware is “locked” into fixed pools for flexibility, so the economic effect will be limited (maybe that’s why research into heterogeneous hardware generally is so slow).
- ^
There’s also a latency aspect, but experts in the mid-to-late layers can be prefetched early in the forward pass, while those in the early layers can be predicted by the MTP heads on the previous token. Although both approaches come with some bandwidth penalty as the predictions will never be perfect
- ^
The study the authors link indicates that Opus 4.1 from a year ago is marginally better than 4o and those both are way better than Opus 4.6 and GPT-5.4 from this year. If we believe these results, this doesn’t demonstrate any progress towards the “superpersuasion” but quite to the contrary
Except there has been no progress towards this kind of “superpersuasion” since about late 2024 despite impressive progress in benchmarks otherwise. If anything, there has been a regress because many humans on the Internet became more attentive to the signs of AI text (and thus more inclined to ignore unsolicited AI attempts to persuade).
The reason is quite obvious to me: there’s no scalable way to measure how persuasive was a certain LLM response, and thus it’s impossible to hill-climb this skill with post-training (and it doesn’t come for free with pre-training either). Note that social media reach and similar metrics don’t substitute for that
Thank you for quite an interesting reply (sorry, I missed that part of Footnote 6 since it was partially obscured by 7).
BTW, have you seen the recent news about High Bandwidth Flash? Seems very beneficial for prefill, and perhaps could find some use for decode with smart engineering, especially with some latency compromises. The disadvantages might include lower energy efficiency though. Not sure how easy this would be to incorporate into your estimates
Something like this has been my expectation since approximately announcement of Devin in spring 2024, with a major caveat: policymakers won’t push for economically costly measures until some people die from misaligned AIs (and I don’t mean suicides), but the issue is certainly unsolvable with “cheap” measures, meaning people will have to die, and that’s still not a guarantee =(
The statement seems to be false though, as math is not actually a key ingredient of AI research.
LLM architecture is empirically-driven engineering not mathematics, it’s best characterized by a popular quote from Noam Shazeer’s 2020 SwiGLU paper:
We offer no explanation as to why these architectures seem to work; we attribute their success, as all else, to divine benevolence.
Only two optimizers have been adopted in 14 years since AlexNet, Adam in 2014 and Muon in 2024. Even putting aside the discussion of the extent to which their development was actually math-driven, the next good optimizer is not to be expected in years, this is a very minor part of AI research.
Reading the loss curves is not anymore mathematics than reading the performance telemetry of a, say, experimental gas turbine or rocket engine, which is obviously pure engineering.
Sure, mathematicians could be retrained into excellent AI researchers but one shouldn’t create an illusion that they are going to continue doing math for a living if they move into AI
I’m guessing Opus 4.8 to be based on the same pretrain as Opus 4.0
I think that’s actually impossible, as the tokenizer got changed between Opus 4.6 and 4.7
Sonnet CoT summaries stopped appearing for me in web today, could you please check as well?
I think this is mainly a crack down on distillation. OpenAI has shown no reasoning summaries in ChatGPT for about a month, and they are distilled much less if at all
Very interesting and kinda unsurprising, I guess! I expected that because these skills won’t develop without training data which is hard to come by.
Hope people who still expect jagged capabilities to somehow go away and models to magically generalize from math and software engineering to IRL tasks see this and update
expensive autocannon in the path of a flying explosive
Autocannons are not particularly expensive, when compared with their own targeting systems, and especially when compared with direct energy weapons of similar range and effectiveness. Heavy machine guns are even cheaper (albeit with diminished range), and medium (7.62-mm) MGs are cheaper still.
The link refers to heavy field howitzers which are actually quite costly, although their barrels are not.
This is essentially the same problem as distributing bioweapons from theater-range ballistic missiles: solvable for large industrial states but hard, expensive and requires a lot of testing
Most natural gas distribution networks are vulnerable to overpressurization which might cause city-wide fires and hundreds of casualties, a few terrorists can organize that by physically capturing relevant infrastructure. But for that they need technical competences which they lack, as no contemporary violent non-state actors really reach the sophistication of Aum Shinrikyo.
Your plan, however, requires not only engineering skills but time and money and production facilities. It’s not impossible that another AS-like apocalyptic sect will emerge in the future but it seems unlikely. Hezbollah might have the means to develop capability in the near future (as Yair Halberstadt has rightfully noted below) and it may be a natural outgrowth of their current drone program but they won’t be able to keep the program secret from Israeli intelligence
If you put a dangerous wild animal in a flimsy cubicle, it escapes and causes some damage in the real world, it can’t be held accountable but you will. Same for an AI agent
What happens if instead of using any generations you initialize “mechanically” with just a different number of ending sentences, maybe from 1 to 5? Or even just several last tokens?
Russia, Belarus, Iran, Venezuela and Myanmar are pariah countries under sanctions, their regimes control very limited amount of compute which they have to prioritize and manage strategically. Romania and Bulgaria are literal EU countries, EU will manage that. The rest are integrated into global economy and will be pressured by the rich countries which are priority targets for cybercrime to shut down illicit inference.
The current list is not informative for the near-future where cybercriminals use GLM-5.3 (possibly slightly finetuned, or perhaps another model with better performance) with a specialized harness at scale as Vast.ai or any other large service will have to crack down on cybercrime as soon as the latter scales up