The study the authors link indicates that Opus 4.1 from a year ago is marginally better than 4o and those both are way better than Opus 4.6 and GPT-5.4 from this year. If we believe these results, this doesn’t demonstrate any progress towards the “superpersuasion” but quite to the contrary
Petropolitan
Except there has been no progress towards this kind of “superpersuasion” since about late 2024 despite impressive progress in benchmarks otherwise. If anything, there has been a regress because many humans on the Internet became more attentive to the signs of AI text (and thus more inclined to ignore unsolicited AI attempts to persuade).
The reason is quite obvious to me: there’s no scalable way to measure how persuasive was a certain LLM response, and thus it’s impossible to hill-climb this skill with post-training (and it doesn’t come for free with pre-training either). Note that social media reach and similar metrics don’t substitute for that
Thank you for quite an interesting reply (sorry, I missed that part of Footnote 6 since it was partially obscured by 7).
BTW, have you seen the recent news about High Bandwidth Flash? Seems very beneficial for prefill, and perhaps could find some use for decode with smart engineering, especially with some latency compromises. The disadvantages might include lower energy efficiency though. Not sure how easy this would be to incorporate into your estimates
Something like this has been my expectation since approximately announcement of Devin in spring 2024, with a major caveat: policymakers won’t push for economically costly measures until some people die from misaligned AIs (and I don’t mean suicides), but the issue is certainly unsolvable with “cheap” measures, meaning people will have to die, and that’s still not a guarantee =(
The statement seems to be false though, as math is not actually a key ingredient of AI research.
LLM architecture is empirically-driven engineering not mathematics, it’s best characterized by a popular quote from Noam Shazeer’s 2020 SwiGLU paper:
We offer no explanation as to why these architectures seem to work; we attribute their success, as all else, to divine benevolence.
Only two optimizers have been adopted in 14 years since AlexNet, Adam in 2014 and Muon in 2024. Even putting aside the discussion of the extent to which their development was actually math-driven, the next good optimizer is not to be expected in years, this is a very minor part of AI research.
Reading the loss curves is not anymore mathematics than reading the performance telemetry of a, say, experimental gas turbine or rocket engine, which is obviously pure engineering.
Sure, mathematicians could be retrained into excellent AI researchers but one shouldn’t create an illusion that they are going to continue doing math for a living if they move into AI
I’m guessing Opus 4.8 to be based on the same pretrain as Opus 4.0
I think that’s actually impossible, as the tokenizer got changed between Opus 4.6 and 4.7
Sonnet CoT summaries stopped appearing for me in web today, could you please check as well?
I think this is mainly a crack down on distillation. OpenAI has shown no reasoning summaries in ChatGPT for about a month, and they are distilled much less if at all
Very interesting and kinda unsurprising, I guess! I expected that because these skills won’t develop without training data which is hard to come by.
Hope people who still expect jagged capabilities to somehow go away and models to magically generalize from math and software engineering to IRL tasks see this and update
expensive autocannon in the path of a flying explosive
Autocannons are not particularly expensive, when compared with their own targeting systems, and especially when compared with direct energy weapons of similar range and effectiveness. Heavy machine guns are even cheaper (albeit with diminished range), and medium (7.62-mm) MGs are cheaper still.
The link refers to heavy field howitzers which are actually quite costly, although their barrels are not.
This is essentially the same problem as distributing bioweapons from theater-range ballistic missiles: solvable for large industrial states but hard, expensive and requires a lot of testing
Most natural gas distribution networks are vulnerable to overpressurization which might cause city-wide fires and hundreds of casualties, a few terrorists can organize that by physically capturing relevant infrastructure. But for that they need technical competences which they lack, as no contemporary violent non-state actors really reach the sophistication of Aum Shinrikyo.
Your plan, however, requires not only engineering skills but time and money and production facilities. It’s not impossible that another AS-like apocalyptic sect will emerge in the future but it seems unlikely. Hezbollah might have the means to develop capability in the near future (as Yair Halberstadt has rightfully noted below) and it may be a natural outgrowth of their current drone program but they won’t be able to keep the program secret from Israeli intelligence
If you put a dangerous wild animal in a flimsy cubicle, it escapes and causes some damage in the real world, it can’t be held accountable but you will. Same for an AI agent
What happens if instead of using any generations you initialize “mechanically” with just a different number of ending sentences, maybe from 1 to 5? Or even just several last tokens?
Do you think you could try to retrieve yesterday’s CoT today?
I have stumbled upon the problem that an important file Sonnet 5 fetched at my request as a part of the reasoning process got discarded in the web interface after a dialog turn. I presume these files are likely managed as a part of CoT, and reasoning from previous turns gets purged for reasons of economy as soon as KV cache is deleted
FWIW, Gemma 3 12B doesn’t appear to express similar “thoughts” in a comparable way, although it has ” pretending” and some obscenities on some of the tokens of its response https://www.neuronpedia.org/jlens/cmrcms0kk000004l7cyqf1uu4
From my limited experience on Neuronpedia, “sort of buffer for tokens that might need to be emitted soon” and “mix of potential-next-token buffering as well as something higher-order” are actually quite good descriptions at least for two models presented! I would recommend trying the “free chat” option yourself.
If we are to take the aforementioned “concept vector variance” figures at face value, signal-to-noise ratio of the J-lens is ~1:15 even when the layers are carefully selected by experienced researchers. The ratio for the logit lens has to be worse yet, that’s probably why no one discovered this phenomenon with that instrument earlier.
OTOH, if the labs starts hill-climbing on this in some way, the signal-to-noise ratio might quite likely go down, but on the other, people could come up with more reliable methods of detecting the “workspace” based on this work (e. g., by down-weighting the tokens which actually appear in the later generation or semantically similar to them, or ones detected by something like the Future Lens https://ar5iv.labs.arxiv.org/html/2311.04897 etc.)
You spent practically all of your reply critiquing sections 3 and 4 of the essay which I don’t actually endorse[1], entirely ignoring the question of how a model trained in a containerized RLVR environment will magically generalize to unstructured real-world tasks which dominate the world economy (I looked it up: out of gross world product of ~$126T according to IMF, IT services and software contribute to less than $4T according to Gartner) and all of Patel’s arguments why it won’t
- ^
I agree that RLVR does teach deep skills much better than alternatives tried so far, but that fact only leads to increased jaggedness not generalization!
- ^
Mythos already got cleared for ~100 American companies and their foreign employees https://www.semafor.com/article/06/27/2026/us-releases-powerful-anthropic-model-mythos-to-some-us-companies
Let me clarify what I had in mind by “smart engineering”: sure, 8-high HBF is slower than 8-high HBM but you can supplement part of 12-high HBM with, say, 6-high HBF and get same or better bandwidth for a lower chip price but higher electricity consumption.[1]
As for the prefill, if you can allocate a certain share of hardware to prefill-only, you can minimize HBM on that hardware. However, there are usually ~3x less prefill than decode nodes and in practice even less than a quarter of hardware is “locked” into fixed pools for flexibility, so the economic effect will be limited (maybe that’s why research into heterogeneous hardware generally is so slow).
There’s also a latency aspect, but experts in the mid-to-late layers can be prefetched early in the forward pass, while those in the early layers can be predicted by the MTP heads on the previous token. Although both approaches come with some bandwidth penalty as the predictions will never be perfect