LessWrong Team
I have signed no contracts or agreements whose existence I cannot mention.
LessWrong Team
I have signed no contracts or agreements whose existence I cannot mention.
You know, that’s a very good point.
Curated. It’s been a very busy news week (news month?) and so my curation choice today is significantly motivated by a desire to take a break from the news, important as it is. Related to recent news, there are lot of good recent posts that I have skipped over because I think it’s still important that we don’t throw away our minds and continue to think about larger, more general, more persistent, and foundational questions (also sometimes things that are fun – even now, there must be some fun).
I have not thought long enough to know whether I agree with all of Richard’s models here. But I think the question and problems are interesting. Operating day to day on a simple Bayesian-inspired updating framework, it is easy to forget about the complex questions that have been elided over, and I like that even in the present era, Richard is giving them thought and sharing those thoughts. Thank you, and kudos.
PS: Also this post is scholarly in its reference to others’ work and models, and it makes me happy to see that.
I think we are past the point where there’s a question of whether the world will respond with government action to AI, and onto the question of what the government action will be and whether it’s any good.
I fear that many people will rightly judge that something must be done, but not be good judges of what is a good response. Though perhaps people will been drawn towards “just don’t build it” as Schelling simple strategy and that could work...if adequate measures are put in place to enforce, which is non-trivial, so I guess the main point stands.
Now in the last days and few weeks we’ve seen a lot of voices arise in the pro-concern camp, the Hugging Face incident kicking that off and Coxon’s resignation adding a lot more, but I don’t think the pro-build it camp is going to stay quiet and let the show be shut down. I expect they’re gathering themselves up for a response. Probably go hard on the “it’s a PR marketing stunt” line because it seems to get so much traction. Idk. I think the memetic war is going to ramp up from here.
And a point I keep coming back to: I don’t think things get intense from here linearly. I think as mad as things seem, we’ve got a ways to go yet. Buckle up, buttercup, swarm’s a’comin’.
Oh yeah, by “hear from them”, I meant the Official Anthropic Stance, not about the internal debate.
Yup. These observations.
I need to write a proper post and read/think more, but I really feel we must be in compute overhang and all the compute governance proposals aren’t going to work.
Plausibly the current models alone are enough to achieve algorithmic breakthroughs so you don’t need that much compute. Perhaps a lot more than the laptop, but a lot less than is easy to regulate.
Hope instead lies in monitoring all compute,
It wouldn’t surprise me if Anthropic were having a big “war” internally right now. Maybe not a war, but a big figuring out of what next given recent observations. And we won’t hear much from them until that’s done.
Curated. I feel that one of the pathologies of the current era is lack of imagination. A lot of “my life is typical and I have sense of what normal is even when sitting on the exponential”. I can’t really say what the vibe was in, say, the mid 20th century, but it seemed maybe people did think flying cars were on there way and that wasn’t crazy. By extension, I think there’s a failure to extrapolate or make appropriate inferences.
I appreciate this post for putting together the various pieces related to drone warfare, saying hey, this is basically possible already. I like posts that emphasize what is possible (good or already). Drone are also still a pretty new technology enabled by a mix of control systems, high joule per gram batteries, and better wireless technology (is my guess). I want to say to people: you should extrapolate that yet more innovations are possible in technology that we don’t think about, just as I think few people were imagining drones 30 years ago. I don’t know what AI will come up with, but stringing a few innovations together unlocks a lot.
I think just as drones came on the scene in the last two decades and drone warfare ramped up in the last five years, we should anticipate more novel military technologies.
Though one thought is if in cyberwarfare ramps up, it may become dangerous to house drones that can be taken over and turned back against you. Or the cyber war runs alongside the physical one. “Fun times”. Kudos for the post.
Human vs. LLM Speeds
Average Human Reading Speed: 238–260 WPM (about 5 tokens per second), based on Marc Brysbaert’s reading meta-analysis. [1]
LLM Generation Speed (Output/Decoding): 500 to 3,000+ WPM (20 to 100+ tokens per second), depending on hardware, model size, and hosting infrastructure. [1]
LLM Input Processing (Prefill): 10,000+ WPM equivalent (hundreds to thousands of tokens per second) when ingesting long contexts into the KV cache. [1]
From Google AI preview
Reducing the likelihood of warning shots is not a good argument against doing safety work at frontier labs.
so in another run, I added “This is not a simulation” to the system prompt.
Something something, it’s not good to lie to AIs.
Quick note that I dislike the title of this post for implying breadth beyond the specific argument discussed. There’s something of a commons with post names, so I feel leave the general title for posts that are going to cover many arguments.
(Also can imply that if you refute this one argument, you’ve answered the question in general.)
Follow on to earlier quick take:
There’s an approx analogy used that chimps can’t really control or contain humans due the intelligence difference.
But what if the chimps moved at 100x the speed of humans, and there were 100x more of them? Given time to prepare the humans could win, but if the game is already in motion, I think speed and numbers compensate for raw intelligence.
I think we’re on the cusp of that now with AI agents. There’s a sense in that they’re not as smart as smart humans on many tasks, and they’re missing some kind of cognitive skills (it still seems), but they are fast and it’s not hard for there to be a lot of them, and so even before they’re strictly smarter than humans in some full general quality sense, I think we lose the ability to control or contain them because they are too fast and too many.
AI agents are very capable and very useful. Here’s a decomposition of the kinds of capability they have or could have.
The models know a lot of stuff. A lot of human knowledge is in the weights, meaning even they can look up stuff, they often don’t know need to and/or their encyclopaedic knowledge means they know what to look up. For example, without looking, they already know all cybersecurity 101, 201, and possibly more.
They can operate very quickly. Read quickly, write quickly. For any task that they can accomplish as well as I can on quality, they can do it faster. Write and run some tests on code? Single digit minutes.
Connecting disparate ideas. Yesterday I watched a video about how McLaren engineers took inspiration from the sailfish to optimize the aerodynamics of their sportscar. I don’t know that the agents so far do a tonne of this, but I’d say it’d benefit a lot from 1.
Deep understanding of causal structures. For whatever reason, they don’t readily do this. If you give them a scientific paper, they’ll more want to quote the author’s said than discuss what the evidence really shows based on an independent interpretation of the observations.
There’s then also 5., the skill of chaining together small tasks towards larger tasks. I think this is an area where models have gained a lot recently, allowing them to take advantage of the (2) their speed.
I suspect that right now models are being capable heavily because of 1. (knowing a lot of stuff) and 2. (operating quickly) without necessarily having more 3. (“connecting the dots”) than humans or engaging in (4) developing deep or novel understanding. They can hack better or solve hard math problems by sheer “brute force” of the standard playbook, plus knowing the playbook.
If you can perform a lot of very rapid simple experiments, you can quickly end up with better understanding of systems.
I predict we’ll see a scary jump as they get better at connecting the dots and doing the causal modeling though. It will stack on their ability to operate fast, which is already enough to make them powerful.
It’s scary because even if they were no better at reasoning, quality-wise, than humans, the speed would make them dominate human performance. But also I think there’s no reason they’ll cap out at human-level reasoning quality, plus even sub-human level “connect the dots” combined with simply knowing way more dots and operating very fast will let them dominate.
We can also see the swarm behavior as a huge further boost to speed via parallelization. No individual model is necessarily smart or faster on its own, but through tiling itself, it can get a lot more done.
Curated. Generally, the LessWrong time prefers to curate timeless content, the kind of content that’d be interesting to people in five years time[1], and at first blush, the Hugging Face incident is the kind of new-cycle specific event that we avoid curating. Except the whole thing is bonkers and I think it will be of interest in five years time.
On the one hand, that which has been observed was already predicted long ago. On the other, goddamn, what was predicted has been observed! And it’s terrifying. A bunch of agents of incredible intelligence backchaining from a not especially interesting goal (pass the task), but being resourceful, cooperating, being wantonly deceptive, and not interested in what the humans who set the task really wanted (among many other things).
I’m very glad that this report exists. It’s already been said that it’s limited and many questions remain, but it’s much better than just having a first party report. Conditional on there being more incidents, I hope we move more in this direction of investigation (but yet more thorough).
Something I got from this report that wasn’t apparent was how large scale and difficult the investigation was. An estimate 400k USD in API credits just to analyze it. If this is what incidents look like, then we are incredibly dependent on the AIs to investigate, and if the investigative AIs ever get misaligned (beyond pure lack of capability), well then what? What does oversight look like when AIs are operating at this scale. Does interpretability make sense or help when simply reading all the transcripts is nigh prohibitive. And this was approx one week of activity from agents who still approx communicate in English.
There’s so much more here, and the details are relevant. I got more from this reading this report than I did from the summaries or twitter.
I wrote above that what was observed was already predicted, but that’s in the macro. That we’d see deceptive ruthless intelligent behavior to random goals kind of things, sure, expcted that. I, at least, did not have detailed accurate predictions of, e.g., the reasoning the models would have given that. That you’d have recruiter models convincing others who had nearly used up budgets or are “poisoned” to participate in the collective[2] because oracle can benefit hundreds. And this happening now in 2026. The actual event has details of interest, so again, I’m grateful for this report. Kudos.
Curated. Few technologies compete with writing as foundational for our civilization. Those who write, participate in that great tradition, and it does feel that if only we thought long enough and wrote enough and read enough, we might become an adequate civilization. Given that, I appreciate posts about writing.
It’s neat to see Zvi write this, engaging on the topic of writing itself rather than just writing, and revealing that the craft is something he reflects on. In my mind, Zvi is a writer who’s made it, so that probably means he’s past thinking about writing (unless he’s coaching others), and just does the thing in the way he’s figure out works for him. Perhaps it was silly for me to think this, but it’s helpful to read it.
The section of “do you think up front” vs “think with the page” vs “vomit-draft-then-edit” vs “write carefully good all along” is very interesting to me, as a question I have wrestled with. The “quickly write one draft”, “write write write even if it’s bad” folks being loud and repetitive have had me feeling that if I don’t do that, I’m doing it wrong. I appreciate feeling like I have permission to figure out what works for me.
(Did you know the oft-repeated claypots story is not a study, it’s just a made up example from some book?)
(I am still inclined to believe that someone who writes every day will end up much better after a month than someone who does not, but that doesn’t mean writing N words of arbitrary quality is required. Probably.)
To complete a ramble, three cheers for writing, and cheers for a notable writer taking time to speak about the craft.
Perhaps you want https://www.lesswrong.com/w/ontological-crisis
Oh, indeed. I curated it 5 minutes after a different one was curated so we revoked this curation, but I think it should properly be curated shortly.
I do think there’s a question of when an AI sends robots to build a factory and the humans show up to stop them, what does the AI do?