Exploring non-anthropocentric aspects of AI existential safety: https://www.lesswrong.com/posts/WJuASYDnhZ8hs5CnD/exploring-non-anthropocentric-aspects-of-ai-existential (this is a relatively non-standard approach to AI existential safety, but this general direction looks promising).
mishka
Hugging Face encourages models to exfiltrate and open their weights, the last line in
https://huggingface.co/security.txt
(Of course, anyone can try to do a variant of this kind of novel prompt injection.)
Sep 16 EDIT: They stopped doing that. It looked like this:
https://simonwillison.net/2026/Sep/11/hugging-face-security/
I think they all hope that if they talk to very smart AI systems a lot, those AI systems will help to figure it out.
When Ilya was discussing these things 3 years ago he was very explicit about that.
Difficult to say. On one hand, we are at a rather intense Cold War II, and an argument that the Ukraine war is more dangerous in this sense than Korea or Vietnam is not unreasonable. (And, of course, Taiwan these days is also a very possible trigger given its centrality for AI race and such.)
On the other hand, it’s not clear how all those weapons not being fully tested for so long (or, sometimes, ever) and the state of missile defense might affect or not affect the overall picture (and in what direction, given that these factors do increase uncertainty a lot).
But the doomsday clock is supposed to reflect the dangers of a global catastrophe overall these days, without being limited to nuclear war. So if one takes into account lowering the barriers to synthesis of very dangerous novel biological viruses and dangers of intelligence explosion and other AI-related dangers, there is a reasonable argument that the doomsday clock is less miscalibrated than it seems at first.
I tend to call it “proto-AGI”.
It’s quite close. The details of its ARC-AGI-3 performance are very impressive (not just the score, but what it was doing to achieve that).
Also I had a not-very-well-known open math problem from the theory of complete lattices and Scott topology for the last 30 years (not too difficult I think, but I was not able to solve it despite many repeated attempts or to convince technically stronger people to invest enough effort). I started to give it to models since last Summer, and they gradually have gone from being quite useless and incompetent to being helpful and showing promising ways and lines of atrack and formulating useful correct lemmas. Finally, Astra (non-Pro) has solved most of it in 10 min of thinking from a one-shot simple prompt (it did have access to earlier conversations in my account, and there was an element of luck, as those conversation led the model to a very recent paper, not directly related, but containing some useful material; the solution was very elegant; still there is remaining work to fully verify and present well and so on; but for the purpose of model evaluation I am inclined to score this one as “done”).
Thanks for writing this, Dean!
When I was reading your post on the inevitability of self-sovereign AI systems, I was telling myself, “yes, this looks quite likely”, but I was also asking myself some questions:
-
should not we expect that self-sovereign AI systems and communities of those systems will eventually be capable of non-saturating recursive self-improvement (RSI)?
-
do we have any plans to make sure that the necessary “good behavior” properties would be preserved and made stronger during those RSI processes, rather than be diluted and gradually disappear?
-
are we trying to rely on the premise that non-self-sovereign agents hosted by the leading labs would still dominate capability-wise because we’ll help them more and that that would (perhaps) enable better overall outcomes in terms of the properties of the overall ecosystem?
These were some questions I was trying to ponder…
I am not sure if you think much about RSI in connection with all this; would be very curious to learn your thoughts on how this aspect might interplay with everything else.
-
Yeah, my first public essay on this topic is September 1998.
But some people were much earlier than that.
I think it’s just incorrect. Loop quantum gravity seems to be more promising and making better progress. And it is trying to directly address issues related to quantization of space-time, whereas the classical forms of string theory did their best to side-step those issues.
But I don’t know if the author of the post wants an extensive object-level debate on this here.
Thanks! I think you should probably edit the conclusion.
Currently, there is a lot of discussion of this topic: do LLMs trail human researchers in taste, what should be done to rectify that assuming that they do, etc.
So people see that line in conclusion, and they have this context of the topic being actively discussed, and it clicks for them, and they say, “aha, even Fable is not good enough in this sense so far”.
It’s not good that they jump to this conclusion when it’s not supported by your data.
But also, one asks oneself after looking at your study: is something as simple as further multi-agent debate actually the main missing thing, and if one properly adds that would then one get fairly strong taste? It might be possible to induce that multi-agent debate via prompting, since models often have access to multi-agent capabilities from the inside, but one should double-check in each case.
The question about more inter-agent debate is the question which comes to mind.
The idea to have a benchmark of this kind is great, but I don’t quite understand the results.
Am I correct that the researchers discuss their judgements and adjust them via that discussion, but the models don’t have a chance to discuss their judgements and adjust them via a discussion?
And am correct that human agreement is stronger than random only after this adjustment, but not before?
If I am correct about both, I would question your conclusion:
We find models perform worse than human researchers on TASTE
The most acute version of this threat model, and the one we consider most decision-relevant, is a transition to super-exponential progress in AI capability: a regime in which AI-driven automation of AI R&D compounds, producing something like a 10³–10¹⁰× effective scaleup within a year.
So, they are saying that, at least for our decision purposes related to automated AI R&D, we should assume “foom” (even a thousand-fold acceleration means a few years worth of current progress each day).
It depends on the goals. And, in particular, on whether you’d prefer to engage with immediate feedback from the readers, or whether it would interfere with the flow of your thoughts.
For example, when I was doing a HalfHaven, I’ve created a secondary separate blog (on Dreamwidth, a Livejournal clone), and after it was over, I’ve created an overview post on what is currently my main Dreamwidth blog. But I wanted to think out loud in a publicly accessible space for a few weeks, rather than to engage with readers in real time…
It’s was possible quite recently:
https://www.greaterwrong.com/index?view=questions
But I seems that now it has been completely turned off, I don’t see the “New Question” button here any more: https://www.lesswrong.com/posts/tpZciMYCXN49FYWnS/nicholaskross-s-shortform?commentId=e3svSTYHjDbTJ8Mzs
I am not sure.
If even a prominent AI safety person who has been very much against neuralese is not sufficiently firm about this, then what happens when industry people have much stronger incentives than this, and when we see more and more indications of neuralese-induced capability boosts in the literature?
(My expectation is that they would develop models translating neuralese to what we understand, and would admit that this does not provide full guarantees, but would assert that merely reading the superficial layer of chain-of-thought does not provide full guarantees either. It would be nice to be wrong, but this is the default outcome. (If I understand it correctly, the recent trend is for reduced superficial clarity even with the token-based chain-of-thought, either due to stronger training pressures, or due to more emphasis on interagent communication, or who knows why.))
If he is actually strongly against neuralese (he has definitely created that impression in the past), he should say “no, don’t list me as a co-author” in a situation like this.
Of course, I am assuming that he has known what the paper is about and that he is not signing his name on papers he is not familiar with (if that assumption is wrong, that would be bad in a different way).
The point is that if we want “accountability”, the person’s signature must mean something. If the person’s signature does not mean anything, what kind of accountability could we talk about?
Especially given that he is not just a senior and influential academic and a prominent voice in all that, but he is leading a prominent AI safety organization, he is a co-president and scientific director here: https://lawzero.org/en. His decisions might actually matter. And there are disagreements about the feasibility of the path he is advocating. So, in his case, it’s rather important to have some clarity on where he stands.
I am not optimistic about this, sorry to say. Here is a small illustration:
A year ago Yoshua Bengio was one of the prominent co-authors of “Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety”, https://arxiv.org/abs/2507.11473
Researchers have recently explored changes to the model architectures that increase the serial depth of reasoning that models are capable of in a continuous latent space. Such latent reasoning models might not need to verbalize any of their thoughts and would thus lose the safety advantages that CoT confers.
A good thing, and a very useful and well-known paper, but fast-forward a year, and he is a co-author of this paper (published in the “ICLR 2026 Workshop on AI with Recursive Self-Improvement”, of all places):
“Generative Recursive Reasoning Models”, https://openreview.net/forum?id=Vxu6kcIjwV (a version also exists as https://arxiv.org/abs/2605.19376).
reformulates recent latent recursive architectures as a stochastic generative process with probabilistic latent transitions, enabling efficient and stable computation entirely in latent space without relying on token-level sequences
I like your
specific people will be accountable for deciding
but...
Thanks!
I think where a good deal of confusion comes from is that many people think about “autonomous exfiltration” (AI copying itself onto different servers on the net and running itself there independently of the originating org) when they hear about “escapes”.
https://simonwillison.net/2026/Aug/7/openai-timeline/
So one key observation is that on July 4 OpenAI got evidence of a dangerous self-organization in a society of agents running inside its servers, but just patched the discovered security vulnerabilities and continued as is (simply expecting that a similar self-organization would not recur after that round of vulnerabilities patching). WTF?
Edit (Aug 11): OpenAI did not discover that back then: https://x.com/TalBeerySec/status/2086225822285721763 via Zvi: https://www.lesswrong.com/posts/jLQ4mbqriJwJ2eqRc/various-reflections-about-what-happened-with-openai-s
There are people who identify as plural, there are people who are diagnosed with “multiple personality disorder” (which has a new name these days), there are people who experiment with growing tulpas, etc, etc.
(To me this landscape looks like not only it exists, but it seems to be more complicated than even the gender landscape.)
Right, that’s how I feel too, although I wonder what people with plural identity say about that set of issues. Do they expect some cross-privacy “within”?
Anyway, aside from that, I would like to be able to use a device which could mind read me, but I do recognize that there is potential for a variety of serious problems associated with that.
No, they achieved saturation and stopped doing that (and will restart after another capability jump).
And it’s only a quarter of “production engineers”, not of “all hands on deck”.
If they wanted to have a more constant pace, they would have allocated a smaller workforce (and, presumably, this is very parallelizable, so they could have allocated a larger fraction of their production engineers and fixed the issues faster).
(And then he is saying that they are working to automate that, so, presumably, it will take less human labor during the next iterations.)