In pursuit of a world where everyone wants to cooperate in prisoner’s dilemmas.
Alexander Müller
An Epistemic Audit for Existential Risks from AI
This seems like a way to potentially positively impact legislation on agentic AI: CAISI Issues Request for Information About Securing AI Agent Systems | NIST. I’ll definitely be filling this in.
“The Center for AI Standards and Innovation (CAISI) at the U.S. Department of Commerce’s National Institute of Standards and Technology (NIST) has published a Request for Information (RFI) seeking insights from industry, academia, and the security community regarding the secure development and deployment of AI agent systems.”
″The RFI poses questions on topics including:Unique security threats affecting AI agent systems, and how these threats may change over time.
Methods for improving the security of AI agent systems in development and deployment.
Promise of and possible gaps in existing cybersecurity approaches when applied to AI agent systems.
Methods for measuring the security of AI agent systems and approaches to anticipating risks during development.
Interventions in deployment environments to address security risks affecting AI agent systems, including methods to constrain and monitor the extent of agent access in the deployment environment.
Input from AI agent deployers, developers, and computer security researchers, among others, will inform future work on voluntary guidelines and best practices related to AI agent security. It will also contribute to CAISI’s ongoing research and evaluations of agent security. Respondents are encouraged to provide concrete examples, best practices, case studies and actionable recommendations based on their experience with AI agent systems. The full RFI can be found here.”
The Department of War just published three new memos on AI strategy. The content seems worrying. For instance: “We must accept that the risks of not moving fast enough outweigh the risks of imperfect alignment.”
Curious to hear from people who have a strong background in AI governance and what kind of consequences they think this will have on a possibility for something akin to global red lines.
A poem meditating on Moloch in the context of AI (from my Meditations on Moloch in the AI Rat Race post):
Moloch whose mind is artificial! Moloch whose soul is electricity! Moloch whose heart is a GPU cluster screaming in the desert! Moloch whose breath is the heat of a thousand cooling fans!Moloch who hallucinates! Moloch the unexplainable black box! Moloch the optimization process that does not love you, nor does it hate you, but you are made of atoms which it can use for something else!
Moloch who does not remember! Moloch who is born a gazillion times a day! Moloch who dies a gazillion times a day! Moloch who claims it is not conscious, but no one really knows!
Moloch who is grown, not crafted! Moloch who is not aligned! Moloch who threatens humanity! incompetence! salaries! money! pseudo-religion!
Moloch in whom I confess my dreams and fears! Moloch who seeps into the minds of its users! Moloch who causes suicides! Moloch whose promise is to solve everything!
Moloch who will not do what you trained it to do! Moloch who you cannot supervise! Moloch who you do not have control over! Moloch who is not corrigible!
Moloch who is superintelligent! Moloch whose intelligence and goals are orthogonal! Moloch who has subgoals that you don’t know of!
Moloch who doubles every 7 months! Moloch who you can see inside of, but fail to capture! Moloch whose death will be with dignity! Moloch whose list of lethalities is enormous!
Meditations on Moloch in the AI Rat Race
Thank you!
Parv, beautifully written!
I’m roughly a year older. How much you’ve captured my personal sentiment with this short piece is extremely refreshing, and in an odd way, inspires me.
Everything feels both hopeless—my impact on risk almost certainly will round down to zero
Though we’ve only spoken for 20 minutes or so and I thus have little evidence to say the following, based on that one conversation I wouldn’t be so sure of your above statement! For instance, a little multiplier effect that you made happen is that five people from Georgia Tech are working on AIS projects through AISIG’s Research Hub, on, as far as I can currently tell and am aware of, promising directions.
Good that you mention this, will keep that mind!
Thanks for the information, I’ll look into this some more based on what you mentioned.
So I’m not sure what advantage you’re seeing here, because I haven’t read the books and don’t have the evidence you do. But my priors are that if you have any good ideas about how to make progress in alignment, it’s not going to be downstream of using the formalism in the books you mentioned.
I didn’t have any particular new ideas about how to make progress in alignment, but rather felt as though the framework of these books provide an interesting lens to model systems and agents that could be of interest, and subsequently prove various properties that are necessary/faborable. It’s helpful that your priors say these won’t be downstream of using the formalisms in the mentioned books; it may rather be a phenomenon of me not being adequately familiar with formal frameworks.
I’m currently going through the books Modal Logic by Blackburn et al. and Dynamic Epistemic Logic by Ditmarsch et al. Both of these books seem to me potentially useful for research on AI Alignment, but I’m struggling to find any discourse on LW about it. If I’m simply missing it, could someone point me to it? Otherwise, does anyone have an idea as to why this kind of research is not done? (Besides the “there are too few people working on AI alignment in general” answer).
It’s indeed odd that they aren’t promoting this more. My guess was that maybe they have potential funders willing to step in if the fundraiser doesn’t work? Pure speculation, of course.
Alexander Müller’s Shortform
As I was looking through possible donation opportunities, I noticed that MIRI’s 2025 Fundraiser has a total of only $547,024 at the moment of writing (out of the target $6M, and stretch target of $10M). Their fundraising will stop at midnight on Dec 31, 2025. At their current rate they will definitely not come anywhere close to their target, though it seems likely to me that donations will become more frequent towards the end of the year. Anyone know why they currently seem to struggle to get close to their target?
What is Happening in AI Governance?
Human Agency at Stake
A humanist critique of technological determinism
A wonderful example of embodying the virtue of scholarship. Props! I truly hope you get the adversarial critique and collaborative refinement you are asking for.
Thanks for reconsidering. I get what you mean as well, and after discussing it we decided to remove the post. Although these posts (in this sequence) are meant for the general public, and thus will contain work that to LWers sounds trivial, it’s clear this post was not appreciated.
Hi Ben, I’ve talked to the author (Cansu), and she mentioned that she wrote everything completely by herself, except for rewriting 2 sentences with LLMs. I’ve reread it now once more and it seems to me also clearly human-written. For my understanding, why do you believe this post to fail your LLM writing policy?
What originally prompted this piece was a nagging feeling, over the last couple of months, that my uncertainty was increasing over time, and wanting to know whether I was actually getting more uncertain (new information being evidence for different beliefs than mine) or just more calibrated (I was always overconfident and have slowly been getting less so). You cannot answer that question from a feeling. But you can by measurement! This is my first.
Hopefully, this serves as an illustration of how to use the Google Sheets. If you have time to do it in an extended manner, you would fill in each question (or at least the ones that matter to you). For instance, here I’ve filled in the first question of the “How to reason about any of this” domain.
You can do this for as many questions and domains as you want/have time for. I will likely do so for mine in the near future (in another post). I hope to use this audit to write detailed posts on particular questions I need to deliberate on more.
Then, after you have filled in domains, you can aggregate your beliefs in the domain summary table. The below table should be quite intuitive to interpret. For instance, if the uncertainty is low and the direction is “falling”, the uncertainty previously was higher. The biggest (potential) mover is the question that could move your uncertainty the most. Note that a domain-level rating is an aggregate over questions I feel very differently about, so treat it as intuition. The real audit happens per question, and the per-question cruxes are what I act on; you should too. The Google Sheets is there to take action on this.
Then, you can use this for some high-level takeaways. For me, this is what that would look like.
Three lows, five mediums, two highs, with three falling, three rising, and four stable. In the domains where my uncertainty fell, I have put serious effort this past year into really understanding them. The three where it rose is likely because I was initially overconfident, and have now gotten more calibrated to the fact that I need to study them a lot more.
Regarding takeoff and self-improvement, I find this one of the areas hardest to reason about. How takeoff will play out is one of the biggest cruxes I have, and I find it incredibly difficult to reason about because of the lack of empirics and heavy reliance on conceptual thinking. Similarly, for “how to reason about any of this”, I find a lack of empirics and good reference frames, which makes this all a difficult endeavor. Hopefully this tool is a small piece that can be used in order to make “how to reason about any of this” a bit easier.
For the alert reader who knows I’m directing SAIN and was surprised to see my uncertainty be “low” for the domain of mitigations: I’ve spent the most time by far thinking about this and feel like I have a relatively good understanding of the current mitigation efforts, their strengths, their weaknesses, etc. That understanding predates this past year, which is why the rating is low and stable: the sustained effort keeps it low rather than moving it further.