newsletter.safe.ai
Dan H
ML Safety Newsletter #20: AI Wellbeing, Classifier Jailbreaking and Honest Pushback Benchmarking
AISN #71: Cyberattacks & Datacenter Moratorium Bill
AI Safety Newsletter #70: Automated Warfare and AI Layoffs
AI Safety Newsletter #69: Department of War, Anthropic, and National Security
MLSN #18: Adversarial Diffusion, Activation Oracles, Weird Generalization
Thank you to Neel for writing this. Most people pivot quietly.
I’ve been most skeptical of mechanistic interpretability for years. I excluded interpretability in Unsolved Problems in ML Safety for this reason. Other fields like d/acc (Systemic Safety) were included though, all the way back in 2021.
Here’s are some earlier criticisms: https://www.lesswrong.com/posts/5HtDzRAk7ePWsiL2L/open-problems-in-ai-x-risk-pais-5#Transparency
More recent commentary: https://ai-frontiers.org/articles/the-misguided-quest-for-mechanistic-ai-interpretability
I think the community should reflect on its genius worship culture (in the case of Olah, a close friend of the inner circle) and epistemics: the approach was so dominant for years, and I think this outcome was entirely foreseeable.
MLSN #17: Measuring General AI Abilities and Mitigating Deception
AISN #65: Measuring Automation and Superintelligence Moratorium Letter
This dynamic is captured in IABIED’s story and this paper from 2023: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4445706
AISN#64: New AGI Definition and Senate Bill Would Establish Liability for AI Harms
AISN #63: California’s SB-53 Passes the Legislature
just rehearsed variations on the arguments Eliezer/MIRI already deployed
I think they’re improved and simplified.
My favorite chapter is “Chapter 5: Its Favorite Things.”
A key part of this strategy is deterrence, “Mutually Assured Compute Destruction” which gets its own section. It doesn’t mention the generalization Mutually Assured AI Malfunction (MAIM) from Schmidt, Wang, and me last year. This also spends 1⁄3 of the MAIM discussion on verification and how to do this in a multilateral way.
Meanwhile it cites other works like A Narrow Path. I even left this feedback to you all at AI Futures before this was released. This would constitute plagiarism in any other context. It’s a bewildering unforced error—it’s extremely related, it’s a certainly a nontrivial idea, and I told you all this recently—I hope you all fix it.