AI security via formal methods
Quinn
i’m relatively shocked that anthropics / bostrom simulation got so vindicated. It looked like goofy inside baseball only relevant to philosophers at the time, then it became important to reason about how AIs might guess whether they’re in training or deployment
i use the Tab Nodes Tree extension, which i quite like, but i’ll try native.
am i stupid or is my profile not showing my total karma anymore?
i’m confused about giving career advice to people trying to get into AI safety. I do not want to be known as a well connected guy who will tell you the words of power so you can manipulate the funding streams and fellowships into launching your career, and then you just go into a frontier company anyway.
my EA career advisor friend has a story of someone saying out loud “safety is the easiest way to get into AI [and i want to work in AI for normal money/status reasons]”. [EDIT: she would like me to clarify that this was literally once out of hundreds of calls]. This isn’t just CitizenTen’s the vultures are circling (though that post will make a comeback with Anthropic DAFs), this is a new thing.
Anyways: i’m dispositionally here to help, it would take a lot of effort for me to not be an attack surface for these folks, but i’m considering becoming more discerning.
What do you guys think? How do you vet peoples’ authenticity? Is taking on conformity risk (i.e. by judging people based on how old their LW account is) the lesser of two evils? My friend and I joke that I have a better relationship with OpenPhil than he does because i’m a kidney donor.
Quickly:
AI-accelerated fuzzers, which rapidly self-improve their fuzz harnesses at runtime.
Forall R&D was funded to investigate this hypothesis and came back with a negative result. My current belief is that this isn’t super powerful. I do gotta work on a writeup to convey that pessimism to others!
AI Security is Harm Reduction
makes sense—my razor for “this is like the evil social media notification slop version of LLMs” vs “i’m uplifting in a way that is good for me” is whether or not I can defend an artifact by myself on a whiteboard. I almost never meet this bar. So when I use it to teach me some new tool, then assemble that tool into a solution to a problem, I would ideally slow down enough to let my own alone-at-a-whiteboard command over the material catch up. Perhaps we agree that finding time to do this is the analogue of not letting social media notifications fry your brain.
(I think my work is less conceptually complicated or interesting than your work)
I thought about adding commentary but decided the quote’s recontextualization speaks for itself.
That shortform (which John may or may not have remembered writing) was living in my brain ever since I read it, as a poweruser who occasionally has pangs of “what if enfeeblement?” and so on. So to me, the post about how things shifted as of the last few months feels something shaped like “the one guy we could count on to resist enfeeblement is now uplifted” which feels like a “sign of the times” of some sort.
A year ago, you wrote:
People would sometimes say something like “John, you should really get a smartphone, you’ll fall behind without one” and my gut response was roughly “No, I’m staying in place, and the rest of you are moving backwards”.
And in hindsight, boy howdy do I endorse that attitude! Past John’s gut was right on the money with that one.
I notice that I have an extremely similar gut feeling about LLMs today. Like, when I look at the people who are relatively early adopters, making relatively heavy use of LLMs… I do not feel like I’ll fall behind if I don’t leverage them more. I feel like the people using them a lot are mostly moving backwards, and I’m staying in place.
brief memo on microFROs
focused research organizations (FROs) were conceived (by Convergent et al) as “startup nonprofits” on a deadline, they’re planned to exist for 4-7 years or so to fight scope creep for specific missions.
One relatively obvious objection is with the world changing so much, how can you plan 5 years of “focus” up front? 5 years may be too long to successfully articulate a focus that makes sense for even that long.
I would propose microFROs, ~90-200 day sprints in the spirit of a FRO that spin up, kick ass, and spin down.
I don’t think Leveson style safety engineering is a great fit for alignment, but I do think its a fit for a lot of AI security efforts (like SL5 datacenters).
June-July 2026 AI Security via Formal Methods
Flipping the eval on its head
great work, love the “low lift” vibes
registering a prediction: you will regret not putting a “short title” field when you’re managing a spreadsheet later
(there’s obviously lots of cheap boutique software you could vibe up to make the backend easier, as you know)
Patching ~All Security-Relevant Open-Source Software? [niplav 2025]
Common application for AI safety funding
just a note: i think this already exists, its called a Google Doc. My last couple grants were shopped around to a few funders, i would make a copy and change the title to reference the target funder when i’d send it around (but in one case a funder received the “Manifund draft” version of the gdoc and funded it anyway)
Apr-May 2026 AI Security via Formal Methods
given all the known unknowns and unknown unknowns around this stuff, plausible enough to me that anthropic RSI techs are in the same bind.
Like yes based on public information the story that openai RSI techs are in this bind is much cleaner, but I expect like 80-90% of the relevant knowledge to be private.
GDM, idk—they seem less RSI pilled than the other two (but again, what do I know)
is “bad writing” a meatcert?
I feel like half the reasons people tell me they like my newsletter is that I basically write how I talk, which is a voice, which is a “meat certificate” (reasons my readers believe I’m not using AI too much), which has gotta be an important part of my relationship with readers since its dead easy to make a slop current events newsletter anymore.
The thing is—I have memories dating back to the beforetimes of people telling me my verbal tics were “bad writing”. Often people line edit my writing to turn it into “good writing” and it hobbles the voice. Slop itself and especially our reactions to it are kinda proof that “good writing” in its extremes is grotesque or offputting.
I’m wondering how far up this goes.
the effortpost version of this would have examples, i’m jotting down a quick note instead
i’m hiring someone with a logician/formal methods skillset who has lastmile energy for cobbling together a bunch of parts off the shelf and getting it into the hands of a user, for SL5 work. Maybe workload verification work soon.
DM me for JD