AI security via formal methods
Quinn
I don’t think Leveson style safety engineering is a great fit for alignment, but I do think its a fit for a lot of AI security efforts (like SL5 datacenters).
June-July 2026 AI Security via Formal Methods
Flipping the eval on its head
great work, love the “low lift” vibes
registering a prediction: you will regret not putting a “short title” field when you’re managing a spreadsheet later
(there’s obviously lots of cheap boutique software you could vibe up to make the backend easier, as you know)
Patching ~All Security-Relevant Open-Source Software? [niplav 2025]
Common application for AI safety funding
just a note: i think this already exists, its called a Google Doc. My last couple grants were shopped around to a few funders, i would make a copy and change the title to reference the target funder when i’d send it around (but in one case a funder received the “Manifund draft” version of the gdoc and funded it anyway)
Apr-May 2026 AI Security via Formal Methods
given all the known unknowns and unknown unknowns around this stuff, plausible enough to me that anthropic RSI techs are in the same bind.
Like yes based on public information the story that openai RSI techs are in this bind is much cleaner, but I expect like 80-90% of the relevant knowledge to be private.
GDM, idk—they seem less RSI pilled than the other two (but again, what do I know)
is “bad writing” a meatcert?
I feel like half the reasons people tell me they like my newsletter is that I basically write how I talk, which is a voice, which is a “meat certificate” (reasons my readers believe I’m not using AI too much), which has gotta be an important part of my relationship with readers since its dead easy to make a slop current events newsletter anymore.
The thing is—I have memories dating back to the beforetimes of people telling me my verbal tics were “bad writing”. Often people line edit my writing to turn it into “good writing” and it hobbles the voice. Slop itself and especially our reactions to it are kinda proof that “good writing” in its extremes is grotesque or offputting.
I’m wondering how far up this goes.
the effortpost version of this would have examples, i’m jotting down a quick note instead
My answer is that I think CAIS-like scenarios necessarily precede agents-running-amok scenarios, and furthermore
I wonder how to grade this november 2022 forecast? (CAIS as in comprehensive AI services). it seems like i got it really bad, if you compare the openclaw pandora’s box to some of the stuff i recall in the CAIS document about narrow products proliferation?
Apart has a Secure Program Synthesis fellowship coming up, mentor application deadline is May 5th
https://apartresearch.com/fellowships/the-secure-program-synthesis-fellowship
Focus Areas
Specification Elicitation: Develop tools and workflows for extracting formal specifications from ambiguous, distributed, or implicit sources (e.g., documentation, legacy systems, human stakeholders). Projects may include structured editors, GUIs, or pipelines that translate informal requirements into formal representations (e.g., Lean), building on approaches like “SpecIDE.”
Specification validation: Design methods to verify that extracted specifications are correct and complete. This includes techniques for testing, cross-checking, or formally validating whether a specification accurately captures intended system behavior.
Spec-Driven Development & Evaluation: Explore workflows where specifications generate multiple candidate implementations, which are then evaluated against each other. Projects may involve building infrastructure to compare robustness, correctness, or performance across implementations derived from the same spec. For further reading, see: Approximately Aligned Decoding, Zero-DoF Programming
Adversarial Robustness for FM & QA Tools: Investigate failure modes in formal methods pipelines, LLM-assisted tooling, and QA systems. Projects may focus on adversarial inputs, robustness guarantees, or identifying weaknesses in automated reasoning and verification systems.
...
What’s in it for the mentors?:
A team executing your research vision. You propose the project; an Apart project manager and mentees run it with you. You’re directing the research, not running the project solo.
A dedicated project manager. An Apart Research Project Manager (RPM) handles the operational side of running a research team, so that you can focus on the research direction.
A working trial of potential hires. Historically, organizations running programs like this have used them functionally as work trials. Over four months you’ll work closely with mentees who’ve already passed our bar, giving you a clear read on who might be a fit for your team.
Project resources for the team: compute, and API credits.
Possibly a stipend. By default mentor roles aren’t paid, but if compensation would enable your participation, indicate this in the application form.
please email any questions to
secure-program-synthesis-fellowship@apartresearch.comthere will be an associated hackathon late may
[exploding note] Apply to Mentor Secure Program Synthesis Fellowship by May 5th
EDIT/update: i just remembered a key detail of the trial which is that one of the reasons we won was a key prosecution witness (who said things which, to the best of my knowledge, were largely made up) got character-assassinated (the defense found evidence of mob ties) and that’s part of why his testimony got thrown out. So character assassination played a role there, just not directly about the defendants.
jumping from press releases by lawyers to how criminal trials work is a jump. Clearly, inside the courtroom could be different from the press releases!
I was a defendant in a highly politicized free speech case and the lawyers on both sides made a couple gestures at character but the judge ignored them. overall the process was razor focused on establishing timelines and facts.
Yes obviously before and after the press releases focused on morality and character, which wasn’t really the crux of the outcome.
sometimes i wonder if the extent to which you’re a live player is your proximity to a frontier company. Feels that way sometimes! I think there’s a third thing he could mean, which is that the orgs that aren’t anthropic or close partners aren’t live players [1] .
(note: even if interacting with frontier companies is a reliable way to increase your magnitude, that doesn’t mean you’d be able to reliably predict the sign).
- ↩︎
seems false in governance, awareness/public opinion/forecasting, policy. Pretty plausible for technical work though
- ↩︎
Can We Secure AI With Formal Methods? January-March 2026
How to Solve Secure Program Synthesis
i’m pretty confident that a lot of jobs will be resistant to robots for a while, i expect there to be diminishing returns on raw cognitive skill when there’s the long feedback loops of robot invention and manufacture. Lots of people have to fix a mechanical widget in a very dusty or muddy or wet environment, people scuba-weld, etc.
What I’m confused about is will those wages, if the white collar wages go to zero, go up or down? I expect them to be better off than 90% of SWEs, but also society at large is getting a lot richer, and more people will be willing to do those dirty jobs. Lots of factors! I don’t know what to think.
can someone explain why Simulator theory was written about end of 2022, people semi forgot or moved on, then Personas made a big splash recently?
brief memo on microFROs
focused research organizations (FROs) were conceived (by Convergent et al) as “startup nonprofits” on a deadline, they’re planned to exist for 4-7 years or so to fight scope creep for specific missions.
One relatively obvious objection is with the world changing so much, how can you plan 5 years of “focus” up front? 5 years may be too long to successfully articulate a focus that makes sense for even that long.
I would propose microFROs, ~90-200 day sprints in the spirit of a FRO that spin up, kick ass, and spin down.