Principal Investigator representing Leela AI at NIST’s AI Consortium. Recently contributed a formal response to NIST’s RFI on AI agent security, covering autonomous action risk spectrums, positive alignment as a security paradigm, and emergent goal formation in agentic systems.
Previous life: over 25 years in IC/silicon design at HP, National Semiconductor, and AMD — an experience that shapes how I think about verification and alignment (if you’ve ever tried to prove a chip correct before tapeout, you understand why I’m skeptical of post-hoc safety testing). PhD in machine learning (CSU 2022), MS from MIT 1989.
Research interests: AI safety and alignment, positive behavioral benchmarks beyond harm avoidance, world models as a path to robust agency, and the parallels between silicon verification and AI evaluation. Interested in how we build systems that are aligned by construction rather than aligned by patch.
Based in Fort Collins, CO
Overall I like this plan and think it raises good points and shows possible paths. Thanks for the great work.
1) The pause time shown is from 2035 to 2040, which feels OK. One concern I have with longer pauses is the usual innovation of HW and SW, which tend generally to range from 100 to 1000X per decade each. So a 5 year pause means all non-controlled AI get 100-1000X more compute per dollar—systems that cost $1 billion to train and run well drop to 10 or even 1 million, which adds to the problem of enforcing a ban. You may want to note this cost reduction in the 2037 or 2038 discussion of Plan A.
2) I agree that working to build an aligned AGI that is aligned well enough to survive recursive self improvement through to ASI is key and we should be starting work on it now. I would recommend citing Hinton’s comments about maternal instincts and Contemplative Wisdom for Superalignment as possible examples.
3) I think it’s worth mentioning the potential that AI’s either achieve consciousness or it becomes clear that we can choose this through design. If AI’s ‘naturally’ achieve subjective experience as they pass through true AGI, then their welfare becomes important. If it turns out that with selective training or design we can ‘choose’ whether AGI is conscious, we should talk about and plan for how we collectively want to handle that. Of course this topic could become a huge discussion post of its own, and I don’t think you need to detail out all the ways this plays out, but I think it should be mentioned with a paragraph.