I’m an Assistant Professor at Carnegie Mellon’s Machine Learning Department. I’m also a core faculty member in CMU’s Neuroscience Institute, and hold a courtesy appointment in the Robotics Institute.
My lab works at the intersection of neuroscience & AI to reverse-engineer animal intelligence and build the next generation of autonomous agents, responsibly and safely.
Learn more here: https://cs.cmu.edu/~anayebi
I very much agree regarding corrigibility being the most tractable safety target for alignment. I discuss some formal results proving this claim, in case it’s of interest here: https://www.lesswrong.com/posts/M5owRcacptnkxwD2u/from-barriers-to-alignment-to-the-first-formal-corrigibility-1