Currently doing agent foundations at Resolution. In 2023 I was on Vivek’s team at MIRI, before that I did MATS 2, and before that I did a CS and Neuroscience undergrad (thesis on statistical learning theory).
The best summary of my AI & alignment beliefs is my corrigibility basin of attraction post.
I’m interested in doing in-depth dialogues to find cruxes. It takes a while but it’s a great way to highlight broken parts of my beliefs and good practice for communicating clearly in general. Message me if you are interested in doing this.
Xavier Roberts-Gaal sent me a Bayesian analysis of this data that I like more than my own:
https://drive.google.com/file/d/1OucvxNv6q4fTB3Qm-u5923xhD2Ei8mPZ/view?usp=sharing