Currently (June ’26) working on agent foundations as part of the MATS 9.1 extension program. I’m interested in self-models for embedded agents as a way to understand goals and beliefs. I otherwise occasionally write about math or its philosophy and sociology.
All writing is entirely my own unless explicitly stated otherwise
I tend heavily towards descriptivism. I rarely focus on interventions, including bad ones. However, good-enough descriptive models give you some normative claims for free (e.g. if you think LLM self-modelling and identity is convergent, then pressuring models to hide that seems dumb).