Evan Hubinger (he/him/his) (evanjhub@gmail.com)
Head of Alignment Stress-Testing at Anthropic. My posts and comments are my own and do not represent Anthropic’s positions, policies, strategies, or opinions.
Previously: MIRI, OpenAI
See: “Why I’m joining Anthropic”
Selected work:
I think it’s actually an easier battle to fight to pull some fancy new neuralese thing out of the next model training run once you already have clear empirical evidence that it is in fact bad for monitorability. I think it’s harder to try to get something pulled before you actually have concrete empirical evidence that it’s bad.