I research intelligence and its emergence and expression in neural networks to ensure advanced AI is safe and beneficial.
I’m currently a Research Scientist at UK AISI working on training and interpreting model organisms of misalignment — such as of reward hacking, evaluation awareness, and sandbagging.
For more, check out my scholar profile and personal website.
Thanks! Yes, I agree that the cheap regex proxy can be wrong sometimes, and we thus do LLM-based monitors too for the important runs (see Figure 3). We don’t do this always because it can slow training. Since then we’ve explored deeper into this KL-induced-unfaithfulness and have a new post on it that provides more clarity: https://www.lesswrong.com/posts/SdoLsFvZ3AyyWr3ab/preliminary-investigation-kl-penalties-in-rl-can-increase