Really not sure what heuristic leads you to count people working on ARC-Theory working on an ambitious, speculative version of interp as working on alignment but not any of the people working to build from current interp paradigms. Similarly, anyone working on e.g. making models more honest in prod models is in fact learning a bunch of lessons about what scalable oversight looks like (albeit not publishing, which i agree is sad). Or doing any science of misalignment, or doing any empirical character work, or experimenting with making models adhere to a spec, or carefully understanding their generalisation patterns, or just trying to understand what the actual objects that we are creating right now are??
It seems like having any current interaction with frontier models is seen as disqualifying for actually doing alignment work?
Is that really the case in general outside the two recent high profile cases? I remember someone arguing recently instead that counterexamples were the true ‘creative’ aspect of maths and that proving theorems was more amenable to just churning hard work with obvious goals and thus maybe more amenable to AI.