If we successfully slow down the speed of AI development, but do not perform any ‘Science of Alignment’ in the meantime, what will we gain? Unless you think an outright global ban on superintelligence is tenable (we will need a global-catastrophe-level warning shot to generate the political will for this IMO, and even then hard for this hold durably), then we have to use the time wisely to advance alignment science.
A reasonable position might be that all prosaic alignment attempts are doomed, so we need to allocate our time / resources to non-prosaic alignment (eg ARC), which has next to zero dual use at the moment. I think that’s a reasonable opinion, but not what I want to bet all our marbles on.
Thanks for this post! As a newcomer to AI safety, and seeing some of these debates, reading this helps clarify my thinking of what kind of work I should be doing.
I’m curious where you would place mech interp in this ontology? Two of its major applications seem to be:
Building a deeper understanding of what the models are doing and how they work
Designing more robust evaluations which can catch misalignment that blackbox methods would have missed.
So it seems much more on the ‘Science’ side, but I am hesitant to label all of it as Science? I guess I could imagine a point where mech interp becomes a dominant alignment eval method, such that improving mech interp methods just leads to greater confidence the models are aligned, unlocking continued acceleration of capbalities? Clearly we are not in this world yet, but could be there in 1 year. Still if these methods were good enough that this is actually earned (rather than false) confidence in the models alignment, it might not be entirely unjustified.
Anyway since mech interp seems more on the ‘Science’ side by default, perhaps its also a place where people who have an inclination towards ‘Engineering’ style hill climbing / metric improvement could direct their skills? Engineering performing Mech Interp methods being useful for Science