Thanks for this post! As a newcomer to AI safety, and seeing some of these debates, reading this helps clarify my thinking of what kind of work I should be doing.
I’m curious where you would place mech interp in this ontology? Two of its major applications seem to be:
Building a deeper understanding of what the models are doing and how they work
Designing more robust evaluations which can catch misalignment that blackbox methods would have missed.
So it seems much more on the ‘Science’ side, but I am hesitant to label all of it as Science? I guess I could imagine a point where mech interp becomes a dominant alignment eval method, such that improving mech interp methods just leads to greater confidence the models are aligned, unlocking continued acceleration of capbalities? Clearly we are not in this world yet, but could be there in 1 year. Still if these methods were good enough that this is actually earned (rather than false) confidence in the models alignment, it might not be entirely unjustified.
Anyway since mech interp seems more on the ‘Science’ side by default, perhaps its also a place where people who have an inclination towards ‘Engineering’ style hill climbing / metric improvement could direct their skills? Engineering performing Mech Interp methods being useful for Science
Thanks for this post! As a newcomer to AI safety, and seeing some of these debates, reading this helps clarify my thinking of what kind of work I should be doing.
I’m curious where you would place mech interp in this ontology? Two of its major applications seem to be:
Building a deeper understanding of what the models are doing and how they work
Designing more robust evaluations which can catch misalignment that blackbox methods would have missed.
So it seems much more on the ‘Science’ side, but I am hesitant to label all of it as Science? I guess I could imagine a point where mech interp becomes a dominant alignment eval method, such that improving mech interp methods just leads to greater confidence the models are aligned, unlocking continued acceleration of capbalities? Clearly we are not in this world yet, but could be there in 1 year. Still if these methods were good enough that this is actually earned (rather than false) confidence in the models alignment, it might not be entirely unjustified.
Anyway since mech interp seems more on the ‘Science’ side by default, perhaps its also a place where people who have an inclination towards ‘Engineering’ style hill climbing / metric improvement could direct their skills? Engineering performing Mech Interp methods being useful for Science