Thanks for this post! As a newcomer to AI safety, and seeing some of these debates, reading this helps clarify my thinking of what kind of work I should be doing.
I’m curious where you would place mech interp in this ontology? Two of its major applications seem to be:
Building a deeper understanding of what the models are doing and how they work
Designing more robust evaluations which can catch misalignment that blackbox methods would have missed.
So it seems much more on the ‘Science’ side, but I am hesitant to label all of it as Science? I guess I could imagine a point where mech interp becomes a dominant alignment eval method, such that improving mech interp methods just leads to greater confidence the models are aligned, unlocking continued acceleration of capbalities? Clearly we are not in this world yet, but could be there in 1 year. Still if these methods were good enough that this is actually earned (rather than false) confidence in the models alignment, it might not be entirely unjustified.
Anyway since mech interp seems more on the ‘Science’ side by default, perhaps its also a place where people who have an inclination towards ‘Engineering’ style hill climbing / metric improvement could direct their skills? Engineering performing Mech Interp methods being useful for Science
As someone who has very recently pivoted (still pivoting really) to AI safety I pretty strongly disagree with this take.
It seems the real bottleneck right now is organizational capacity to absorb talented people and get them working on important problems in a semi-organized way. Its clear there is so much research to be done and a lot of low hanging fruit to be plucked. We absolutely need more people doing this work. The funding is there and growing fast, expect even more after frontier labs IPO. An announcement like Project Tailwind is a bat signal that the funding is there, the constraint is new orgs to use it effectively. Existing labs can only take on so many new people at a time while staying focused.
But the good news is as that as these organizations and the AI safety ecosystem grows, they will start to be able to absorb more people per year. Like this year we have X number of mentors for AI safety research, they take say 2X mentees. Say 50% of the mentees are really good and become independent researchers, then 6 months from now we will have 2X mentors, so 4X mentees can enter the field.
If you are capable already of being a quite independent researcher, perhaps because you already have a decent amount of research experience in another field, its super valuable for you to pivot because you will not eat up much mentorship time and will still be productive. Even as newcomer with probably not the greatest research taste in this area yet I see the opportunity for so many projects that are interesting, many that I could do mostly independently (writing up my first one now). This would have no drain on organization capacity and so is definitely a net positive.
And I don’t want to be rude, but when it comes to research, there are levels to this shit.
Someone just out of undergrad, even if they are very bright and talented needs a decent amount of handholding to be directed to productive research activities. In my previous field, after a PhD I would say ~33% of people become truly independent researchers who can generate their own good ideas, ~33% have the technical skills but don’t (yet) have the vision to craft their own research agenda, and ~33% never really had the sauce so to speak and needed a ton of hand holding the whole time. (This includes many people coming from top universities, research is just quite different than being a good student). Unsure how this shakes out in the powerful AI age we now in, but I think it may raise the floor but also increase the differentiation in output even more.
Even for senior, established researchers (ie professors), there is a pretty decent gap in productivity. In my previous field of ML+physics probably 50% of the good ideas were coming from the top 5% of the community. Its not that everyone else was doing useless things but they were often not pushing the boundaries in the same way, and more so filling gaps in with incremental work (still valuable! but lower impact).
So getting some exceptionally talented person to switch into AI safety can have huge value. And unfortunately that possibility makes it worth advising as many people as possible to switch into AI safety, to have a chance to get a super talented person doing something impactful. Even if that means some more junior people struggle to find a good home for a while.