> the bar at MATS has raised every program for 4 years now
What?! Something terrible must be going on in your mechanisms for evaluating people (which to be clear, isn’t surprising, indeed, you are the central target of the optimization that is happening here, but like, to me it illustrates the risks here quite cleanly).
It is very very obvious to me that median MATS participant quality has gone down continuously for the last few cohorts. I thought this was somewhat clear to y’all and you thought it was worth the tradeoff of having bigger cohorts, but you thinking it has “gone up continuously” shows a huge disconnect.
Like, these days at the end of a MATS program half of the people couldn’t really tell you why AI might be an existential risk at all. Their eyes glaze over when you try to talk about AI strategy. IDK, maybe these people are better ML researchers, but obviously they are worse contributors to the field than the people in the early cohorts.
One thing to note about the first two MATS cohorts is that they occurred before the FTX crash (and pre-ChatGPT!). [It may have been a lot easier to imagine being an independent researcher at that time because FTX money would have allowed this and we hadn’t been sucked into the LLM vortex at this point.]
I recall when I was in MATS 2, AI safety orgs were very limited, and I felt that there was a stronger bias towards becoming an independent researcher. Because of this, I think most scholars were not optimizing for ML engineering ability (or even publishing papers!), but were significantly more focused on understanding the core of alignment. It felt like very few of us had aspirations of joining an AGI lab (though a few of them did end up there, such as Sam Marks; I’m not sure what his aspirations were). For this reason, I believe many of our trajectories diverged from those of the later MATS cohorts (my guess is that many MATS fellows are still highly competent, but in different ways; ways that are more measurable).
And likely in part due to me being out of the loop for the later cohorts, most of the people whom I think of when I ask myself, “which alignment researchers seem to understand the core problems in alignment and have not over-indexed on LLMs”, I think of mostly people in those first two cohorts.
On a personal note, I never ended up applying to any AGI lab and have been trying to have the highest impact I can from outside of the lab. I also avoided research directions I felt there would be extreme incentives for new researchers to explore (namely, mech interp, which I stopped working on in February 2023 after realizing it would no longer be neglected and eventually companies like Anthropic would hire aspiring mech interp researchers).
Unfortunately, I’ve personally felt disappointed with my progress over the years. Though I think it’s obviously harder to have an impact if you are constantly exploring new directions like I have been (had I stuck with mech interp, I might be leading a team in that research direction at this point).
On the other hand, there’s another concern I’ve been wary of in the context of AI safety startups (which is what I’m currently exploring) and research in general: following the short-term success gradient. In startups, you can start with a noble vision and then become increasingly pressured away from the initial vision simply because you are pursuing the customer gradient and “building what people want.” If your goal is large-scale (venture) success, then it only makes sense. You need customers and traction for your Series A after all. Even in research, there’s only so much fucking around you can do until people want something legible from you.
Anyway, despite not having started a successful AI safety startup at this point, at least part of it has come from taking my time in finding which mountain I want to climb and avoiding locking myself into a path that doesn’t end up making progress on the core technical problems in alignment.
I mainly didn’t do it because I thought Ryan wrote a useful post, and I didn’t want to derail (what I felt was supposed to be) the conversation further. But maybe you’re right, and it would be fine.
On the other hand, there’s another concern I’ve been wary of in the context of AI safety startups (which is what I’m currently exploring) and research in general: following the short-term success gradient. In startups, you can start with a noble vision and then become increasingly pressured away from the initial vision simply because you are pursuing the customer gradient and “building what people want.” If your goal is large-scale (venture) success, then it only makes sense. You need customers and traction for your Series A after all. Even in research, there’s only so much fucking around you can do until people want something legible from you.
This is my biggest concern with d/acc style techno-optimism, it seems to assume that genuinely defensive technologies can compete economically with offensive ones (all it takes is the right founders, seed funding etc.).
Whereas my impression is that any kind of ethical/ideological commitment immediately puts a startup at a massive structural disadvantage against those who chose simply to give the market what it wants (acceleration).
This is my biggest concern with d/acc style techno-optimism, it seems to assume that genuinely defensive technologies can compete economically with offensive ones (all it takes is the right founders, seed funding etc.).
Does it assume that? There are many ways for governments to adjust for d/acc tech being less innately appealing by intervening on market incentives, for example, through subsidies, tax credits, benefits for those who adopt these products, etc. Doing that may for various reasons be more tractable than command-and-control regulation. But either way, doing either (incentivising or mandating) seems easier once the tech actually exists and is somewhat proven, so you may want founders to start d/acc projects even if you think they would not become profitable in the free market and even if you want to mandate that tech eventually.
(That is not to say that there is a lot of useful d/acc tech that awaits being created, and that if implemented would make a major difference. I just think that, if there is, then that tech being able to compete economically isn’t necessarily a huge problem.)
You are right that I am being a bit reductive. Maybe it would be better to say it assumes some kind of ideal combination of innovation, markets and technocratic governance would be enough to prevent catastrophe?
And to be clear I do think its much better for people to be working on defensive technologies, than not to. And its not impossible that the right combination of defensive entrepreneurs and technocratic government incentives could genuinely solve a problem.
But I think this kind of faith in business as usual but a bit better can lead to a kind of complacency where you conflate working on good things with actually making a difference.
I also just want to point out that there should be a base rate here that’s higher context in the beginning since before MATS and similar there weren’t really that many AI Safety training programs.
So the intiial people that you get will automatically be higher context because the sample is taken from people who have already worked on it/learnt about it for a while. This should go down over time due to the higher context individuals being taken in?
(I don’t know how large this effect would be but I would just want to point it out.)
Habryka responding to Ryan Kidd:
One thing to note about the first two MATS cohorts is that they occurred before the FTX crash (and pre-ChatGPT!). [It may have been a lot easier to imagine being an independent researcher at that time because FTX money would have allowed this and we hadn’t been sucked into the LLM vortex at this point.]
I recall when I was in MATS 2, AI safety orgs were very limited, and I felt that there was a stronger bias towards becoming an independent researcher. Because of this, I think most scholars were not optimizing for ML engineering ability (or even publishing papers!), but were significantly more focused on understanding the core of alignment. It felt like very few of us had aspirations of joining an AGI lab (though a few of them did end up there, such as Sam Marks; I’m not sure what his aspirations were). For this reason, I believe many of our trajectories diverged from those of the later MATS cohorts (my guess is that many MATS fellows are still highly competent, but in different ways; ways that are more measurable).
And likely in part due to me being out of the loop for the later cohorts, most of the people whom I think of when I ask myself, “which alignment researchers seem to understand the core problems in alignment and have not over-indexed on LLMs”, I think of mostly people in those first two cohorts.
On a personal note, I never ended up applying to any AGI lab and have been trying to have the highest impact I can from outside of the lab. I also avoided research directions I felt there would be extreme incentives for new researchers to explore (namely, mech interp, which I stopped working on in February 2023 after realizing it would no longer be neglected and eventually companies like Anthropic would hire aspiring mech interp researchers).
Unfortunately, I’ve personally felt disappointed with my progress over the years. Though I think it’s obviously harder to have an impact if you are constantly exploring new directions like I have been (had I stuck with mech interp, I might be leading a team in that research direction at this point).
On the other hand, there’s another concern I’ve been wary of in the context of AI safety startups (which is what I’m currently exploring) and research in general: following the short-term success gradient. In startups, you can start with a noble vision and then become increasingly pressured away from the initial vision simply because you are pursuing the customer gradient and “building what people want.” If your goal is large-scale (venture) success, then it only makes sense. You need customers and traction for your Series A after all. Even in research, there’s only so much fucking around you can do until people want something legible from you.
Anyway, despite not having started a successful AI safety startup at this point, at least part of it has come from taking my time in finding which mountain I want to climb and avoiding locking myself into a path that doesn’t end up making progress on the core technical problems in alignment.
I think this would be very useful to have posted in the original thread.
I mainly didn’t do it because I thought Ryan wrote a useful post, and I didn’t want to derail (what I felt was supposed to be) the conversation further. But maybe you’re right, and it would be fine.
This is my biggest concern with d/acc style techno-optimism, it seems to assume that genuinely defensive technologies can compete economically with offensive ones (all it takes is the right founders, seed funding etc.).
Whereas my impression is that any kind of ethical/ideological commitment immediately puts a startup at a massive structural disadvantage against those who chose simply to give the market what it wants (acceleration).
Does it assume that? There are many ways for governments to adjust for d/acc tech being less innately appealing by intervening on market incentives, for example, through subsidies, tax credits, benefits for those who adopt these products, etc. Doing that may for various reasons be more tractable than command-and-control regulation. But either way, doing either (incentivising or mandating) seems easier once the tech actually exists and is somewhat proven, so you may want founders to start d/acc projects even if you think they would not become profitable in the free market and even if you want to mandate that tech eventually.
(That is not to say that there is a lot of useful d/acc tech that awaits being created, and that if implemented would make a major difference. I just think that, if there is, then that tech being able to compete economically isn’t necessarily a huge problem.)
You are right that I am being a bit reductive. Maybe it would be better to say it assumes some kind of ideal combination of innovation, markets and technocratic governance would be enough to prevent catastrophe?
And to be clear I do think its much better for people to be working on defensive technologies, than not to. And its not impossible that the right combination of defensive entrepreneurs and technocratic government incentives could genuinely solve a problem.
But I think this kind of faith in business as usual but a bit better can lead to a kind of complacency where you conflate working on good things with actually making a difference.
I also just want to point out that there should be a base rate here that’s higher context in the beginning since before MATS and similar there weren’t really that many AI Safety training programs.
So the intiial people that you get will automatically be higher context because the sample is taken from people who have already worked on it/learnt about it for a while. This should go down over time due to the higher context individuals being taken in?
(I don’t know how large this effect would be but I would just want to point it out.)