AI Safety Researchers are only a small part of the AI Capabilities pool but are much bigger part of the AIS research pool. So unless the research lacks a plausible theory of change for safety, its impact of safety is probably more important than its impact on capabilities. See this entire comment thread from Paul Christiano’s 2021 AMA for interesting discussion on whether work is net positive.
You mention that Yudkowsky probably accelerated timelines[1] but he talked about AGI extensively back in 2008 when it was not taken seriously by most so it was a much larger share of the AGI discussion and thus could accelerate it more (also he had an unusually large audience of talented people who could work on AI capabilities). Now, it’s very much in the water supply and taken seriously by many frontier lab employees.
I would guess that he helped build the AI Safety field more than offsets the impact he had to timelines (although this is not very important to think about imo). It’s not clear what the purpose of extra time is if there are few people working on existential AI safety.
I think this argument might have been plausible before e.g. Anthropic started (and was definitely much more plausible before OpenAI started) but these two organizations alone make up a huge proportion of AI capabilities progress over the last decade and are very intimately intertwined with AI safety researchers.
Yeah, but I think my “AGI is now in the water supply” argument from above can explain this and still predict different future results. In my model the safety community ended up being intertwined with frontier labs and causing accelerations was a result of the rest of the ML community generally not being AGI-pilled. So the safety community had a somewhat small fraction of the overall capabilities work, but a considerably larger fraction of the capabilities work specifically by AGI-pilled people and therefore ended up helping create AGI labs. Beth Barnes’ comment from the aforementioned 2021 AMA predicted quite a bit of acceleration from 2021 safety researchers, and seems to have aged well.
Now that AGI is very much in the water supply for labs and many ML researchers, I expect much less acceleration from marginal people convinced by a friend to work on AI Safety. Admittedly, this is not super clear-cut to me and I’m open to be convinced otherwise but I’d need some more theoretical arguments or empirical data.
I think the term “AGI-pilled” fails to distinguish a fundamental difference, something like “is this person trying to think for themselves” vs “is this person following prestige gradients/social incentives”.
My concern is that people who are trying to think for themselves about what advanced AI will look like drive a lot of acceleration, by laying out a conceptual roadmap for what comes next (and then in many cases actually just working on that conceptual roadmap).
Whereas ML researchers who are now willing to talk about AGI are mostly doing so because it’s socially acceptable, and therefore will do a much worse job of innovative thinking about the future of AI.
(For example, most ML researchers—including those at DeepMind and OpenAI—who you might call “AGI-pilled” don’t take recursive self-improvement or superintelligence seriously, because that’s still socially weird. Whereas Anthropic people tend to take those things significantly more seriously, which helped them to catch up with/arguably overtake DeepMind and OpenAI starting from very far behind.)
It’s interesting that you bring up this comment from Beth btw because I think she’s been making a very similar mistake in underestimating how much METR is contributing to acceleration.
So the safety community had a somewhat small fraction of the overall capabilities work, but a considerably larger fraction of the capabilities work specifically by AGI-pilled people and therefore ended up helping create AGI labs.
I think it’s important to include “helping create AGI labs” as capabilities work, which would mean the safety community actually did a very significant fraction of the overall capabilities work from the last decade. Without accounting for that, all our estimates about the capabilities externalities of different work will be very off.
AI Safety Researchers are only a small part of the AI Capabilities pool but are much bigger part of the AIS research pool. So unless the research lacks a plausible theory of change for safety, its impact of safety is probably more important than its impact on capabilities. See this entire comment thread from Paul Christiano’s 2021 AMA for interesting discussion on whether work is net positive.
You mention that Yudkowsky probably accelerated timelines[1] but he talked about AGI extensively back in 2008 when it was not taken seriously by most so it was a much larger share of the AGI discussion and thus could accelerate it more (also he had an unusually large audience of talented people who could work on AI capabilities). Now, it’s very much in the water supply and taken seriously by many frontier lab employees.
I would guess that he helped build the AI Safety field more than offsets the impact he had to timelines (although this is not very important to think about imo). It’s not clear what the purpose of extra time is if there are few people working on existential AI safety.
I think this argument might have been plausible before e.g. Anthropic started (and was definitely much more plausible before OpenAI started) but these two organizations alone make up a huge proportion of AI capabilities progress over the last decade and are very intimately intertwined with AI safety researchers.
Yeah, but I think my “AGI is now in the water supply” argument from above can explain this and still predict different future results. In my model the safety community ended up being intertwined with frontier labs and causing accelerations was a result of the rest of the ML community generally not being AGI-pilled. So the safety community had a somewhat small fraction of the overall capabilities work, but a considerably larger fraction of the capabilities work specifically by AGI-pilled people and therefore ended up helping create AGI labs. Beth Barnes’ comment from the aforementioned 2021 AMA predicted quite a bit of acceleration from 2021 safety researchers, and seems to have aged well.
Now that AGI is very much in the water supply for labs and many ML researchers, I expect much less acceleration from marginal people convinced by a friend to work on AI Safety. Admittedly, this is not super clear-cut to me and I’m open to be convinced otherwise but I’d need some more theoretical arguments or empirical data.
I think the term “AGI-pilled” fails to distinguish a fundamental difference, something like “is this person trying to think for themselves” vs “is this person following prestige gradients/social incentives”.
My concern is that people who are trying to think for themselves about what advanced AI will look like drive a lot of acceleration, by laying out a conceptual roadmap for what comes next (and then in many cases actually just working on that conceptual roadmap).
Whereas ML researchers who are now willing to talk about AGI are mostly doing so because it’s socially acceptable, and therefore will do a much worse job of innovative thinking about the future of AI.
(For example, most ML researchers—including those at DeepMind and OpenAI—who you might call “AGI-pilled” don’t take recursive self-improvement or superintelligence seriously, because that’s still socially weird. Whereas Anthropic people tend to take those things significantly more seriously, which helped them to catch up with/arguably overtake DeepMind and OpenAI starting from very far behind.)
It’s interesting that you bring up this comment from Beth btw because I think she’s been making a very similar mistake in underestimating how much METR is contributing to acceleration.
I think it’s important to include “helping create AGI labs” as capabilities work, which would mean the safety community actually did a very significant fraction of the overall capabilities work from the last decade. Without accounting for that, all our estimates about the capabilities externalities of different work will be very off.