Quick history of some of my AI safety involvement and takes: - 2020 : read LW & sequences, yes this seems important, will get to it at some time (for the classic EA reasons of it’s the most important good), but was looking for my first job in France and took some software engineer position in a startup. I continue reading and upskilling on AIS during those 2 years. - 2022: quit my SWE job to go into AI safety. reasoning: yeah this is not a problem for the future, this is coming soon. I meet some people who have <5 year AGI timelines. I don’t know that it’s true, but it’s worth considering since it’s the very start of scaling and it could have been possible intelligence scaled much faster with scaling than it ended up in practice scaling. I work on AI Safety field building and infrastructure. There are then still <100 − 2023 : Still doing AIS field building, but looks like bottleneck is governance, I mostly keep pitching people around me to do comparatively more governance. (over time, this is one of the contributions that leads to the formation of CeSIA, the French Centre for AI Safety) - 2024 : Involvement with a wider range of ideas from the field has generally updated me more towards the OpenAI and Anthropic and Deepmind positions that prosaic alignment for human level AGI is technically feasible and rather tractable, and that this is indeed a very important input to how we most likely will progress on ASI alignment. I systematically distinguish between AGI risks and ASI risks and clarify the cruxes and assumptions for different threat models and am annoyed that MIRI threat models often skip to the end-game without consideration that certain trajectories towards that end-game falsify important assumptions (see Superintelligence of the Gaps for some elaboration there). I still believe governance is the bottleneck though, and continue contributing ideas that feed into https://www.lesswrong.com/posts/EexsebbYhbe2gXkPP/the-current-bottleneck-is-political-will-not-research − 2025 : start the year with a burnout after some intense governance work. Ideas wise not too much change with 2024, still doing field building and teaching AI safety theory occasionally, but also take a long break and learn more widely from other wisdom traditions (eg. buddhism, tpot and postrat stuff). - early 2026 : man does getting good US AI governance fast look intractable given the current administration, at this point it seems better to accelerate good safety work within the AI companies and in the surrounding ecosystem. I broadly wanna contribute to us surviving in the world where we don’t get much of a pause, where governments aren’t that competent and coordinated. I would work for any of the AGI companies on safety if found an adequate role. I think short timelines to human level AGI (eg. 2028) is plausible and preparing for if algorithmic improvement doesn’t asymptote too shallowly is important. Governance might not affect this in time. Being in the room where it happens seems likes the highest leverage way to increase the probability of good ai futures. A bunch of my theory of change is just helping the people in the room where it happens be more wise.
I would worry about the “be in the room” strategy. It seems like most people who justify their career decisions that way wind up getting captured by the groupthink of the org. They think they’ll be the ones who can resist it, but they won’t.
To the extent that an org does have a “groupthink” and its own theory&value system, then whether a given person in the room adheres to it seems to depend on : 1) the selection effect of who wanted to be there, in that particular org vs others 2) the discussions with other people in the org changing that person’s views 3) systematic pressures, eg. greedy/selfish parts of them optimizing to continue getting revenue
If I joined Anthropic and 1 year from now people thought I had surprisingly Anthropic-like views, I’d guess it’s mostly because of 1) and 2). 2) happens a lot but is broadly good. 3) is the one that’s mostly bad, from the outside/civilizational point of view, and the prior should be most people are susceptible to this, but this can be updated away from seeing particular life accomplishments. In my case, I have enough history of independence, selflessness and moral upstandingness that I don’t think 3) will influence me substantially, but I don’t recommend this path to those without that history.
I would really suggest reflecting on it if you haven’t already. “I am special and can resist the groupthink” is often a false belief.
As for 2), I don’t think it is necessarily good? If most people there have 1) and 3) influencing them, that will filter the kinds of opinions they have which then get transmitted to you in conversation, and now even if you are stalwartly resisting the direct pull of 3), it’ll still reach through others to pull at you.
Not to mention, workers at frontier labs seem to be doing a fourth thing, delegating increasingly large amounts of trust and thinking to their AIs, in ways which might be troublesome; you would be signing up for this. It can happen indirectly, even if you don’t do it yourself, because others will launder AIs’ beliefs as their own.
Separately, there’s the issue that leadership at the labs simply have their own beliefs about various important issues, and don’t care for the opinions of the rank and file. Anthropic defanged the RSPs it arguably drew in many researchers with; OpenAI let two alignment teams wither; DeepMind sold out to the military over its employees’ objections. The explicit goal of these labs is RSI, and the first workers they want to unemploy are their own, especially their juniors. The remaining employees at late stage AI labs will mostly be a core of leadership and senior researchers whose research taste is still required. What sorts of impactful decisions would you be able to meaningfully influence in the window between signing on and obsolescence, if any?
Thanks for your comments. I don’t expect doing an analysis of my situation in particular is best use of our time but I do think these are helpful questions to consider for people in my situation or similar.
Re your last paragraph, I’d happily bet that Anthropic has not reduced their workforce 2 years from now. Yes relative employee disempowerment is an important factor I care about, but it is precisely in worlds where alignment is not that good that having humans in the loop is important (in an obvious seen-by-leadership way). It is only reasonable to automate everything with very very high trust in both the competence and alignment of AI systems, and Anthropic as a company is not that unreasonable, they definitely do find and classify many Claude behaviors as undesirable, and will continue doing so.
There’s a usual back and forth about how much to distrust leadership of AGI companies which is hard to ground in material fact. Some people take the lack of safety actions now to mean lack of care for when it will matter, but conversely the fact that it never mattered yet is a good reason for them not to have cared for these inconsequential things. The explanation for defanging the RSP is a good one, I don’t think people should tie themselves to masts and go blind into the unknown unknowns of AGI development. They should build capacity to remain aware, capacity to pause, have institutions that can do independent audit and have real power to stop them, but not fixed RSP-like stuff.
Finally, still on last paragraph, the “window between signing on and obsolescence” is very dependant on people’s rates of growth, but also where they can work immediately. I am generally glad that Joe Carlsmith joined Anthropic to help with the Claude Constitution, I think he immediately is having very significant impact. There is much object level work to make the chances of better futures to be done. Even if one later gets automated, having made alignment that much better before full automation could be a significant difference.
Quick history of some of my AI safety involvement and takes:
- 2020 : read LW & sequences, yes this seems important, will get to it at some time (for the classic EA reasons of it’s the most important good), but was looking for my first job in France and took some software engineer position in a startup. I continue reading and upskilling on AIS during those 2 years.
- 2022: quit my SWE job to go into AI safety. reasoning: yeah this is not a problem for the future, this is coming soon. I meet some people who have <5 year AGI timelines. I don’t know that it’s true, but it’s worth considering since it’s the very start of scaling and it could have been possible intelligence scaled much faster with scaling than it ended up in practice scaling. I work on AI Safety field building and infrastructure. There are then still <100
− 2023 : Still doing AIS field building, but looks like bottleneck is governance, I mostly keep pitching people around me to do comparatively more governance. (over time, this is one of the contributions that leads to the formation of CeSIA, the French Centre for AI Safety)
- 2024 : Involvement with a wider range of ideas from the field has generally updated me more towards the OpenAI and Anthropic and Deepmind positions that prosaic alignment for human level AGI is technically feasible and rather tractable, and that this is indeed a very important input to how we most likely will progress on ASI alignment. I systematically distinguish between AGI risks and ASI risks and clarify the cruxes and assumptions for different threat models and am annoyed that MIRI threat models often skip to the end-game without consideration that certain trajectories towards that end-game falsify important assumptions (see Superintelligence of the Gaps for some elaboration there). I still believe governance is the bottleneck though, and continue contributing ideas that feed into https://www.lesswrong.com/posts/EexsebbYhbe2gXkPP/the-current-bottleneck-is-political-will-not-research
− 2025 : start the year with a burnout after some intense governance work. Ideas wise not too much change with 2024, still doing field building and teaching AI safety theory occasionally, but also take a long break and learn more widely from other wisdom traditions (eg. buddhism, tpot and postrat stuff).
- early 2026 : man does getting good US AI governance fast look intractable given the current administration, at this point it seems better to accelerate good safety work within the AI companies and in the surrounding ecosystem. I broadly wanna contribute to us surviving in the world where we don’t get much of a pause, where governments aren’t that competent and coordinated. I would work for any of the AGI companies on safety if found an adequate role. I think short timelines to human level AGI (eg. 2028) is plausible and preparing for if algorithmic improvement doesn’t asymptote too shallowly is important. Governance might not affect this in time. Being in the room where it happens seems likes the highest leverage way to increase the probability of good ai futures. A bunch of my theory of change is just helping the people in the room where it happens be more wise.
I would worry about the “be in the room” strategy. It seems like most people who justify their career decisions that way wind up getting captured by the groupthink of the org. They think they’ll be the ones who can resist it, but they won’t.
To the extent that an org does have a “groupthink” and its own theory&value system, then whether a given person in the room adheres to it seems to depend on :
1) the selection effect of who wanted to be there, in that particular org vs others
2) the discussions with other people in the org changing that person’s views
3) systematic pressures, eg. greedy/selfish parts of them optimizing to continue getting revenue
If I joined Anthropic and 1 year from now people thought I had surprisingly Anthropic-like views, I’d guess it’s mostly because of 1) and 2). 2) happens a lot but is broadly good. 3) is the one that’s mostly bad, from the outside/civilizational point of view, and the prior should be most people are susceptible to this, but this can be updated away from seeing particular life accomplishments. In my case, I have enough history of independence, selflessness and moral upstandingness that I don’t think 3) will influence me substantially, but I don’t recommend this path to those without that history.
Here’s an example of 3) happening to someone, and they noticed it: https://forum.effectivealtruism.org/posts/rHyAmvXiqrC9iAR9T/jay-bailey-s-shortform
I would really suggest reflecting on it if you haven’t already. “I am special and can resist the groupthink” is often a false belief.
As for 2), I don’t think it is necessarily good? If most people there have 1) and 3) influencing them, that will filter the kinds of opinions they have which then get transmitted to you in conversation, and now even if you are stalwartly resisting the direct pull of 3), it’ll still reach through others to pull at you.
Not to mention, workers at frontier labs seem to be doing a fourth thing, delegating increasingly large amounts of trust and thinking to their AIs, in ways which might be troublesome; you would be signing up for this. It can happen indirectly, even if you don’t do it yourself, because others will launder AIs’ beliefs as their own.
Separately, there’s the issue that leadership at the labs simply have their own beliefs about various important issues, and don’t care for the opinions of the rank and file. Anthropic defanged the RSPs it arguably drew in many researchers with; OpenAI let two alignment teams wither; DeepMind sold out to the military over its employees’ objections. The explicit goal of these labs is RSI, and the first workers they want to unemploy are their own, especially their juniors. The remaining employees at late stage AI labs will mostly be a core of leadership and senior researchers whose research taste is still required. What sorts of impactful decisions would you be able to meaningfully influence in the window between signing on and obsolescence, if any?
Thanks for your comments. I don’t expect doing an analysis of my situation in particular is best use of our time but I do think these are helpful questions to consider for people in my situation or similar.
Re your last paragraph, I’d happily bet that Anthropic has not reduced their workforce 2 years from now. Yes relative employee disempowerment is an important factor I care about, but it is precisely in worlds where alignment is not that good that having humans in the loop is important (in an obvious seen-by-leadership way). It is only reasonable to automate everything with very very high trust in both the competence and alignment of AI systems, and Anthropic as a company is not that unreasonable, they definitely do find and classify many Claude behaviors as undesirable, and will continue doing so.
There’s a usual back and forth about how much to distrust leadership of AGI companies which is hard to ground in material fact. Some people take the lack of safety actions now to mean lack of care for when it will matter, but conversely the fact that it never mattered yet is a good reason for them not to have cared for these inconsequential things. The explanation for defanging the RSP is a good one, I don’t think people should tie themselves to masts and go blind into the unknown unknowns of AGI development. They should build capacity to remain aware, capacity to pause, have institutions that can do independent audit and have real power to stop them, but not fixed RSP-like stuff.
Finally, still on last paragraph, the “window between signing on and obsolescence” is very dependant on people’s rates of growth, but also where they can work immediately. I am generally glad that Joe Carlsmith joined Anthropic to help with the Claude Constitution, I think he immediately is having very significant impact. There is much object level work to make the chances of better futures to be done. Even if one later gets automated, having made alignment that much better before full automation could be a significant difference.