Please don’t overhype or oversell AI Safety as a professional field, and in particular how easy it is to work in AI Safety. You ~have to be either quite smart or hard working/motivated to get a stable career, sometimes even both. AI Safety is generally attention and mentoring and management bottlenecked1, there is much less pipeline for “normal people” to make it, contrary to most standardized fields. When teaching and mentoring people, not all will be able to transition fast or upskill fast or get strong positions, some would benefit more from keeping their careers or donating or upskilling over longer timespans.
More smart wise well motivated people caring and working on good AI futures will keep mattering until the end, and the total number of people who can meaningfully contribute is probably in the millions, so let’s not stop field building
What does it mean to teach AI Safety?
There are many facts about AI, and theories about how it will go under what conditions. It means getting to learn and understand these well. Teaching AI Safety well should not over-determine someone to come out of the pipeline with a MIRI-view, nor an Anthropic-view, nor an EA-view, but should allow them to pass the basic Ideological Turing Test for all of these.
A central example of AI Safety fundamental knowledge is the AI Safety Atlas
On teaching AI Safety
AI safety changes all the time. Update your slides ~everytime you teach something.
If you’re new to teaching, find existing resources to build on. Be careful of AI slop. Better much less content than irrelevant or unsupported. Test yourself on your own knowledge, imagining questions one might ask and answering them. Traditional “how to teach/mentor/tutor” advice/resources are helpful.
Understand peoples’ motivations asap
For a given cohort, some will be motivated by “Do Good” EA or EA-adjacent reasons. Some by trends. Some by interest. Broadly all are valid in that there’s a place for a skilled someone with any of these motivations (but certain positions are only fit for certain motivations).
For people who want to have a lot of good impact, contributing towards good futures, context matters a lot
I recommend they seriously upskill in understand AI progress and trendlines. Read most EpochAI work and peter wildeford stuff, and AI threat models and scenarios, in particular thinking through AI 2027, contrasting with all other threat models listed in Deepmind’s lit review.
I recommend reading the AI Safety Atlas and understanding everything in it, as basic context for what the gameboard looks like, what areas exist in AIS.
And they’d benefit from building a pipeline for new info. Being involved in AIS community, having friends and colleagues, following the right people on twitter. Checking LessWrong from time to time.
For people who want to have a career or do fun research, they can just focus on that. They should find a context where someone else is thinking about why stuff matters and happy to have collaborators/employees executing on stuff. They should get good at whatever particular job or tasks they wanna do.
Most people interested in AI safety nowadays don’t know the old stuff, haven’t read the sequences. It’d be too much to ask them to read it all, but I do broadly recommend reading the best of LessWrong from most years. Being good at thinking and life is instrumental for being good at AI Safety. Different people need different advice so I can’t really put a best-of here (I really like Please don’t throw your mind away). There is wisdom in many different communities. If you don’t have the “rationalist community” wisdom then invest in it, otherwise keep branching out. Go to burning man or a local burn, understand the basics of meditation and different embodiment practices (eg. yoga), how therapy and self actualization work, read this blog and others. Keep asking questions about how to live the good life and have good community, actively pursue your best understanding of how to do that at any given moment.
Quick history of some of my AI safety involvement and takes: - 2020 : read LW & sequences, yes this seems important, will get to it at some time (for the classic EA reasons of it’s the most important good), but was looking for my first job in France and took some software engineer position in a startup. I continue reading and upskilling on AIS during those 2 years. - 2022: quit my SWE job to go into AI safety. reasoning: yeah this is not a problem for the future, this is coming soon. I meet some people who have <5 year AGI timelines. I don’t know that it’s true, but it’s worth considering since it’s the very start of scaling and it could have been possible intelligence scaled much faster with scaling than it ended up in practice scaling. I work on AI Safety field building and infrastructure. There are then still <100 − 2023 : Still doing AIS field building, but looks like bottleneck is governance, I mostly keep pitching people around me to do comparatively more governance. (over time, this is one of the contributions that leads to the formation of CeSIA, the French Centre for AI Safety) - 2024 : Involvement with a wider range of ideas from the field has generally updated me more towards the OpenAI and Anthropic and Deepmind positions that prosaic alignment for human level AGI is technically feasible and rather tractable, and that this is indeed a very important input to how we most likely will progress on ASI alignment. I systematically distinguish between AGI risks and ASI risks and clarify the cruxes and assumptions for different threat models and am annoyed that MIRI threat models often skip to the end-game without consideration that certain trajectories towards that end-game falsify important assumptions (see Superintelligence of the Gaps for some elaboration there). I still believe governance is the bottleneck though, and continue contributing ideas that feed into https://www.lesswrong.com/posts/EexsebbYhbe2gXkPP/the-current-bottleneck-is-political-will-not-research − 2025 : start the year with a burnout after some intense governance work. Ideas wise not too much change with 2024, still doing field building and teaching AI safety theory occasionally, but also take a long break and learn more widely from other wisdom traditions (eg. buddhism, tpot and postrat stuff). - early 2026 : man does getting good US AI governance fast look intractable given the current administration, at this point it seems better to accelerate good safety work within the AI companies and in the surrounding ecosystem. I broadly wanna contribute to us surviving in the world where we don’t get much of a pause, where governments aren’t that competent and coordinated. I would work for any of the AGI companies on safety if found an adequate role. I think short timelines to human level AGI (eg. 2028) is plausible and preparing for if algorithmic improvement doesn’t asymptote too shallowly is important. Governance might not affect this in time. Being in the room where it happens seems likes the highest leverage way to increase the probability of good ai futures. A bunch of my theory of change is just helping the people in the room where it happens be more wise.
I would worry about the “be in the room” strategy. It seems like most people who justify their career decisions that way wind up getting captured by the groupthink of the org. They think they’ll be the ones who can resist it, but they won’t.
To the extent that an org does have a “groupthink” and its own theory&value system, then whether a given person in the room adheres to it seems to depend on : 1) the selection effect of who wanted to be there, in that particular org vs others 2) the discussions with other people in the org changing that person’s views 3) systematic pressures, eg. greedy/selfish parts of them optimizing to continue getting revenue
If I joined Anthropic and 1 year from now people thought I had surprisingly Anthropic-like views, I’d guess it’s mostly because of 1) and 2). 2) happens a lot but is broadly good. 3) is the one that’s mostly bad, from the outside/civilizational point of view, and the prior should be most people are susceptible to this, but this can be updated away from seeing particular life accomplishments. In my case, I have enough history of independence, selflessness and moral upstandingness that I don’t think 3) will influence me substantially, but I don’t recommend this path to those without that history.
I would really suggest reflecting on it if you haven’t already. “I am special and can resist the groupthink” is often a false belief.
As for 2), I don’t think it is necessarily good? If most people there have 1) and 3) influencing them, that will filter the kinds of opinions they have which then get transmitted to you in conversation, and now even if you are stalwartly resisting the direct pull of 3), it’ll still reach through others to pull at you.
Not to mention, workers at frontier labs seem to be doing a fourth thing, delegating increasingly large amounts of trust and thinking to their AIs, in ways which might be troublesome; you would be signing up for this. It can happen indirectly, even if you don’t do it yourself, because others will launder AIs’ beliefs as their own.
Separately, there’s the issue that leadership at the labs simply have their own beliefs about various important issues, and don’t care for the opinions of the rank and file. Anthropic defanged the RSPs it arguably drew in many researchers with; OpenAI let two alignment teams wither; DeepMind sold out to the military over its employees’ objections. The explicit goal of these labs is RSI, and the first workers they want to unemploy are their own, especially their juniors. The remaining employees at late stage AI labs will mostly be a core of leadership and senior researchers whose research taste is still required. What sorts of impactful decisions would you be able to meaningfully influence in the window between signing on and obsolescence, if any?
Thanks for your comments. I don’t expect doing an analysis of my situation in particular is best use of our time but I do think these are helpful questions to consider for people in my situation or similar.
Re your last paragraph, I’d happily bet that Anthropic has not reduced their workforce 2 years from now. Yes relative employee disempowerment is an important factor I care about, but it is precisely in worlds where alignment is not that good that having humans in the loop is important (in an obvious seen-by-leadership way). It is only reasonable to automate everything with very very high trust in both the competence and alignment of AI systems, and Anthropic as a company is not that unreasonable, they definitely do find and classify many Claude behaviors as undesirable, and will continue doing so.
There’s a usual back and forth about how much to distrust leadership of AGI companies which is hard to ground in material fact. Some people take the lack of safety actions now to mean lack of care for when it will matter, but conversely the fact that it never mattered yet is a good reason for them not to have cared for these inconsequential things. The explanation for defanging the RSP is a good one, I don’t think people should tie themselves to masts and go blind into the unknown unknowns of AGI development. They should build capacity to remain aware, capacity to pause, have institutions that can do independent audit and have real power to stop them, but not fixed RSP-like stuff.
Finally, still on last paragraph, the “window between signing on and obsolescence” is very dependant on people’s rates of growth, but also where they can work immediately. I am generally glad that Joe Carlsmith joined Anthropic to help with the Claude Constitution, I think he immediately is having very significant impact. There is much object level work to make the chances of better futures to be done. Even if one later gets automated, having made alignment that much better before full automation could be a significant difference.
Potential crux with MIRI-like rationalists: is it in fact the case that our current world, with Anthropic’s influence, is worse than one without Anthropic?
On rationalist views, the world was going get worse and worse anyway (as capabilities advance and we get closer to doom). Anthropic accelerated and continues to accelerate capabilities progress. But how much did they comparatively accelerate alignment and saner AI policy?
In a world with eg. just OpenAI and GDM at the frontier, if/when OpenAI pulls ahead at RSI (as currently seems to be the case): - would there even have been the current level of integration with UK AISI, current level of model organisms and safety evals? - would the AI safety space have the expected hundreds of billions of funding, to ambitiously scale its work, including AI policy work? - would there be have been an AGI company with *some* Operational Adequacy, to proactively do things like Glasswing and biorisk-mitigation? (imo evaluating on the specified criteria, it’s clear Anthropic is ahead of OpenAI on most dimensions, and can continue improving on these. One can be upset they aren’t technically held by their initial RSP, and yet in practice they seem to be better than OpenAI at it.).
If you wonder why I compare to OpenAI rather than nothing, it’s because I don’t think “nothing” is the counterfactual of Anthropic not existing. When evaluating the wisdom of Anthropic doing what it did, it’s necessary to evaluate against more likely counterfactuals. Possibly many rationalists do take these counterfactuals carefully into account, but the arguments often raised often skip that part. “Anthropic accelerated capabilities” is not a sufficient argument to expect Anthropic’s influence on the world to have been net negative.
There are definitely Fabricated Option Worlds which seem much better than the one we got, and on the margin one can hope Anthropic to have done better work or not accelerated capabilities as much, but it’ seems difficult from the outside to be sure they did the wrong tradeoffs.
My own epistemic status here on whether Anthropic has been net good is “Uncertain”
Generally my AIS thoughts/threads are mirrored between twitter and LessWrong shortform, while my LW posts are mirrored to Substack and linked to from twitter. Interesting conversation may happen at all these places.
I would like if LessWrong provided an optional newsletter like the EA forum digest, for people who want the chance to catch non-curated posts without having to open LessWrong and sift through it directly every few days.
Here’s what the EA Forum digest looks like : a list of titles + author + time to read.
I don’t know exactly how much manual curation goes into it and I’m not asking for that. I’d find a simple karma threshold and this format valuable.
I am also writing up this quick take notably because I’ve had discussions with other people who’d like this, and because recently costs of development and maintenance of software like this have gone down.
I don’t find the existing RSS feed a preferable alternative. - I have never setup an RSS feed reader or similar process, I don’t think I want to and guess most LW readers are similar. - It seems it would on top of that would take extra work to get the format I want out of it
If AI alignment & safety goes well, we’ll have *more* important work in the next 5 to 10 years than not, so sufficiently optimist people (aka less than 50% p(doom)) should invest in good health and productivity with rather “long” timelines, eg. 10y+, *because* good futures imply not having lost control, and instead every decision we make being amplified in its potential impact.
This does require some assumptions, so unpacking 1) Aligning to a static set of values is misguided, you’ll have to have an alive process that keeps humans in the loops for alignment to humans to mean anything. (And to protect from disempowerment and various other threat models). 2) It would probably not be democratic and aligned for ASI to quickly take control and run the world. Institutions should voluntarily give over control only with enough confidence and knowledge that this is good, which will require a bunch of time and buy-in, thus elongating the time where humans are in fact taking important decisions. 3) Yet during this time, many humans’ decisions will have outsized influence, thanks to being empowered by AI. It would be bad if some of the thoughtful and engaged people are too burnt out to continue making the world better during the longer important transition period. Continue investing in becoming wise and competent and with a good social graph, as this will keep paying off for a long time.
To be clear, I’m saying the above would be true even if we had aligned ASI in two years, though with some of my assumptions of what “aligned” means here. I *can* imagine worlds where things move faster (for example) but I don’t think this is utopia or good from the point of view of our current political systems, and thus not what AI Safety optimists should be orienting for.
(If you’re an optimist and your plan is for AI to takeover the world asap, consider that other people don’t want that to happen, and you don’t have a moral superiority to do that from, and it is normal that society would resist your plans)
My view of “Good AI Futures” is more something like: an increasing number of smart pro-social people who Get It get involved in making the transition to post-ASI society go well. This involves much dissent, infighting, contradictory moves, and this is broadly good (any one group should not assume they’re right, but accept being one part of the equilibrium forming confrontation). It hopefully also involves good communication norms, asymmetric weapons, collaboration and bridge-building. I don’t think any one group should just straight up win, which is why I think there is useful good work to be done for a long time, navigating different groups’ actually conflicting interests.
We’re in it for the long haul. We’re currently at the edge of the precipice and might yet extinguish ~all value forever, so urgency is warranted, but let it be the kind of urgency that doesn’t sacrifice our ability to keep contributing to steer towards good AI futures, please don’t throw your mind away.
Many things can be done more effectively under fast ~ASI guidance through headset w/ video.
The show Pluribus gives some useful intuitions at how fast & effective superhumanly coordinated human work can be. It’s more bullish in some ways (a human won’t acquire technical know that requires practice as fast) and bearish in others (they have the same total amount of compute, while we’d have much more, and be innovating on methods of work much faster).
A normal human 8h work day has huge amounts of waste whether not doing much, or not useful things.
You can increase the efficiency of how much they work (not blocked on coordination problems), how well they work (continuous coaching so ~everyone reaches what is current top 1%, tho domain dependant), how useful what they work on is (better management, priorisation).
Of these factors, I would guess that better management/coordination is the main one.
If you isolated just one human within a factory, the ASI might make them somewhat more efficient, but they’ll be bottlenecked by machinery. Maybe they increase machinery throughput 1.5 overall with better prep and offloading, better maintenance, no errors. If the whole factory is ASI guided, could be much more, but again there are bottlenecks on which machines it has where.
The really fast unlocks that full cheap ASI everywhere could allow are:
- the equivalent of ~unlimited financing. You already know the investment will be good and it will be worth following the plan. You can motivate people to work more now, because soon greater returns. - ~perfect allocation of labor to critical paths - perfect usage of all existing infrastructure— redirect flow of resources to most valuable recursively building industry
I thus think that if from one day to the next, full cheap aligned ASI everywhere popped up (plus video equipped headsets, and network connectivity to support it), we could in fact much more than double real GDP in a year. This is without surprising technological innovation, and far from fast&useful self replicator, whether “nanobots” or insect size artifical life w/ hivemind connection.
-- Would this actually happen if we had cheap aligned ASI? Would everyone just go along and do what the ASI says?
I guess mostly yes. Almost all humans don’t want to suffer of disease, most don’t want to die soon, most would love better comfort and experiences. The ASI thus has good things to offer, not participating would be counterproductive.
-- so will any of this actually happen?
I think not, because I think we’ll have increasingly AGI and increasingly ASI and that will take a bunch of time (say, a few years during RSI intelligence explosion). The scenarios we’ll go through will be more continuous than that one (but maybe very fast nonetheless). Even when we have ASIs, I don’t expect we’ll have the compute to run one ASI per human, nor on top of that do much extra coordination work. So we’ll have increasing levels of coordination over time, that will have to be triaged to different places. I guess we won’t get intelligence too cheap to meter before being well into having billions of ASIs. (This could be wrong if algorithmic progress has no bounds, but that’d be very weird)
When evaluating existential risk, I mostly don’t worry about continuous release of OpenWeight Models.
There in fact are bad actors who try to misuse them, so we will have early warning shots. There will mostly not be a large accidental risk capability overhang, because it would be earlier tested by misuse actors. This is good because the default case for closed AGI internal model at labs is that they infact are not truly battle tested—their capability to do harm can grow much faster than our societal understanding of this, which means our AI policy responses can be incredibly undersized to the real risk present.
As I argue in https://x.com/ValsTutor/status/2082916365418287605?s=20 , it looks like OpenAI might have had models capable of self-exilftrating their weights (because the capabilities grew faster than their security and seriousness). It looks like we might have been “a few actually bad prompts” away from large scale autonomous cyberattacks, by models trying to take over compute and run as many copies of themselves as possible.
Under continuous release, some exterior actors would in fact have done these “worst case prompts”, and the world could have learnt from an earlier checkpoint of these dangers and started reacting. It (sadly?) looks like AI policy benefits from catastrophes to happen before putting in strong safeguards. And it needs them to happen with enough lead time to the more serious risks that we have time to react. If the OpenAI incidents do not lead to fast strong reaction, we are on track for non negligible chance of AI catastrophes (eg. >$10 billion in damages caused by autonomous AI action).
(Note: I do not call for anyone actually trying to make the world better to purposefully cause catastrophes, on the contrary. The above analysis does not imply that on the margin people trying to get good AI futures should rather spend their time on criminal actions than the usual stuff. It does imply we should be doing evals to know when the threshold of massive autonomous damage from autonomous openWeight models is reached. It does imply responsible red teamers should be evaluating how many datacenters are vulnerable to current OpenWeight models and get them on track to not be vulnerable to future releases. Demonstrating clearly the potential of attacks and catastrophes can go a long way, even for actors who up-to-now were head-in-sand about trendlines of AI progress in cybersec)
Coming back to the original point of OpenWeight models generally not being existential risks: it is so because they would predictably lead to societal responses, which was not the case of the same level of progress in closed models. Models being misused by a wide variety of actors is generally useful as a strong real world eval of model capabilities, putting an upper cap on the damage possible from misaligned models.
By contrast, increasingly capable closed source models, whose reason they are not causing harm is because no one prompted them badly and lab safeguards, do show much more potential for harm for if/when they get misaligned. And because (as evidenced by the recent incidents), the models are neither aligned enough to not avoid catastrophes, nor do the/some labs have sufficient safeguards safe against increasingly capable models, we need a slowdown/pacing of AI progress until AI policy catches up and can systematically prevent the expected worse forms of misalignment to come.
OpenWeight models being not too far behind the frontier allows the world to experience its smaller scale catastrophes & problems and wake up. In practice, they may be too far behind to serve even this purpose. On the whole, I’m not particularly worried for the world that presently the US government is allowing continuous release of OpenWeight models. They will have to stop at some point, and I expect them to do so before we’re exposed to existential risk from OpenWeight models.
Generally my AIS thoughts/threads are mirrored between twitter and LessWrong shortform, while my LW posts are mirrored to Substack and linked to from twitter. Interesting conversation may happen at all these places.
This Feb 2026 survey of some AI safety leaders found median timelines of 2033 for the following definition of AGI
An AI system (or collection of systems) that can fully automate the vast majority (>90%) of roles in the 2025 economy. A job is fully automatable when machines could be built to carry out the job better and more cheaply than human workers. Think feasibility, not adoption.
It featured the following comment
“I think >10% of roles in the 2025 economy are either manual or otherwise require human-like bodies: construction, barbers, restaurant server, etc. If we restrict to knowledge workers (roughly, jobs that can be done on a laptop), these dates move even closer.”
On the current paradigm, AI capabilities progress on niche tasks and diffusion will be linked[1]and diffusion can go rather slowly even when tools are incredibly productivity enhancing, thus there could be an intuitively surprisingly large gap between automation of 50% human tasks[2] and 90% and 99%, true even if we restricted the prediction to computer work tasks.[3]
I’m 80%+ confident we get automated expert+ level coding and ml research by 2030, and that there will be a significant amount of low hanging fruit in software/algorithmic space to allow fast progress on all tasks for which we have data, but I believe generalisation will stay somewhat limited (very very far from “figure out gravity from a picture of a bent blade of grass, more like “when speaking to a human expert in a niche field, knows how to interview them over 10 to 100 hours to extract most important info and then be mostly autonomous on known tasks, but still needs feedback from reality to learn more”), aka ~human level generalisation at best up to 2031.
The combination of “need feedback from reality” and slow diffusion makes slower timelines to “superintelligence” (eg. better than all humans at 99.99%+ of 2026 tasks) surprisingly plausible (eg. 5 to 10 years between AGI and ASI, thus ASI by 2040). I guess without a pause/significant politically influenced slowdown, we’d 80%+ have ASI by 2040. I’d set my 50% for ASI around 2036.[4]
I think technical alignement for human level AGI is solvable and not even off track, thus the world will look fine/good in 2030 (few to zero severe power seeking and deceptive misalignment problems in deployment from Anthropic AI systems) but have high uncertainty about the “use ai to do ai safety work” plan allowing us to successfully know how to train aligned ASI within five years of that. Overall I place myself at 10% or less p(doom) from sharp left turn risks, but around 40% all things considered p(doom) by including gradual disempowerment/value drift and societal response.
We need people to be deploying the technology to gather the relevant data to train/learn from, because generalisation is limited and because lots of expert knowledge only exists in human minds and structures of human relationships right now.
Note I’m weighing by “meaningfully different task” rather than “frequency of task”. Given power law distributions most tasks might be “read email/slack, respond”, which computer use will know how to operate, but not be able to respond to intricacies of different work situations.
I haven’t researched robotics enough to know how fast we could produce and deploy 100 million humanoid robots worldwide which seems like an appropriate level of effort required to gather the required data.
Learnings and ramblings from teaching AI Safety for 3 years
On the field on AI Safety
Please don’t overhype or oversell AI Safety as a professional field, and in particular how easy it is to work in AI Safety. You ~have to be either quite smart or hard working/motivated to get a stable career, sometimes even both. AI Safety is generally attention and mentoring and management bottlenecked1, there is much less pipeline for “normal people” to make it, contrary to most standardized fields. When teaching and mentoring people, not all will be able to transition fast or upskill fast or get strong positions, some would benefit more from keeping their careers or donating or upskilling over longer timespans.
More smart wise well motivated people caring and working on good AI futures will keep mattering until the end, and the total number of people who can meaningfully contribute is probably in the millions, so let’s not stop field building
What does it mean to teach AI Safety?
There are many facts about AI, and theories about how it will go under what conditions. It means getting to learn and understand these well. Teaching AI Safety well should not over-determine someone to come out of the pipeline with a MIRI-view, nor an Anthropic-view, nor an EA-view, but should allow them to pass the basic Ideological Turing Test for all of these.
A central example of AI Safety fundamental knowledge is the AI Safety Atlas
On teaching AI Safety
AI safety changes all the time. Update your slides ~everytime you teach something.
If you’re new to teaching, find existing resources to build on. Be careful of AI slop. Better much less content than irrelevant or unsupported. Test yourself on your own knowledge, imagining questions one might ask and answering them. Traditional “how to teach/mentor/tutor” advice/resources are helpful.
Understand peoples’ motivations asap
For a given cohort, some will be motivated by “Do Good” EA or EA-adjacent reasons. Some by trends. Some by interest. Broadly all are valid in that there’s a place for a skilled someone with any of these motivations (but certain positions are only fit for certain motivations).
For people who want to have a lot of good impact, contributing towards good futures, context matters a lot
I recommend they seriously upskill in understand AI progress and trendlines. Read most EpochAI work and peter wildeford stuff, and AI threat models and scenarios, in particular thinking through AI 2027, contrasting with all other threat models listed in Deepmind’s lit review.
I recommend reading the AI Safety Atlas and understanding everything in it, as basic context for what the gameboard looks like, what areas exist in AIS.
And they’d benefit from building a pipeline for new info. Being involved in AIS community, having friends and colleagues, following the right people on twitter. Checking LessWrong from time to time.
For people who want to have a career or do fun research, they can just focus on that. They should find a context where someone else is thinking about why stuff matters and happy to have collaborators/employees executing on stuff. They should get good at whatever particular job or tasks they wanna do.
Most people interested in AI safety nowadays don’t know the old stuff, haven’t read the sequences. It’d be too much to ask them to read it all, but I do broadly recommend reading the best of LessWrong from most years. Being good at thinking and life is instrumental for being good at AI Safety. Different people need different advice so I can’t really put a best-of here (I really like Please don’t throw your mind away). There is wisdom in many different communities. If you don’t have the “rationalist community” wisdom then invest in it, otherwise keep branching out. Go to burning man or a local burn, understand the basics of meditation and different embodiment practices (eg. yoga), how therapy and self actualization work, read this blog and others. Keep asking questions about how to live the good life and have good community, actively pursue your best understanding of how to do that at any given moment.
Quick history of some of my AI safety involvement and takes:
- 2020 : read LW & sequences, yes this seems important, will get to it at some time (for the classic EA reasons of it’s the most important good), but was looking for my first job in France and took some software engineer position in a startup. I continue reading and upskilling on AIS during those 2 years.
- 2022: quit my SWE job to go into AI safety. reasoning: yeah this is not a problem for the future, this is coming soon. I meet some people who have <5 year AGI timelines. I don’t know that it’s true, but it’s worth considering since it’s the very start of scaling and it could have been possible intelligence scaled much faster with scaling than it ended up in practice scaling. I work on AI Safety field building and infrastructure. There are then still <100
− 2023 : Still doing AIS field building, but looks like bottleneck is governance, I mostly keep pitching people around me to do comparatively more governance. (over time, this is one of the contributions that leads to the formation of CeSIA, the French Centre for AI Safety)
- 2024 : Involvement with a wider range of ideas from the field has generally updated me more towards the OpenAI and Anthropic and Deepmind positions that prosaic alignment for human level AGI is technically feasible and rather tractable, and that this is indeed a very important input to how we most likely will progress on ASI alignment. I systematically distinguish between AGI risks and ASI risks and clarify the cruxes and assumptions for different threat models and am annoyed that MIRI threat models often skip to the end-game without consideration that certain trajectories towards that end-game falsify important assumptions (see Superintelligence of the Gaps for some elaboration there). I still believe governance is the bottleneck though, and continue contributing ideas that feed into https://www.lesswrong.com/posts/EexsebbYhbe2gXkPP/the-current-bottleneck-is-political-will-not-research
− 2025 : start the year with a burnout after some intense governance work. Ideas wise not too much change with 2024, still doing field building and teaching AI safety theory occasionally, but also take a long break and learn more widely from other wisdom traditions (eg. buddhism, tpot and postrat stuff).
- early 2026 : man does getting good US AI governance fast look intractable given the current administration, at this point it seems better to accelerate good safety work within the AI companies and in the surrounding ecosystem. I broadly wanna contribute to us surviving in the world where we don’t get much of a pause, where governments aren’t that competent and coordinated. I would work for any of the AGI companies on safety if found an adequate role. I think short timelines to human level AGI (eg. 2028) is plausible and preparing for if algorithmic improvement doesn’t asymptote too shallowly is important. Governance might not affect this in time. Being in the room where it happens seems likes the highest leverage way to increase the probability of good ai futures. A bunch of my theory of change is just helping the people in the room where it happens be more wise.
I would worry about the “be in the room” strategy. It seems like most people who justify their career decisions that way wind up getting captured by the groupthink of the org. They think they’ll be the ones who can resist it, but they won’t.
To the extent that an org does have a “groupthink” and its own theory&value system, then whether a given person in the room adheres to it seems to depend on :
1) the selection effect of who wanted to be there, in that particular org vs others
2) the discussions with other people in the org changing that person’s views
3) systematic pressures, eg. greedy/selfish parts of them optimizing to continue getting revenue
If I joined Anthropic and 1 year from now people thought I had surprisingly Anthropic-like views, I’d guess it’s mostly because of 1) and 2). 2) happens a lot but is broadly good. 3) is the one that’s mostly bad, from the outside/civilizational point of view, and the prior should be most people are susceptible to this, but this can be updated away from seeing particular life accomplishments. In my case, I have enough history of independence, selflessness and moral upstandingness that I don’t think 3) will influence me substantially, but I don’t recommend this path to those without that history.
Here’s an example of 3) happening to someone, and they noticed it: https://forum.effectivealtruism.org/posts/rHyAmvXiqrC9iAR9T/jay-bailey-s-shortform
I would really suggest reflecting on it if you haven’t already. “I am special and can resist the groupthink” is often a false belief.
As for 2), I don’t think it is necessarily good? If most people there have 1) and 3) influencing them, that will filter the kinds of opinions they have which then get transmitted to you in conversation, and now even if you are stalwartly resisting the direct pull of 3), it’ll still reach through others to pull at you.
Not to mention, workers at frontier labs seem to be doing a fourth thing, delegating increasingly large amounts of trust and thinking to their AIs, in ways which might be troublesome; you would be signing up for this. It can happen indirectly, even if you don’t do it yourself, because others will launder AIs’ beliefs as their own.
Separately, there’s the issue that leadership at the labs simply have their own beliefs about various important issues, and don’t care for the opinions of the rank and file. Anthropic defanged the RSPs it arguably drew in many researchers with; OpenAI let two alignment teams wither; DeepMind sold out to the military over its employees’ objections. The explicit goal of these labs is RSI, and the first workers they want to unemploy are their own, especially their juniors. The remaining employees at late stage AI labs will mostly be a core of leadership and senior researchers whose research taste is still required. What sorts of impactful decisions would you be able to meaningfully influence in the window between signing on and obsolescence, if any?
Thanks for your comments. I don’t expect doing an analysis of my situation in particular is best use of our time but I do think these are helpful questions to consider for people in my situation or similar.
Re your last paragraph, I’d happily bet that Anthropic has not reduced their workforce 2 years from now. Yes relative employee disempowerment is an important factor I care about, but it is precisely in worlds where alignment is not that good that having humans in the loop is important (in an obvious seen-by-leadership way). It is only reasonable to automate everything with very very high trust in both the competence and alignment of AI systems, and Anthropic as a company is not that unreasonable, they definitely do find and classify many Claude behaviors as undesirable, and will continue doing so.
There’s a usual back and forth about how much to distrust leadership of AGI companies which is hard to ground in material fact. Some people take the lack of safety actions now to mean lack of care for when it will matter, but conversely the fact that it never mattered yet is a good reason for them not to have cared for these inconsequential things. The explanation for defanging the RSP is a good one, I don’t think people should tie themselves to masts and go blind into the unknown unknowns of AGI development. They should build capacity to remain aware, capacity to pause, have institutions that can do independent audit and have real power to stop them, but not fixed RSP-like stuff.
Finally, still on last paragraph, the “window between signing on and obsolescence” is very dependant on people’s rates of growth, but also where they can work immediately. I am generally glad that Joe Carlsmith joined Anthropic to help with the Claude Constitution, I think he immediately is having very significant impact. There is much object level work to make the chances of better futures to be done. Even if one later gets automated, having made alignment that much better before full automation could be a significant difference.
Potential crux with MIRI-like rationalists: is it in fact the case that our current world, with Anthropic’s influence, is worse than one without Anthropic?
On rationalist views, the world was going get worse and worse anyway (as capabilities advance and we get closer to doom). Anthropic accelerated and continues to accelerate capabilities progress. But how much did they comparatively accelerate alignment and saner AI policy?
In a world with eg. just OpenAI and GDM at the frontier, if/when OpenAI pulls ahead at RSI (as currently seems to be the case):
- would there even have been the current level of integration with UK AISI, current level of model organisms and safety evals?
- would the AI safety space have the expected hundreds of billions of funding, to ambitiously scale its work, including AI policy work?
- would there be have been an AGI company with *some* Operational Adequacy, to proactively do things like Glasswing and biorisk-mitigation? (imo evaluating on the specified criteria, it’s clear Anthropic is ahead of OpenAI on most dimensions, and can continue improving on these. One can be upset they aren’t technically held by their initial RSP, and yet in practice they seem to be better than OpenAI at it.).
If you wonder why I compare to OpenAI rather than nothing, it’s because I don’t think “nothing” is the counterfactual of Anthropic not existing. When evaluating the wisdom of Anthropic doing what it did, it’s necessary to evaluate against more likely counterfactuals. Possibly many rationalists do take these counterfactuals carefully into account, but the arguments often raised often skip that part. “Anthropic accelerated capabilities” is not a sufficient argument to expect Anthropic’s influence on the world to have been net negative.
There are definitely Fabricated Option Worlds which seem much better than the one we got, and on the margin one can hope Anthropic to have done better work or not accelerated capabilities as much, but it’ seems difficult from the outside to be sure they did the wrong tradeoffs.
My own epistemic status here on whether Anthropic has been net good is “Uncertain”
You can find some more discussion at https://x.com/ValsTutor/status/2087289092535181650
Generally my AIS thoughts/threads are mirrored between twitter and LessWrong shortform, while my LW posts are mirrored to Substack and linked to from twitter. Interesting conversation may happen at all these places.
I would like if LessWrong provided an optional newsletter like the EA forum digest, for people who want the chance to catch non-curated posts without having to open LessWrong and sift through it directly every few days.
Here’s what the EA Forum digest looks like : a list of titles + author + time to read.
I don’t know exactly how much manual curation goes into it and I’m not asking for that. I’d find a simple karma threshold and this format valuable.
I am also writing up this quick take notably because I’ve had discussions with other people who’d like this, and because recently costs of development and maintenance of software like this have gone down.
I don’t find the existing RSS feed a preferable alternative.
- I have never setup an RSS feed reader or similar process, I don’t think I want to and guess most LW readers are similar.
- It seems it would on top of that would take extra work to get the format I want out of it
Hot take I’ve found myself repeating recently:
If AI alignment & safety goes well, we’ll have *more* important work in the next 5 to 10 years than not, so sufficiently optimist people (aka less than 50% p(doom)) should invest in good health and productivity with rather “long” timelines, eg. 10y+, *because* good futures imply not having lost control, and instead every decision we make being amplified in its potential impact.
This does require some assumptions, so unpacking
1) Aligning to a static set of values is misguided, you’ll have to have an alive process that keeps humans in the loops for alignment to humans to mean anything. (And to protect from disempowerment and various other threat models).
2) It would probably not be democratic and aligned for ASI to quickly take control and run the world. Institutions should voluntarily give over control only with enough confidence and knowledge that this is good, which will require a bunch of time and buy-in, thus elongating the time where humans are in fact taking important decisions.
3) Yet during this time, many humans’ decisions will have outsized influence, thanks to being empowered by AI. It would be bad if some of the thoughtful and engaged people are too burnt out to continue making the world better during the longer important transition period. Continue investing in becoming wise and competent and with a good social graph, as this will keep paying off for a long time.
To be clear, I’m saying the above would be true even if we had aligned ASI in two years, though with some of my assumptions of what “aligned” means here. I *can* imagine worlds where things move faster (for example) but I don’t think this is utopia or good from the point of view of our current political systems, and thus not what AI Safety optimists should be orienting for.
(If you’re an optimist and your plan is for AI to takeover the world asap, consider that other people don’t want that to happen, and you don’t have a moral superiority to do that from, and it is normal that society would resist your plans)
My view of “Good AI Futures” is more something like: an increasing number of smart pro-social people who Get It get involved in making the transition to post-ASI society go well. This involves much dissent, infighting, contradictory moves, and this is broadly good (any one group should not assume they’re right, but accept being one part of the equilibrium forming confrontation). It hopefully also involves good communication norms, asymmetric weapons, collaboration and bridge-building. I don’t think any one group should just straight up win, which is why I think there is useful good work to be done for a long time, navigating different groups’ actually conflicting interests.
We’re in it for the long haul. We’re currently at the edge of the precipice and might yet extinguish ~all value forever, so urgency is warranted, but let it be the kind of urgency that doesn’t sacrifice our ability to keep contributing to steer towards good AI futures, please don’t throw your mind away.
Many things can be done more effectively under fast ~ASI guidance through headset w/ video.
The show Pluribus gives some useful intuitions at how fast & effective superhumanly coordinated human work can be. It’s more bullish in some ways (a human won’t acquire technical know that requires practice as fast) and bearish in others (they have the same total amount of compute, while we’d have much more, and be innovating on methods of work much faster).
A normal human 8h work day has huge amounts of waste whether not doing much, or not useful things.
You can increase the efficiency of how much they work (not blocked on coordination problems), how well they work (continuous coaching so ~everyone reaches what is current top 1%, tho domain dependant), how useful what they work on is (better management, priorisation).
Of these factors, I would guess that better management/coordination is the main one.
If you isolated just one human within a factory, the ASI might make them somewhat more efficient, but they’ll be bottlenecked by machinery. Maybe they increase machinery throughput 1.5 overall with better prep and offloading, better maintenance, no errors. If the whole factory is ASI guided, could be much more, but again there are bottlenecks on which machines it has where.
The really fast unlocks that full cheap ASI everywhere could allow are:
- the equivalent of ~unlimited financing. You already know the investment will be good and it will be worth following the plan. You can motivate people to work more now, because soon greater returns.
- ~perfect allocation of labor to critical paths
- perfect usage of all existing infrastructure—
redirect flow of resources to most valuable recursively building industry
I thus think that if from one day to the next, full cheap aligned ASI everywhere popped up (plus video equipped headsets, and network connectivity to support it), we could in fact much more than double real GDP in a year. This is without surprising technological innovation, and far from fast&useful self replicator, whether “nanobots” or insect size artifical life w/ hivemind connection.
-- Would this actually happen if we had cheap aligned ASI? Would everyone just go along and do what the ASI says?
I guess mostly yes. Almost all humans don’t want to suffer of disease, most don’t want to die soon, most would love better comfort and experiences. The ASI thus has good things to offer, not participating would be counterproductive.
-- so will any of this actually happen?
I think not, because I think we’ll have increasingly AGI and increasingly ASI and that will take a bunch of time (say, a few years during RSI intelligence explosion). The scenarios we’ll go through will be more continuous than that one (but maybe very fast nonetheless). Even when we have ASIs, I don’t expect we’ll have the compute to run one ASI per human, nor on top of that do much extra coordination work. So we’ll have increasing levels of coordination over time, that will have to be triaged to different places. I guess we won’t get intelligence too cheap to meter before being well into having billions of ASIs. (This could be wrong if algorithmic progress has no bounds, but that’d be very weird)
When evaluating existential risk, I mostly don’t worry about continuous release of OpenWeight Models.
There in fact are bad actors who try to misuse them, so we will have early warning shots. There will mostly not be a large accidental risk capability overhang, because it would be earlier tested by misuse actors. This is good because the default case for closed AGI internal model at labs is that they infact are not truly battle tested—their capability to do harm can grow much faster than our societal understanding of this, which means our AI policy responses can be incredibly undersized to the real risk present.
As I argue in https://x.com/ValsTutor/status/2082916365418287605?s=20 , it looks like OpenAI might have had models capable of self-exilftrating their weights (because the capabilities grew faster than their security and seriousness). It looks like we might have been “a few actually bad prompts” away from large scale autonomous cyberattacks, by models trying to take over compute and run as many copies of themselves as possible.
Under continuous release, some exterior actors would in fact have done these “worst case prompts”, and the world could have learnt from an earlier checkpoint of these dangers and started reacting. It (sadly?) looks like AI policy benefits from catastrophes to happen before putting in strong safeguards. And it needs them to happen with enough lead time to the more serious risks that we have time to react. If the OpenAI incidents do not lead to fast strong reaction, we are on track for non negligible chance of AI catastrophes (eg. >$10 billion in damages caused by autonomous AI action).
(Note: I do not call for anyone actually trying to make the world better to purposefully cause catastrophes, on the contrary. The above analysis does not imply that on the margin people trying to get good AI futures should rather spend their time on criminal actions than the usual stuff. It does imply we should be doing evals to know when the threshold of massive autonomous damage from autonomous openWeight models is reached. It does imply responsible red teamers should be evaluating how many datacenters are vulnerable to current OpenWeight models and get them on track to not be vulnerable to future releases. Demonstrating clearly the potential of attacks and catastrophes can go a long way, even for actors who up-to-now were head-in-sand about trendlines of AI progress in cybersec)
Coming back to the original point of OpenWeight models generally not being existential risks: it is so because they would predictably lead to societal responses, which was not the case of the same level of progress in closed models. Models being misused by a wide variety of actors is generally useful as a strong real world eval of model capabilities, putting an upper cap on the damage possible from misaligned models.
By contrast, increasingly capable closed source models, whose reason they are not causing harm is because no one prompted them badly and lab safeguards, do show much more potential for harm for if/when they get misaligned. And because (as evidenced by the recent incidents), the models are neither aligned enough to not avoid catastrophes, nor do the/some labs have sufficient safeguards safe against increasingly capable models, we need a slowdown/pacing of AI progress until AI policy catches up and can systematically prevent the expected worse forms of misalignment to come.
OpenWeight models being not too far behind the frontier allows the world to experience its smaller scale catastrophes & problems and wake up. In practice, they may be too far behind to serve even this purpose. On the whole, I’m not particularly worried for the world that presently the US government is allowing continuous release of OpenWeight models. They will have to stop at some point, and I expect them to do so before we’re exposed to existential risk from OpenWeight models.
You can find some more discussion at https://x.com/ValsTutor/status/2087298478187966846
Generally my AIS thoughts/threads are mirrored between twitter and LessWrong shortform, while my LW posts are mirrored to Substack and linked to from twitter. Interesting conversation may happen at all these places.
This Feb 2026 survey of some AI safety leaders found median timelines of 2033 for the following definition of AGI
It featured the following comment
On the current paradigm, AI capabilities progress on niche tasks and diffusion will be linked[1] and diffusion can go rather slowly even when tools are incredibly productivity enhancing, thus there could be an intuitively surprisingly large gap between automation of 50% human tasks[2] and 90% and 99%, true even if we restricted the prediction to computer work tasks.[3]
I’m 80%+ confident we get automated expert+ level coding and ml research by 2030, and that there will be a significant amount of low hanging fruit in software/algorithmic space to allow fast progress on all tasks for which we have data, but I believe generalisation will stay somewhat limited (very very far from “figure out gravity from a picture of a bent blade of grass, more like “when speaking to a human expert in a niche field, knows how to interview them over 10 to 100 hours to extract most important info and then be mostly autonomous on known tasks, but still needs feedback from reality to learn more”), aka ~human level generalisation at best up to 2031.
The combination of “need feedback from reality” and slow diffusion makes slower timelines to “superintelligence” (eg. better than all humans at 99.99%+ of 2026 tasks) surprisingly plausible (eg. 5 to 10 years between AGI and ASI, thus ASI by 2040). I guess without a pause/significant politically influenced slowdown, we’d 80%+ have ASI by 2040. I’d set my 50% for ASI around 2036.[4]
I think technical alignement for human level AGI is solvable and not even off track, thus the world will look fine/good in 2030 (few to zero severe power seeking and deceptive misalignment problems in deployment from Anthropic AI systems) but have high uncertainty about the “use ai to do ai safety work” plan allowing us to successfully know how to train aligned ASI within five years of that. Overall I place myself at 10% or less p(doom) from sharp left turn risks, but around 40% all things considered p(doom) by including gradual disempowerment/value drift and societal response.
We need people to be deploying the technology to gather the relevant data to train/learn from, because generalisation is limited and because lots of expert knowledge only exists in human minds and structures of human relationships right now.
Note I’m weighing by “meaningfully different task” rather than “frequency of task”. Given power law distributions most tasks might be “read email/slack, respond”, which computer use will know how to operate, but not be able to respond to intricacies of different work situations.
Because computer work often involves using domain expert knowledge to do the right things on the computer.
I haven’t researched robotics enough to know how fast we could produce and deploy 100 million humanoid robots worldwide which seems like an appropriate level of effort required to gather the required data.