PSA: We can do better
tl;dr: people should understand and think hard about the problems they work on.
We’ve observed that those who work in AI safety (ourselves included) often rely on concerning heuristics when choosing what to work on. Running a conference is probably good, doing pragmatic alignment research might be good, and as long as such objectives don’t breach our internal models of what could contribute to reducing x-risk, these things are “what should be done”. But using such vibesy thought processes don’t always produce “actually impactful work” that would beat a prospective counterfactual. We wrote this post to share our observations and figure out what we should be doing instead.
People don’t know what they’re working on
AI safety is talent constrained. However, simply inflating the field doesn’t solve our bottleneck; rather, we need more people who understand the core arguments of AI safety. You can’t determine how to meaningfully contribute to AI safety without deeply knowing the problem you are trying to solve. Many newer people (us included!) rush into research, fellowships, and the like without building the context necessary for navigating the field.
Agency-maxxing is not always good
Moving fast is good. Moving too fast leads to poor ToC and sloppily executing projects. Many people working in AI safety seem to downweight spending time thinking about the actual problem they are trying to solve and backchaining from it, in favor of moving as fast as possible to get things done quick and dirty. While this enables faster execution and increases the volume of work being done, it doesn’t always move the needle on effectively reducing x-risk.
The problem with force multipliers
We categorize ToCs that focus on enhancing the impact of others as force multipliers (think upskilling, infrastructure work). There’s strong incentives to become a force multiplier: enhancing the impact of others through auxiliary work is (arguably) higher ROI and often more approachable than “directly” working on AI safety.
However, when everyone becomes a force multiplier, no one ends up directly working on the problems we actually care about. We think this work is extremely valuable, but we believe there is comparative advantage in focusing on the direct work, and that a small number of force multipliers can produce the same (if not larger) multiplicative effect as current efforts.
Harshul: In Balatro, your final score is calculated as the number of chips you score (blue) times your mult (red). I’m skeptical of a mult-heavy build (focusing on auxiliary work) without a strong base of chips to actually multiply (direct safety work). I’m more bought into a chip-heavy build (focusing on direct work, which I believe benefits more from added effort) alongside a smaller fraction of the field providing sufficient mult (auxiliary support).
Deferring thinking to others
Our subfield includes a multitude of conflicting agendas: whether or not alignment research is “good” to do, how much or little we should be focusing on political advocacy, how effective comms work is and could be, etc. Such conflicts don’t yet have a resolution. But, in order to “do things”, people will defer to a respected figure as opposed to building their own reasoning for their models. They will defer on what things are good and bad as a way to consolidate their consciousness and start doing things to “make impact”.

As an example, upcoming technical researchers will often defer to what is popular as a proxy for impact. They may look at places like Redwood or Anthropic, and attest to wanting to work there to “do the most impactful work”. But in reality, it’s very much up for debate as to whether a technical researcher should be going to said places in order to produce counterfactual impact (not to mention whether they should be doing technical research at all). But they instead defer to the vibes/popularity these places hold and conflate this with prospective impact.
Streetlighting
AI safety is hard, but not everyone ends up working on the hardest problems. Instead, people focus on more tractable problems that may seem nearly as impactful, but actually fail solving the true bottlenecks at hand. We believe this arises through some combination of people getting turned away from working on hard problems after getting stuck, and people prioritizing fast execution over thinking deeply about their work. When pushed to “just do things”, it’s often easiest to “just do” the easiest things.
(Note: Hard problems aren’t impactful by default, and easy problems aren’t auxiliary by default either. This section applies to the set of hard problems that are actually useful to work on.)
How to avoid these:
Build your models
Pursue the fundamentals: Read early AI Safety/alignment work, sequences, etc.
Read more modern work, stay up to date on meaningful new literature
Reflect on your reading regularly. Analyze when you update and what uncertainties you have after reading a piece.
Construct feedback loops
Seek lots of feedback on your direction/work from different demographics, especially those with conflicting views.
Work in public and let others help you course correct
Try and maintain strong epistemic, and avoid subtle dishonesties
We do not claim to have it all figured out ourselves. We have not thought as deeply as many others, and this post is not meant to be a “call out” of people who have exhibited the aforementioned behaviors. Rather, this is a call for improvement, and a raise in standards. We do not have to sacrifice the understanding of our problems, haphazardly defer, or avoid what is difficult. We can try harder, be smarter, and raise the bar of the work we devote our time to.
It’s really hard to get people to think hard about the core of the problem. So thanks for highlighting this and sharing your journey!
It’s not even just beginners that don’t like to go deep on this; MATS has tried running a strategy curriculum but fellows were apparently like ‘but we’re here now so we want to do the (empirical) research we came here to do’.
I think such thinking is important and neglected, so to get more people to do it, I recently co-created a new advanced AI Safety strategy course with my colleagues at Lens Academy; Forecasting, Modeling, and Shaping AI Futures.
Our motto with this course is something like: “What should people know that they wouldn’t otherwise come to know (because the topics aren’t sexy enough)?”
I might include this article in one of our courses. It seems pretty nice at conveying some important considerations to people.
One thing that is not mentioned here, but I have a suspicion that it is highly relevant, is funding. I.e. we should strongly consider the hypothesis that people are not thinking hard about what is worth working on because they think they will not be able to act on the answer anyway, because funders choose what work is done. (At least, for work that benefits from compute, collaboration, or full-time attention.)
I guess there is an argument that funding decisions are somewhat elastic, and a large population of talented researchers eager to work on a particular direction can point more money in that direction to some extent?
I definitely second that but it’s a chicken and egg situation. Some people jump into fellowships and research to know more about the context so that they can make better decisions. Sometimes we can’t learn more about the field and ourselves by simply reading. It’s a learning journey given that the field is progressively evolving. Also, I’m glad you point out the fact that this field need people to “understand the core arguments of AI safety”. I’m really in favor of people slowing down a little to deeply think about how they can better contribute to AI Safety but can this be achieved given that people now expect efficiency, faster iteration and results due to the capability of AI and also that (major) panic that we are short on time and AGI is apparently gonna happen in a few years? Sometimes it’s so hard to shut the noise and do what is right for the field but there are practical constraints for the talents we are trying to attract. But nonetheless, we should remain hopeful.
It’s quite possible that you’re directionally correct in terms of many newcomers to the field not understanding what they’re working on. But personally, I spent a good amount of time while getting involved in AI safety research thinking about high-level strategy and ToCs for different research directions, and it mainly left me stressed, confused, and demotivated. I found it difficult to get excited about research when I was constantly second-guessing the ToC and wondering whether we’re in the large portion of worlds where whatever I was working on ends up being useless (or even net negative!). I’m not saying it was a complete waste of time, I think I did learn from it and maybe I’m a better researcher now because of it. But, like, I think it can definitely be overdone.
There’s also potentially an argument to be made that not everyone needs to be a strategist. Like, maybe there should be room in the field for people who don’t want to spend time thinking about big-picture strategy, or for outsiders who might not know or care about all the deep theoretical arguments for AI risk but who can still see that building ASI is probably a bad idea and want to help, and maybe it’s fine for those people to defer to others on what to work on. But this does raise thorny issues with avoiding various forms of capture or value drift, which I don’t really know how to deal with… I’ll leave that question to the strategists!
Hmm that’s tough.. And I imagine you’re not the only one to have this reaction. This does seem to be a legitimate downside of thinking through the high-level. It does feel somewhat important (for a good chunk of people) to go through that, though, I think. Like you say, it makes you a better researcher.
Sure, not everyone. Though more people should probably have better strategic takes than is the case now. One reason is that AI Safety is pretty pre-paradigmatic. Another one is that impact is heavy-tailed and I expect people with good strategic takes to be strongly overrepresented in that heavy tail.
Lovely post.
Without citing any specific incidents being pointed at where “this person didn’t think it through” is clearly at play, it’s hard to form more tangible hypotheses about why this is, but when I introspect about this in my situation, these are the reasons I come up with.
1. Thought leaders lead for a reason, usually. Now that I’ve read articles on 80,000 Hours, read a bunch of EA books, and taken the EA introductory and in-depth courses, it’s readily apparent to me that I’m not at the top of the intellectual bell curve. When I read an article by Holden Karnofsky, for example, it’s obvious that he’s done a better quality of work thinking about a given topic than I could. I don’t have any reason to believe that I’m going to be able to come up with anything better, so I default to believing whatever the smarter person has to say.
2. Thinking is hard. I remember the author David Epstein quoting someone, saying that we believe that the brain is designed to think, but it’s actually designed not to think. Biologically speaking, thought is expensive, so our brain comes up with all kinds of shortcuts and heuristics to avoid having to think to much. I rely on others who are better at thinking (or just have more knowledge on a given topic) than me because it’s “cheaper” from an energy perspective.
3. Thinking takes time. I have a job, friends, hobbies, etc., and I tend to prioritize them over re-deriving a Theory of Change from first principles because they’re instrumental to my wellbeing. Unless there’s a clear problem with the Theory of Change that necessitates me going back to square one, I’d rather just get on with my life.
I don’t say all of this to naysay your point; I’m not about to go onto the LessWrong forum and speak out against critical thinking. In fact, I think what you’re saying is very valuable. I just think it’s also worth understanding why people are the way they are.
Your first point is true to some extent but I think one of the points of Tsvi’s first article commented below exposes a problem with it:
”Deferral-based opinions don’t contain the detailed content that generated the opinions, and therefore can’t direct action effectively or update on new evidence correctly.”
Another problem is that if people defer, they will never get to that expert level.
But yeah, overall, I agree with the sentiment of your comment and I notice similar patterns in myself.
Totally agree. It’s the same reason they say that the best way to learn something is to teach it: it forces you to understand something from the ground up. But if rederiving something from scratch takes a lot of time and energy, then there has to be a good incentive for doing it. I think that’s what’s missing: incentive. “People should do difficult thinking to produce more impact” is a fine value to have, but without being able to point to any specific problem being caused by its absence, there’s no motivation for anyone to enact that value. It’s hard to be motivated to solve an abstract problem that may not even exist when there are real ones sitting in front of us.
Yup that tracks. (Mainly thinking out loud to see what I can do with my org:) I wonder if some sort of proof-of-work / proof-of-understanding incentive structure might help. I think one reason people focus on more concrete things in their upskilling journey is that it’s more legible to outsiders / offers more rewards/prestige/status. E.g. if you do some fellowship, that’s good for your resume and good for concrete artefacts. Deep (abstract) thinking doesn’t have similar legibility (yet), i.e. “I thought hard for 100h” probably doesn’t do much on a resume.
For this reason, EA organizations seem to rely less on credentials and more on short answer questions, writing samples, structured interviews, and work tests.
Then again, there is a weakness to just relying on that information. A candidate might be very intelligent and have thought about a problem very deeply but then have no track record for doing the actual work: a “living in the clouds” intellectual. Credentials and achievements are usually used as proxies for past evidence of hard work and general career direction, even if they don’t necessarily filter for critical thinking.
What does ToC stand for here?
Theory of Change
(Cf. regarding deference: https://www.lesswrong.com/posts/ksBcnbfepc4HopHau/dangers-of-deference https://www.lesswrong.com/posts/jzy5qqRuqA9iY7Jxu/the-problem-of-graceful-deference-1 https://www.lesswrong.com/posts/zLG2DnJw6oEZqAgaE/tools-for-deferring-gracefully )
Nice post! Why do you think more people are becoming force multipliers as opposed to working directly on AI safety? It seems to me (rough guess) that far more people work directly on things like technical alignment?