Thanks for the response, and for tolerating my rude pot shots re meta-science, where I’m sure you feel I’ve misunderstood your views (just like I feel like you’re misunderstanding mine). It might be fun to chat in person at some point about the meta-science stuff. Getting to the more immediately cruxy stuff (for me)...
My sense however is that most of them were ways we “got lucky” rather than things we did well. And so I’m particularly wary of the inference “things are going better than expected --> our strategy is good”.
I definitely agree that it seems important to distinguish between getting lucky (and I think there’s been plenty of this) vs. safety work having actually been useful.
The main thing I’ll say is that you should think of “safe” as a totally different predicate for past systems and current systems and future systems. I think if you taboo the word (and also the word “aligned”) and try to figure out which more concrete properties of these systems might generalize to much greater capabilities, that would be far better.
Sure, here’s a sketch of what needs to happen for this plan to work out. Assuming (for simplicity) that we develop AI systems in discrete successive generations AI-1, AI-2, …, we need for each N:
AI-N does not successfully take over
AI-N is sufficiently useful for ensuring that (1) and (2) attain for AI-(N+1). This in particular requires:
AI-N to not sabotage or subvert our ability to do work involved in making sure (1) and (2) attain for AI-(N+1)
AI-N to be sufficiently capable at the work involved in making sure (1) and (2) attain for AI-(N+1)
I agree that the reasons AI-N doesn’t successfully take over vary as a function of N. E.g. right now AIs don’t successfully take over because they’re incapable of doing so, but later these reasons will be sensitive to our past actions in more interesting ways (e.g. how much we’ve hardened the world to cyber attacks, how well-safeguarded a model is, or how interested models are in takeover). This also means that the work that needs to benefit from AI assistance will change over time. (Though I think there’s likely broad patterns that enable us to prepare in advance. E.g. a basic case for scalable oversight work is that many of these types of work will be loaded on fuzzier tasks whose completion we can’t robustly score numerically, e.g. “investigate this incident and write a report that gives me a broadly accurate sense of what’s going on”.)
Maybe a useful analogy is the lead-up to World War 1.
Thanks, I like this analogy. Note though that this mostly bites for people who are hoping to do good by differentially advancing their preferred horse (e.g. accelerating Anthropic because it’s important for Anthropic to beat other developers). It’s not clear if it should apply to people who feel more like (1) medics during WWI, who wish the war would stop but figure it’s good to “participate” in the war effort in ways that mitigate its harms; or (2) (to give a more contentious example that feels more like the current situation for AI) engineers that work on weapons targeting systems, hoping to reduce the civilian collateral but aware that their work will also make their military more effective at waging war.
I’m well aware that there are many people with this theory of impact. I also agree that there are many epistemic distortions that make it look more compelling than it is. (To name another one, it’s easier to observe impact from doing stuff than from not doing stuff, which makes retrospective impact assessments systematically biased towards doing stuff, especially stuff that leads to having more power and therefore more ability to do stuff.)
That said, this isn’t so cruxy for me in particular, because I don’t personally view my main theory of impact as differentially advancing Anthropic. One way to operationalize this is: How would I feel about my research being published? If I think it’d be good for all AI developers to know about my work, then that’s good reason to believe that I wasn’t mainly doing it in order to differentially accelerate Anthropic. (Relatedly, I think requiring labs to allocate substantial compute to research that is made open is one of my favorite proposals for coordinating slow down and greater investment in safety.)
It is true that recently I publish less of my work than before, which is maybe a bad sign for whether that work is well-targeted. Just to name some reasons this happens: a lot of the work is “boring” stuff like “we applied obvious technique, overcoming mundane challenges specific to our training stack”; some of the work is entangled with proprietary information about Anthropic’s training stack; and some of the work also makes models more commercially valuable, such that I’d need to fight about publication with other people in the company who are more into winning the commercial race or worried about the effects of accelerating competitors.
Paul of course has published the various bits of work that you criticize him for (RLHF, starting ARC Evals to measure dangerous capabilities), so I don’t feel like the WWI analogy feels apt. (Of course, he, and I, could still be rationalizing the impact of our work in other ways.)
FWIW I didn’t find your message strident at all (and I’ve really appreciated your patience with me during this exchange). But yes, chatting in person sounds good—I’ll follow up with you over DMs.
That said, your last message has, I think, helped me understand what you might be getting at, so I wanted to respond again to explain (what I think) you’re saying in terms that make sense to me.
Here’s my interpretation of (one part of) your view:
<richard_according_to_sam>
Sam, let’s grant for the sake of argument that your safety research is in fact valuable (in the sense that its improves our ability to navigate transformative AI). Even granting this, there’s still an important negative externality of your work that you’re not tracking that comes via your association with “AI safety” as a brand/community/movement/field.
Namely, by doing your work under the “AI safety” banner, you lend credibility/power to that banner. This is bad. For instance, it’s bad because the “AI safety” banner has been co-opted by AI companies who use it to recruit for capabilities roles that accelerate the race. It’s also bad because it gives many people the cover they need to rationalize work on things that are harmful by your (Sam’s) own lights. Concretely, we might imagine AI company recruiters arguing: “By working for us, you’re indirectly supporting AI safety work like Sam Marks’s” despite recruiting for roles that you (Sam) think are bad for the world and don’t support your work.
Therefore, even if you’re right that your work is good for the world, you (Sam) should at least publicly disaffiliate yourself from the brand “AI safety” or “AI alignment,” in order to not lend support to a dangerous machine that you don’t control.
</richard_according_to_sam>
I think richard-according-to-sam’s perspective makes a lot of sense. My main issues are meta:
This is an argument that I should engage in a social-level dispute. But it’s well-known that humans are too drawn to social-level disputes and overrate their importance.
A while back, I decided to adopt a policy that my public communication should center on object-level communication on topics about which I have expertise. In particular, this means refraining from publicly denouncing group X or cheerleading group Y. When I write things in public, I would like readers to trust that my words can be interpreted in terms of their plain meaning, not as veiled gambits in a power struggle between interest groups. (I’m open to revising this policy, but I’ve been quite happy with it so far.)
This being said, there are some object-level points about prioritization and career choice that I’m happy to comment on:
If you’re someone trying to improve the world, I currently think it’s very unlikely that your best option is to work on accelerating your preferred AI developer.
Here are some options that I think are substantially better: METR, TruthfulAI, Redwood Research, Transluce, CAISI, UK AISI, joining safety teams at AI developers.
I also think it’s plausible that researchers on safety teams at AI developers should leave to work at these^ organizations, though in this case I have less of a blanket recommendation and think it’s more case-by-case.
I think the theory of change “have safety-motivated people work at labs as a form of political leverage within the lab” is currently weak.
I expect that lab employee influence will rapidly diminish as AIs become more useful for AI R&D.
Speaking about the situation at Anthropic in particular, I don’t feel like political leverage is a very important bottleneck.
When safety staff asks for things, I’ve felt that Anthropic leadership has been generally reasonable and responsive to good arguments. (Or at least, if they’ve tricked me into thinking they’re reasonable and responsive, they’ll likely trick you too.) Responsiveness to employee leverage hasn’t felt like a big part of the story to me.
In general, gathering better evidence (e.g. evidence of risk, or evidence that such-and-such intervention is good/bad) feels much more important to me than exerting more pressure.
Also Anthropic already has many capabilities staff who care about safety, so the marginal influence of another one is small.
When discussing career choice with staff at AI developers that you don’t have a strong reason to trust, keep in mind that they might selectively deploy favorable evidence, overstate the impact of your potential work, overstate the positive impact of the developer, or mislead you in other ways. It’s generally good to talk these decisions over with thoughtful people who don’t work at an AI developer.
On the value of “applied” alignment work (i.e. production-facing work that also makes AI systems more commercially valuable):
It’s a bit case-by-case, but this work often funges with capabilities work, in the sense that researchers who aren’t primarily motivated by safety can be reallocated to work on it.
That said, I do think it’s important for many safety researchers to gain experience working with production systems for a bunch of reasons: (1) making sure that safety researchers don’t miss simple/mundane ways to make marginal improvements (e.g. generically improving the quality of RL environments to make them less hackable); (2) disseminating learnings from trying to align production systems, which can often be informative for longer-term agendas (e.g. having a rich sense of how LLM generalization works); (3) gaining practical experience with productionizing safety interventions, to make sure that additional safety interventions can be implemented quickly and properly; etc.
All things considered, I think any individual person should choose between production-facing alignment work and other safety work based on their comparative advantage. Staff at AI developers that can do both should ideally rotate between the two types of work.