This is an email I wrote to @Buck after he gave a talk about his strategy confusions to a Constellation audience. The talk was about how Redwood’s work could have stopped the Hugging Face incident from happening, which would have been very bad for the salience of AI safety, and how he’s confused about that. I think this is a very important consideration which should be shared more prominently, but I also think ALL of Redwood’s strategic considerations should be shared more prominently. Here’s an email I wrote to Buck, lightly edited for clarity.
Hi Buck,
I’m pretty new to the AI safety space. I’ve loosely followed thing for a few years, but only really sunk my teeth in for about five months. Four months ago, I was mistaken about what leading safety orgs believed, and what their motivations were. For example, I thought that Redwood’s control regime was misguided, because it wouldn’t scale to ASI. I now believe (as Redwood does) that control is mostly good because it allows us to mine labor from, say, just barely superhuman AIs which could do tons of work to align ASIs. The point of this anecdote is that I think having transparency about your strategy beliefs as an org at the “center” of the AI ecosystem is really important for directing people who are newer to the space to grapple with the right strategy questions (and not jump in to doing useless stuff). I think many people new to the space pursue projects which are unhelpful because they are working under the worldview of “control good” rather than [insert convoluted Redwood worldview which probably endorses some kind of control research but not others]. I largely avoided this by being confused and being more interested/motivated by strategic and abstract questions than by hairy empirical ones. Probably many people who like thinking about hairy empirical problems and not strategy will just get nerd sniped and not do the best things no matter what you do, but we can at least move the space in the right direction.
Anyways, this summer I noticed I was confused about whether implementing control mechanisms right now is good at all. Warning shots matter so much! And then you gave a talk about this topic which I enjoyed, and my sense in the room was that the talk raised an important question that most people had not deeply considered. I’m very worried about the contribution of “AI safety” towards putting us in a looks good, is bad world [Buck’s framing from his talk]. I think you should make your talk into a brief write up and post it somewhere. In general, I think you should more prominently place your strategy considerations on Redwood’s website and research. I don’t know the best way to do this and I am sure there are significant trade offs I have not considered. But if Redwood isn’t transparent about strategy, who will be! Not just transparent somewhere in a blog post from a few months ago, but clearly transparent with a “do not try this at home unless you understand this blog post” disclaimer next to research that you think is probably good for a very specific reason. I would like the AI safety space to be more confused, more correct, and less headstrong.
Best,
Ben Pomeranz
@Cleo Nardo points out that this also has a lot to do with field building research programs, where people want to start doing research on day three but should probably be considering strategy for a few weeks. He also points out that the control stuff was preceded by posts like “the case for ensuring AIs are controlled,” but I think that these justifications should be featured prominently alongside the research that they justify, because many people will engage with and build off of technical work without looking back in time for relevant blogposts that may precede it. Also, he points out that Redwood’s more recent research direction into conceptual uplift does not come with a matching “the case for conceptual uplift” post.
I think RR are mostly blameless if junior people decided to spend 3 months doing control projects without bothering to read The case for ensuring that powerful AIs are controlled. It’s a 30 min read! You could probably read all the macro-strategy around control in 2-3 days, and maybe a week to absorb it.
The same goes for evals, ambitious mech interp, pragmatic mech interp, scalable oversight, etc. Upskillers working on X should be able to give 5 min answers to questions like: Why did people originally start working on X? Why was X not done before that point? What are the main arguments against X? What’s the crux between supporters of X and supporters of not X? Who are the central figures on both sides? What changes in the strategic landscape make X look better or worse?
I think most of the blame lies with:
[60%] People running upskilling programs, for not fostering a culture where mentees will naturally think “I’m going to spend 3 days reading and thinking about the case for this research” as Step 1 of doing research. My ideal upskilling culture would have mentees fighting with each other about whose research is best, and mentees would feel free to say in Week 3 “I‘ve decided evals suck, I’m joining Alice’s project on scalable oversight”. (I’m exaggerating slightly.) My experience talking to junior people is often like they think these questions are above their pay-grade or something.
[10%] The field-building grant-makers, for not making “mentees understand what’s going on and have good takes” a core desideratum for the upskilling programs. My impression is grantees focus on easy-to-verify successes like mentees joining full-time prestigious orgs, or conference papers, or impressive mentors.
[25%] The mentees themselves. Come on, guys! They are a little blameless though because they’re understandably paranoid about seeming “productive”.
[5%] RR for not stamping “please read this blog post before doing control research” on their papers. RR is pretty much pareto on (1) doing object level work and (2) explaining why they are doing it & addressing dissenters. The other top orgs seem worse at this.
I agree that Redwood has been historically very good at explaining why they are doing what they are doing. However, I do think that the posts making the case for AI control in particular are getting a bit old, and it would be very good to see updates on them in light of everything that happened in the last two years (e.g. Buck’s recent claim that it’s quite possible that it would have been net negative to implement AI control in the past, because it would have prevented the HF incident).
>Buck’s recent claim that it would have been net negative to implement AI control in the past Where did Buck claim this? I think he has stated confusion about whether it would have been net negative, but my understanding is that he has not come down decisively.
This is an email I wrote to @Buck after he gave a talk about his strategy confusions to a Constellation audience. The talk was about how Redwood’s work could have stopped the Hugging Face incident from happening, which would have been very bad for the salience of AI safety, and how he’s confused about that. I think this is a very important consideration which should be shared more prominently, but I also think ALL of Redwood’s strategic considerations should be shared more prominently. Here’s an email I wrote to Buck, lightly edited for clarity.
Hi Buck,
I’m pretty new to the AI safety space. I’ve loosely followed thing for a few years, but only really sunk my teeth in for about five months. Four months ago, I was mistaken about what leading safety orgs believed, and what their motivations were. For example, I thought that Redwood’s control regime was misguided, because it wouldn’t scale to ASI. I now believe (as Redwood does) that control is mostly good because it allows us to mine labor from, say, just barely superhuman AIs which could do tons of work to align ASIs. The point of this anecdote is that I think having transparency about your strategy beliefs as an org at the “center” of the AI ecosystem is really important for directing people who are newer to the space to grapple with the right strategy questions (and not jump in to doing useless stuff). I think many people new to the space pursue projects which are unhelpful because they are working under the worldview of “control good” rather than [insert convoluted Redwood worldview which probably endorses some kind of control research but not others]. I largely avoided this by being confused and being more interested/motivated by strategic and abstract questions than by hairy empirical ones. Probably many people who like thinking about hairy empirical problems and not strategy will just get nerd sniped and not do the best things no matter what you do, but we can at least move the space in the right direction.
Anyways, this summer I noticed I was confused about whether implementing control mechanisms right now is good at all. Warning shots matter so much! And then you gave a talk about this topic which I enjoyed, and my sense in the room was that the talk raised an important question that most people had not deeply considered. I’m very worried about the contribution of “AI safety” towards putting us in a looks good, is bad world [Buck’s framing from his talk]. I think you should make your talk into a brief write up and post it somewhere. In general, I think you should more prominently place your strategy considerations on Redwood’s website and research. I don’t know the best way to do this and I am sure there are significant trade offs I have not considered. But if Redwood isn’t transparent about strategy, who will be! Not just transparent somewhere in a blog post from a few months ago, but clearly transparent with a “do not try this at home unless you understand this blog post” disclaimer next to research that you think is probably good for a very specific reason. I would like the AI safety space to be more confused, more correct, and less headstrong.
Best,
Ben Pomeranz
@Cleo Nardo points out that this also has a lot to do with field building research programs, where people want to start doing research on day three but should probably be considering strategy for a few weeks. He also points out that the control stuff was preceded by posts like “the case for ensuring AIs are controlled,” but I think that these justifications should be featured prominently alongside the research that they justify, because many people will engage with and build off of technical work without looking back in time for relevant blogposts that may precede it. Also, he points out
that Redwood’s more recent research direction into conceptual uplift does not come with a matching “the case for conceptual uplift” post.
I think RR are mostly blameless if junior people decided to spend 3 months doing control projects without bothering to read The case for ensuring that powerful AIs are controlled. It’s a 30 min read! You could probably read all the macro-strategy around control in 2-3 days, and maybe a week to absorb it.
The same goes for evals, ambitious mech interp, pragmatic mech interp, scalable oversight, etc. Upskillers working on X should be able to give 5 min answers to questions like: Why did people originally start working on X? Why was X not done before that point? What are the main arguments against X? What’s the crux between supporters of X and supporters of not X? Who are the central figures on both sides? What changes in the strategic landscape make X look better or worse?
I think most of the blame lies with:
[60%] People running upskilling programs, for not fostering a culture where mentees will naturally think “I’m going to spend 3 days reading and thinking about the case for this research” as Step 1 of doing research. My ideal upskilling culture would have mentees fighting with each other about whose research is best, and mentees would feel free to say in Week 3 “I‘ve decided evals suck, I’m joining Alice’s project on scalable oversight”. (I’m exaggerating slightly.) My experience talking to junior people is often like they think these questions are above their pay-grade or something.
[10%] The field-building grant-makers, for not making “mentees understand what’s going on and have good takes” a core desideratum for the upskilling programs. My impression is grantees focus on easy-to-verify successes like mentees joining full-time prestigious orgs, or conference papers, or impressive mentors.
[25%] The mentees themselves. Come on, guys! They are a little blameless though because they’re understandably paranoid about seeming “productive”.
[5%] RR for not stamping “please read this blog post before doing control research” on their papers. RR is pretty much pareto on (1) doing object level work and (2) explaining why they are doing it & addressing dissenters. The other top orgs seem worse at this.
I agree that Redwood has been historically very good at explaining why they are doing what they are doing. However, I do think that the posts making the case for AI control in particular are getting a bit old, and it would be very good to see updates on them in light of everything that happened in the last two years (e.g. Buck’s recent claim that it’s quite possible that it would have been net negative to implement AI control in the past, because it would have prevented the HF incident).
>Buck’s recent claim that it would have been net negative to implement AI control in the past
Where did Buck claim this? I think he has stated confusion about whether it would have been net negative, but my understanding is that he has not come down decisively.
Sorry, you are right, I misremembered the claim in Alex’s shortform. I’m editing my comment now.
Yep, makes sense. I did not mean to pick on RR specifically, and agree that they are performing way above the field on this.