I think RR are mostly blameless if junior people decided to spend 3 months doing control projects without bothering to read The case for ensuring that powerful AIs are controlled. It’s a 30 min read! You could probably read all the macro-strategy around control in 2-3 days, and maybe a week to absorb it.
The same goes for evals, ambitious mech interp, pragmatic mech interp, scalable oversight, etc. Upskillers working on X should be able to give 5 min answers to questions like: Why did people originally start working on X? Why was X not done before that point? What are the main arguments against X? What’s the crux between supporters of X and supporters of not X? Who are the central figures on both sides? What changes in the strategic landscape make X look better or worse?
I think most of the blame lies with:
[60%] People running upskilling programs, for not fostering a culture where mentees will naturally think “I’m going to spend 3 days reading and thinking about the case for this research” as Step 1 of doing research. My ideal upskilling culture would have mentees fighting with each other about whose research is best, and mentees would feel free to say in Week 3 “I‘ve decided evals suck, I’m joining Alice’s project on scalable oversight”. (I’m exaggerating slightly.) My experience talking to junior people is often like they think these questions are above their pay-grade or something.
[10%] The field-building grant-makers, for not making “mentees understand what’s going on and have good takes” a core desideratum for the upskilling programs. My impression is grantees focus on easy-to-verify successes like mentees joining full-time prestigious orgs, or conference papers, or impressive mentors.
[25%] The mentees themselves. Come on, guys! They are a little blameless though because they’re understandably paranoid about seeming “productive”.
[5%] RR for not stamping “please read this blog post before doing control research” on their papers. RR is pretty much pareto on (1) doing object level work and (2) explaining why they are doing it & addressing dissenters. The other top orgs seem worse at this.
I agree that Redwood has been historically very good at explaining why they are doing what they are doing. However, I do think that the posts making the case for AI control in particular are getting a bit old, and it would be very good to see updates on them in light of everything that happened in the last two years (e.g. Buck’s recent claim that it’s quite possible that it would have been net negative to implement AI control in the past, because it would have prevented the HF incident).
>Buck’s recent claim that it would have been net negative to implement AI control in the past Where did Buck claim this? I think he has stated confusion about whether it would have been net negative, but my understanding is that he has not come down decisively.
I think RR are mostly blameless if junior people decided to spend 3 months doing control projects without bothering to read The case for ensuring that powerful AIs are controlled. It’s a 30 min read! You could probably read all the macro-strategy around control in 2-3 days, and maybe a week to absorb it.
The same goes for evals, ambitious mech interp, pragmatic mech interp, scalable oversight, etc. Upskillers working on X should be able to give 5 min answers to questions like: Why did people originally start working on X? Why was X not done before that point? What are the main arguments against X? What’s the crux between supporters of X and supporters of not X? Who are the central figures on both sides? What changes in the strategic landscape make X look better or worse?
I think most of the blame lies with:
[60%] People running upskilling programs, for not fostering a culture where mentees will naturally think “I’m going to spend 3 days reading and thinking about the case for this research” as Step 1 of doing research. My ideal upskilling culture would have mentees fighting with each other about whose research is best, and mentees would feel free to say in Week 3 “I‘ve decided evals suck, I’m joining Alice’s project on scalable oversight”. (I’m exaggerating slightly.) My experience talking to junior people is often like they think these questions are above their pay-grade or something.
[10%] The field-building grant-makers, for not making “mentees understand what’s going on and have good takes” a core desideratum for the upskilling programs. My impression is grantees focus on easy-to-verify successes like mentees joining full-time prestigious orgs, or conference papers, or impressive mentors.
[25%] The mentees themselves. Come on, guys! They are a little blameless though because they’re understandably paranoid about seeming “productive”.
[5%] RR for not stamping “please read this blog post before doing control research” on their papers. RR is pretty much pareto on (1) doing object level work and (2) explaining why they are doing it & addressing dissenters. The other top orgs seem worse at this.
I agree that Redwood has been historically very good at explaining why they are doing what they are doing. However, I do think that the posts making the case for AI control in particular are getting a bit old, and it would be very good to see updates on them in light of everything that happened in the last two years (e.g. Buck’s recent claim that it’s quite possible that it would have been net negative to implement AI control in the past, because it would have prevented the HF incident).
>Buck’s recent claim that it would have been net negative to implement AI control in the past
Where did Buck claim this? I think he has stated confusion about whether it would have been net negative, but my understanding is that he has not come down decisively.
Sorry, you are right, I misremembered the claim in Alex’s shortform. I’m editing my comment now.
Yep, makes sense. I did not mean to pick on RR specifically, and agree that they are performing way above the field on this.