Rather, it’s that the explicit goal of the alignment community was to differentially advance alignment over capabilities, but instead they ended up advancing capabilities much more effectively than anybody else, while not advancing alignment much. This is a failure to follow the stated goal of colossal proportions, we in fact optimized the opposite of the goal. Why did this happen
My baseline hypothesis is: “because it was much easier to advance capabilities than alignment (especially when capabilities were weak)”. And also: “Drawing more people’s attention to the importance of AGI will inevitably cause some of them to race towards it”. (Either because they weren’t convinced of the safety part, or because they think that ‘advance capabilities to win the race and get more influence later’ is a good strategy for mitigating the safety risks. I think in practice the AI company founders are selected to be a mix of those two.)
This hypothesis seems really important for me, because if it’s true, it’s not clear that mistakes were made. Because it’s not clear that there were alternative feasible routes which would achieved a significantly better ratio of alignment:capabilities progress. (The best way to accelerate alignment progress might inevitably have come with some amount of capabilities progress; and it’s plausible that the latter would always look more impressive in retrospect due to the greater tractability.)
I’m open to there being further interesting things to learn about what caused this, but I think it’s really important to keep this baseline hypothesis in mind and separate out what evidence we have for further mistakes on top of this. Especially if Richard wants to “create common knowledge around something being wrong with the way [alignment leaders] make strategic choices”, it seems important to rule out deflationary hypotheses like this.
My baseline hypothesis is: “because it was much easier to advance capabilities than alignment (especially when capabilities were weak)”. And also: “Drawing more people’s attention to the importance of AGI will inevitably cause some of them to race towards it”. (Either because they weren’t convinced of the safety part, or because they think that ‘advance capabilities to win the race and get more influence later’ is a good strategy for mitigating the safety risks. I think in practice the AI company founders are selected to be a mix of those two.)
This hypothesis seems really important for me, because if it’s true, it’s not clear that mistakes were made. Because it’s not clear that there were alternative feasible routes which would achieved a significantly better ratio of alignment:capabilities progress. (The best way to accelerate alignment progress might inevitably have come with some amount of capabilities progress; and it’s plausible that the latter would always look more impressive in retrospect due to the greater tractability.)
I’m open to there being further interesting things to learn about what caused this, but I think it’s really important to keep this baseline hypothesis in mind and separate out what evidence we have for further mistakes on top of this. Especially if Richard wants to “create common knowledge around something being wrong with the way [alignment leaders] make strategic choices”, it seems important to rule out deflationary hypotheses like this.