To be clear, I’m not taking a strong stance in this sequence on whether AI will go well or badly—that seems up for grabs. My concern is that the alignment community had a plan to make good outcomes more likely (differentially advancing alignment over capabilities) but has mostly pushed the world in the opposite direction, while some parts of it gained a lot of power by doing so.
I should note here that the original goal mentioned in the paper here was to create AIs with desirable goals, and the plan you have outlined is an sub-goal to make the original succeed.
And of course, other community members had other plans of trying to make AIs with desirable goals, though often with some indirection like AI control or automated AI alignment.
The reason I’m bringing it up is that as far as your entire series goes, it seems to rely on the assumption that the AI alignment community had 1 central plan like differentially advancing alignment over capabilities, and pushed the world in the opposite direction instead, when this was not true.
(This is a good example of where the terminal-instrumental goal distinction actually matters, contra your confusion here, and once you do have the distinction, the entire series of posts begins to not work and doesn’t convey the point you wanted to make, if you had a point here.)
To be clear, under some assumptions that some of the community shares, it’s very important to differentially advance alignment over capabilities to prevent negative outcomes, but under other assumptions, this is much less important, and capabilities advancements that have smaller alignment benefits can be positive. In essence, you are treating the community as too much of a monolith.
(This is similar to a point Sarah Constantin expresses here, but said differently/with different implications)
I should note here that the original goal mentioned in the paper here was to create AIs with desirable goals, and the plan you have outlined is an sub-goal to make the original succeed.
And of course, other community members had other plans of trying to make AIs with desirable goals, though often with some indirection like AI control or automated AI alignment.
The reason I’m bringing it up is that as far as your entire series goes, it seems to rely on the assumption that the AI alignment community had 1 central plan like differentially advancing alignment over capabilities, and pushed the world in the opposite direction instead, when this was not true.
(This is a good example of where the terminal-instrumental goal distinction actually matters, contra your confusion here, and once you do have the distinction, the entire series of posts begins to not work and doesn’t convey the point you wanted to make, if you had a point here.)
To be clear, under some assumptions that some of the community shares, it’s very important to differentially advance alignment over capabilities to prevent negative outcomes, but under other assumptions, this is much less important, and capabilities advancements that have smaller alignment benefits can be positive. In essence, you are treating the community as too much of a monolith.
(This is similar to a point Sarah Constantin expresses here, but said differently/with different implications)