… but first we have to get to the part where AIs are inner-aligned to us first.
This is the upshot of the whole series, and the premise of Shear’s work at Softmax. The point is, this hasn’t been achieved, it’s something that needs to be built at a foundational level.
The way I see it, there are two approaches available to us.
Building capacity to fulfil requests into an LLM and then retrofitting alignment guardrails
Creating socialised AI with the primary objective of alignment, then fulfilling requests comes as a matter of course (fulfilling a request is a subset of aligning oneself with the wants of another).
This is where we get to in the series (the final two posts aren’t published yet). I acknowledge it’s speculative, and as you seem to suggest, alignment may be a fool’s errand, but we don’t really have a choice but to entertain possibilities if we’re interested in continued existence (with autonomy).
A note on your original comment: this is not a primer on alignment, it assumes a basic knowledge of the alignment problem, which I covered in the first post of my first series on alignment. I’m assuming readers here at LessWrong understand that the alignment problem is born out of the differences between human sensibilities and machine capabilities.
This is the upshot of the whole series, and the premise of Shear’s work at Softmax. The point is, this hasn’t been achieved, it’s something that needs to be built at a foundational level.
The way I see it, there are two approaches available to us.
Building capacity to fulfil requests into an LLM and then retrofitting alignment guardrails
Creating socialised AI with the primary objective of alignment, then fulfilling requests comes as a matter of course (fulfilling a request is a subset of aligning oneself with the wants of another).
This is where we get to in the series (the final two posts aren’t published yet). I acknowledge it’s speculative, and as you seem to suggest, alignment may be a fool’s errand, but we don’t really have a choice but to entertain possibilities if we’re interested in continued existence (with autonomy).
A note on your original comment: this is not a primer on alignment, it assumes a basic knowledge of the alignment problem, which I covered in the first post of my first series on alignment. I’m assuming readers here at LessWrong understand that the alignment problem is born out of the differences between human sensibilities and machine capabilities.