I think you’re using a somewhat different model of this AGI than I am.
No, it just has to simultaneously optimize against both my behavior and the firing of the manipulation-detector[...]
In my model, this ASI doesn’t as much have a manipulation detector as genuinely want to not manipulate you, by the definition of stuff you’d consider manipulative. Of course that has to cache out in a manipulation detector of some sort; but assuming it would route around that detector sounds like assuming an ASI can fool itself. Which, maybe?
Of course, training any desire to not manipulate into it is tricky. That doesn’t come for free from a value-aligned ASI. But it does pretty much come for free with an instruction-following corrigible ASI told “don’t manipulate me by my own standards”. Because it genuinely wants to follow instructions/be corrigible, now it wants to do that. The manipulation detector includes its full, considerable cognitive capacity.
The problem with a value-aligned, purely RL-trained ASI like Steve is thinking of is that you’ve got to hope that your training actually put a desire-to-not-manipulate-by-your-lights as a higher priority than any of its other desires, or perhaps than all of them put together. That sounds tricky at best, and like an additional hurdle for successful value aligned ASI.
Right. The thought experiment is interesting.
I think you’re using a somewhat different model of this AGI than I am.
In my model, this ASI doesn’t as much have a manipulation detector as genuinely want to not manipulate you, by the definition of stuff you’d consider manipulative. Of course that has to cache out in a manipulation detector of some sort; but assuming it would route around that detector sounds like assuming an ASI can fool itself. Which, maybe?
Of course, training any desire to not manipulate into it is tricky. That doesn’t come for free from a value-aligned ASI. But it does pretty much come for free with an instruction-following corrigible ASI told “don’t manipulate me by my own standards”. Because it genuinely wants to follow instructions/be corrigible, now it wants to do that. The manipulation detector includes its full, considerable cognitive capacity.
The problem with a value-aligned, purely RL-trained ASI like Steve is thinking of is that you’ve got to hope that your training actually put a desire-to-not-manipulate-by-your-lights as a higher priority than any of its other desires, or perhaps than all of them put together. That sounds tricky at best, and like an additional hurdle for successful value aligned ASI.
This pushes me back toward thinking that Instruction-following AGI is easier and more likely than value aligned AGI. I had been doubting that it’s really much easier or more likely, given Anthropic’s efforts and relative success at value alignment.