IMO we eventually do need to bet on value extrapolation to get the good future (possibly by some system that includes humans).
The question is then how much to empower near-future AIs vs human institutions in the critical period. Maybe somewhere in the middle is good, given how value aligned current AIs seem?
Probably the relevant reference class is more human companies than individual evil humans (assuming AI labs have reasonable internal oversight processes)? And existing human companies seem relatively morally neutral in a way that makes me not want to trust them with infinite cosmic power.
How much moral agency would it be good to give current AIs, ignoring effects on future AIs & also assuming societal buy-in? This should be possible to study by thinking about existing AIs in existing organizations.
Is corrigibility actually an easier alignment target?
I think it would be fine for a democratic process to make this big, hard to reverse decision but it seems more morally dodgy for a developer to do this unilaterally.
I doubt any democratic process will be well-informed & trustworthy (‘actually democratic’ or something) enough to make this decision during the critical period (such that it has an effect on the chance of survival, if this is a good strategy to improve odds of survival). Also, whatever developers do they are empowering someone just by disturbing the status quo.
Some thoughts:
IMO we eventually do need to bet on value extrapolation to get the good future (possibly by some system that includes humans).
The question is then how much to empower near-future AIs vs human institutions in the critical period. Maybe somewhere in the middle is good, given how value aligned current AIs seem?
Probably the relevant reference class is more human companies than individual evil humans (assuming AI labs have reasonable internal oversight processes)? And existing human companies seem relatively morally neutral in a way that makes me not want to trust them with infinite cosmic power.
How much moral agency would it be good to give current AIs, ignoring effects on future AIs & also assuming societal buy-in? This should be possible to study by thinking about existing AIs in existing organizations.
Is corrigibility actually an easier alignment target?
I doubt any democratic process will be well-informed & trustworthy (‘actually democratic’ or something) enough to make this decision during the critical period (such that it has an effect on the chance of survival, if this is a good strategy to improve odds of survival). Also, whatever developers do they are empowering someone just by disturbing the status quo.