I think the idea of “preserving human agency” in a world with AGI is even less tractable than you’re gesturing at, but in a structural way. Imagine having a trusted translator navigate a bus full of strangers on a trip in Italy. I don’t think there’s a “true” trip that captures all reflexively preferred destinations that can be “discovered” (considering the country is large, and your trip is very short). Now imagine the translator reading all of the passengers’ minds. There’s a lot of path dependence in what the translator prefers and thinks the group values; I picture a space of possible post-AGI institutions or aesthetic preferences which is massive and we’re losing optionality as we move down the tech tree.
The end state is an optimally aligned ASI that can accurately predict and influence our thoughts; there’s nothing like human agency left there. The preferences we baked into the AI at development time are the only thing influencing the future.
I think the idea of “preserving human agency” in a world with AGI is even less tractable than you’re gesturing at, but in a structural way. Imagine having a trusted translator navigate a bus full of strangers on a trip in Italy. I don’t think there’s a “true” trip that captures all reflexively preferred destinations that can be “discovered” (considering the country is large, and your trip is very short). Now imagine the translator reading all of the passengers’ minds. There’s a lot of path dependence in what the translator prefers and thinks the group values; I picture a space of possible post-AGI institutions or aesthetic preferences which is massive and we’re losing optionality as we move down the tech tree.
The end state is an optimally aligned ASI that can accurately predict and influence our thoughts; there’s nothing like human agency left there. The preferences we baked into the AI at development time are the only thing influencing the future.