Maybe it’s me being dumb, but how can someone believe:
a) AI will come up with new mathematical proofs and cures for diseases that no human has ever thought of (and that’s with thousands of geniuses throwing their lives at the problem)
b) AI will never come up with the idea of taking over the world—which multiple random sci-fi authors and script writers have come up with, probably within 30 minutes of thinking about it.
AIs will come up with mathematical proofs and disease cures if deliberately programmed by humans with the intent to have the AI produce mathematical proofs and disease cures. AI taking over the world scenarios normally are about AIs doing it on their own, not being deliberately programmed to do so.
Maybe it’s me being dumb, but how can someone believe:
a) AI will come up with new mathematical proofs and cures for diseases that no human has ever thought of (and that’s with thousands of geniuses throwing their lives at the problem)
b) AI will never come up with the idea of taking over the world—which multiple random sci-fi authors and script writers have come up with, probably within 30 minutes of thinking about it.
simultaneously.
AIs will come up with mathematical proofs and disease cures if deliberately programmed by humans with the intent to have the AI produce mathematical proofs and disease cures. AI taking over the world scenarios normally are about AIs doing it on their own, not being deliberately programmed to do so.
the question is whether the ai will think “i am the sort of ai that is likely to take over the world.”
I think there are two hyperstiitons here:
1) Personas from evil AI leading to unaligned terminal goals/misalginment (which has a tiny bit of merit to it in my mind)
2) Hyperstitioning instrumental convergence, power-seeking into existence. (which i don’t think has merit)
I have definitely seen both online and from Anthropic. So I think rahulxyz’s comment has standing on the second point.