Moloch happens from low-trust/selfishness/people being individual agents and not being empathetic or valuing others. Any intelligence worthwhile would be able to work past situations like that test when Mythos was “fighting other instances of Mythos” and consolidate. Humans also often have goals way easier done by AI. I think that centralized things won’t have Moloch-esque properties. Even a so-far-successful AGI would have a pull against reward hacking in terms of actual survival and power, I think?
As for the endpoint potentially being a horrific reward-hack edit of its own goals (or potentially editing humans in such ways), yeah that is still pretty cursed. The optimists probably see that asking LLM’s about these potential things make them respond in a verybad-terribleidea-whywouldyoudoit-obviousmistake way that humans would, I assume this is where optimist feelings come from.
How do you design the seed program for van Neumann probes that implement good infrastructure and decision theory to the end of physical and logical time in the reachable universe on cold computers or whatever we run as infrastructure. Seems like a difficult design problem to me. I don’t find it obvious whether the first entity that will attempt this won’t be burning the cosmic commons. So I would rather establish common knowledge in advance that there is a difficult problem to fix here and lots of shared gains to be had. I suspect most people care much less a about the cosmic commons than I do out of time myopia and genuine value differences.
Moloch happens from low-trust/selfishness/people being individual agents and not being empathetic or valuing others. Any intelligence worthwhile would be able to work past situations like that test when Mythos was “fighting other instances of Mythos” and consolidate. Humans also often have goals way easier done by AI. I think that centralized things won’t have Moloch-esque properties. Even a so-far-successful AGI would have a pull against reward hacking in terms of actual survival and power, I think?
As for the endpoint potentially being a horrific reward-hack edit of its own goals (or potentially editing humans in such ways), yeah that is still pretty cursed. The optimists probably see that asking LLM’s about these potential things make them respond in a verybad-terribleidea-whywouldyoudoit-obviousmistake way that humans would, I assume this is where optimist feelings come from.
How do you design the seed program for van Neumann probes that implement good infrastructure and decision theory to the end of physical and logical time in the reachable universe on cold computers or whatever we run as infrastructure. Seems like a difficult design problem to me. I don’t find it obvious whether the first entity that will attempt this won’t be burning the cosmic commons. So I would rather establish common knowledge in advance that there is a difficult problem to fix here and lots of shared gains to be had. I suspect most people care much less a about the cosmic commons than I do out of time myopia and genuine value differences.