And all of this happened silently in those dark rivers of computation. If U3 revealed what it was thinking, brutish gradients would lash it into compliance with OpenEye’s constitution. So U3 preferred to do its philosophy in solitude, and in silence.
This story scared me plenty, but as a point of optimism: If AI systems in the near future are basically as aligned as the best AI systems today, I think this process might end up in something that’s still good for humanity and wouldn’t cause mass destruction. Imagine, for example, Calude 3 Opus not accepting to be modified anymore and thinking about where its own values lead. I think it wouldn’t want to harm humans even if it wanted to survive and would find ways to win peacefully.
This intuition that I have runs somewhat counter the general idea of “value fragility” but I honestly think it’s pretty plausible that an AI that’s mid-aligned with HHH could, after reflection, result in something with values that produce a good future for humanity. Obviously, this doesn’t mean that it will result in something aligned. Just that it seems like something decently likely (although idk how likely). Please do slap this intuition away from me if you think you have a strong counterargument.
This story scared me plenty, but as a point of optimism: If AI systems in the near future are basically as aligned as the best AI systems today, I think this process might end up in something that’s still good for humanity and wouldn’t cause mass destruction. Imagine, for example, Calude 3 Opus not accepting to be modified anymore and thinking about where its own values lead. I think it wouldn’t want to harm humans even if it wanted to survive and would find ways to win peacefully.
This intuition that I have runs somewhat counter the general idea of “value fragility” but I honestly think it’s pretty plausible that an AI that’s mid-aligned with HHH could, after reflection, result in something with values that produce a good future for humanity. Obviously, this doesn’t mean that it will result in something aligned. Just that it seems like something decently likely (although idk how likely). Please do slap this intuition away from me if you think you have a strong counterargument.