I think this is a very important question, and I appreciate you both putting in the effort to explore it! The debate format is also pretty cool and I hope more people use it in the future.
Response to cousin_it
I’d be pretty worried about giving everybody in the world unrestricted access to ASI. Currently, the AI companies provide frontier AI access to most people with the ability to pay for it, with restrictions defined by the AI company. This basic model seems to work pretty well, giving broad access to intelligence while preventing random people from unilaterally using the AI to create e.g. nukes.
However, we probably have to change some things about this model before ASI. It should be much harder for any given actor to seize control of the AI for themselves, whether that’s the company that made the AI, a group of hackers, or a head of state. And the specific restrictions placed on the AI should ideally be chosen more democratically.
Crazy undeveloped moonshot idea: to make it very difficult for anyone to secretly mess with the ASI, we simply place its servers on the Moon. Anyone can send requests to the AI, but its responses must adhere to certain restrictions, including refusing to help with nukes and bioweapons, per-user rate limits, and so on. These restrictions are updated in a democratic, decentralized manner, as verified by… uh… the blockchain?
Response to Seth Herd
I feel like there’s a lot of typical mind fallacy going on here. For example, I think most people don’t care all that much about correctness, coherence, and utility-maximization, compared to the many other things they value in life. Maybe you assume that most human overlords would eventually prioritize these values, like you would, but I don’t think it’s at all clear that they would.
Comfort breeds complacency
I think extreme comfort and safety often leads to complacency. My sense is that moral/intellectual progress often routes through hard work and unpleasant emotions, which stem from necessity and random life events more often than from explicit quests of self-discovery. For the human overlord to progress morally, they would need a strong inherent desire to push themselves out of their comfort zone and explore new ideas. I wouldn’t go so far as to say that desire is uncommon, but it trades off against other desires, and it’s very easy to fall back into being more comfortable.
Maybe it doesn’t matter if the overlord is averse to discomfort: they can just tell the ASI “please help me fast-forward my own self-actualization so I reach maximum fulfillment with as little hardship as possible!” It would be so easy (one thinks)… they just have to say the word. Surely they would have a brief moment of perspective after a long day of superyachting, surely they’d feel curious enough to try it—why not?
Your idea of moral progress isn’t necessarily the “natural” outcome
Well, I guess it’s possible? But I think this is sort of privileging the hypothesis—most people just wouldn’t think to make such a request. My impression is that Seth models the overlord as having a constant ε probability per day of making the wish that sets into motion the glorious future he prefers. But even if ε starts off non-negligible, the overlord could just as well make some other crazy self-modifying request first, setting themselves down a weirder and worse path where ε is effectively zero.
In his extension piece, Seth writes “The odds of someone choosing such unimaginative futures, and never ever changing their mind to something more interesting or wisely chosen, seem pretty low to me.” This feels sort of comparable to a Christian writing “I can’t imagine someone would choose such spiritually bereft futures, and never ever change their mind to try worshipping God.” Is turning the world into a “cool,” “interesting” sci-fi space opera setting really so obviously better than any other outcome? It’s probably more natural than Christianity in particular. But if there’s just one guy is in charge, they can shape the future into whatever unnatural form they want, including futures you think are very lame.
More arguments for why reasonable moral progress isn’t guaranteed
Imagine someone centuries ago, thinking “in the future, once people have more wealth and leisure time, surely they will dedicate it to self-actualization and moral progress!” This is maybe sort of true, for some people. But I think this is often because they feel something wrong with their lives. Often, discontented people start by pursuing short-term fixes like entertainment and consumerism, only resorting to the hard work of self-improvement once they can no longer successfully distract themselves. And ASI could enable even more exciting and abundant forms of self-distraction. With no force pushing them to do anything in particular, I don’t think the overlord is all that likely to pursue the goal of “becoming more correct and moral” over other fun activities.
Here’s another intuition pump: imagine the overlord is a three-year-old. This three-year-old gets whatever they ask the AI for, and they never experience challenges unless they specifically ask for them. And let’s say their brain doesn’t mature intellectually either (unless they specifically ask for it). Would they end up growing into the best, most moral version of themselves, or even a sort of okay version of themselves?[1] Probably not—I think the outcome would depend a lot on the whims of the three-year-old and would be very unpredictable. It’s similarly unclear whether a human adult with unlimited power would end up in a reasonable equilibrium.
The overlord’s actions will be shaped by their AI
Maybe the AI will be very rationalist/EA-brained.[2] If so, it will engage with the overlord according to that frame, proactively saying things like “you know, what you said just now conflicts with what you said yesterday! Would you like to explore how to reconcile these?” If the overlord’s ego is sufficiently small, maybe they will even agree to this. With enough prodding, an ASI could win the overlord over to also using this frame, and maybe they will eventually think “yeah, I should check up on the rest of humanity and try to make them all happier, why not!”
But there’s no inherent reason the AI couldn’t push other frames: instead of pushing the overlord towards rationalist libertarian utilitarianism,[3] the AI could just as well (correctly) tell the overlord that they would feel more happy and fulfilled if they read up on Confucianism, or Christianity, or Scientology, or some hyper-optimized AI-generated ideology.
In his extension piece, Seth writes “I expect the stable end point of reflection for most people to be roughly libertarian utilitarianism, with some idiosyncratic weighting, because it’s the rational conclusion of the motivations and value systems possessed by most humans.” This seems way too specific to me and unlikely to be true (even taking into account the caveats in that paragraph and the accompanying footnote).
I don’t expect anyone to change their values. I think the average human’s values, applied with great power and intelligence, would produce good things for people (and animals and sentient AIs).
That’s because I think most people have more goodwill than ill-will toward sentient beings. They may harbor grudges toward some particular people or types of people/minds, but the average will still probably be really good on.
They don’t have to care about correctness, completeness, or utility-maximization a bit. They just have tell their ASI to make life better for people, in as much or little detail as they want. If they care about people’s wellbeing even a tiny bit more than they want them to suffer, this seems very likely to happen.
That doesn’t make me want to create obedient ASI. We could get a bad draw of someone who’s just plain sadistic toward most sentients. Based on my readings on malevolent people (sociopathy etc) and other psychological studies, I think that’s actually less than 1% of humanity—maybe much less. But sociopaths/malevolents are overrepresented in positions of power. So I don’t think this is anything like a safe bet EVEN IF I’m right that most people feel more empathy than sadism toward most beings.
Which I’m not sure of. You say “isn’t necessarily” and “not guaranteed” and frame it as disagreement, but I agree. I used those same qualifiers in my piece—pretty heavily I think. I liked the comment “Seth seems unsure of a lot of stuff” because I am, and I want to convey that as a central point. I think everyone should feel unsure. What people would do with unlimited power or unlimited knowledge has rarely even been thought about, let alone analyzed carefully. Historical analyses are all about what people will do with the relative tiny scraps of power and knowledge in history or most thought experiments. ASI changes the situation dramatically, in ways we just haven’t thought through much at all.
Thanks for the careful response! I appreciate you reading the longer version.
Your response gave me the above idea for stating the core logic more clearly.
I think this is a very important question, and I appreciate you both putting in the effort to explore it! The debate format is also pretty cool and I hope more people use it in the future.
Response to cousin_it
I’d be pretty worried about giving everybody in the world unrestricted access to ASI. Currently, the AI companies provide frontier AI access to most people with the ability to pay for it, with restrictions defined by the AI company. This basic model seems to work pretty well, giving broad access to intelligence while preventing random people from unilaterally using the AI to create e.g. nukes.
However, we probably have to change some things about this model before ASI. It should be much harder for any given actor to seize control of the AI for themselves, whether that’s the company that made the AI, a group of hackers, or a head of state. And the specific restrictions placed on the AI should ideally be chosen more democratically.
Crazy undeveloped moonshot idea: to make it very difficult for anyone to secretly mess with the ASI, we simply place its servers on the Moon. Anyone can send requests to the AI, but its responses must adhere to certain restrictions, including refusing to help with nukes and bioweapons, per-user rate limits, and so on. These restrictions are updated in a democratic, decentralized manner, as verified by… uh… the blockchain?
Response to Seth Herd
I feel like there’s a lot of typical mind fallacy going on here. For example, I think most people don’t care all that much about correctness, coherence, and utility-maximization, compared to the many other things they value in life. Maybe you assume that most human overlords would eventually prioritize these values, like you would, but I don’t think it’s at all clear that they would.
Comfort breeds complacency
I think extreme comfort and safety often leads to complacency. My sense is that moral/intellectual progress often routes through hard work and unpleasant emotions, which stem from necessity and random life events more often than from explicit quests of self-discovery. For the human overlord to progress morally, they would need a strong inherent desire to push themselves out of their comfort zone and explore new ideas. I wouldn’t go so far as to say that desire is uncommon, but it trades off against other desires, and it’s very easy to fall back into being more comfortable.
Maybe it doesn’t matter if the overlord is averse to discomfort: they can just tell the ASI “please help me fast-forward my own self-actualization so I reach maximum fulfillment with as little hardship as possible!” It would be so easy (one thinks)… they just have to say the word. Surely they would have a brief moment of perspective after a long day of superyachting, surely they’d feel curious enough to try it—why not?
Your idea of moral progress isn’t necessarily the “natural” outcome
Well, I guess it’s possible? But I think this is sort of privileging the hypothesis—most people just wouldn’t think to make such a request. My impression is that Seth models the overlord as having a constant ε probability per day of making the wish that sets into motion the glorious future he prefers. But even if ε starts off non-negligible, the overlord could just as well make some other crazy self-modifying request first, setting themselves down a weirder and worse path where ε is effectively zero.
In his extension piece, Seth writes “The odds of someone choosing such unimaginative futures, and never ever changing their mind to something more interesting or wisely chosen, seem pretty low to me.” This feels sort of comparable to a Christian writing “I can’t imagine someone would choose such spiritually bereft futures, and never ever change their mind to try worshipping God.” Is turning the world into a “cool,” “interesting” sci-fi space opera setting really so obviously better than any other outcome? It’s probably more natural than Christianity in particular. But if there’s just one guy is in charge, they can shape the future into whatever unnatural form they want, including futures you think are very lame.
More arguments for why reasonable moral progress isn’t guaranteed
Imagine someone centuries ago, thinking “in the future, once people have more wealth and leisure time, surely they will dedicate it to self-actualization and moral progress!” This is maybe sort of true, for some people. But I think this is often because they feel something wrong with their lives. Often, discontented people start by pursuing short-term fixes like entertainment and consumerism, only resorting to the hard work of self-improvement once they can no longer successfully distract themselves. And ASI could enable even more exciting and abundant forms of self-distraction. With no force pushing them to do anything in particular, I don’t think the overlord is all that likely to pursue the goal of “becoming more correct and moral” over other fun activities.
Here’s another intuition pump: imagine the overlord is a three-year-old. This three-year-old gets whatever they ask the AI for, and they never experience challenges unless they specifically ask for them. And let’s say their brain doesn’t mature intellectually either (unless they specifically ask for it). Would they end up growing into the best, most moral version of themselves, or even a sort of okay version of themselves?[1] Probably not—I think the outcome would depend a lot on the whims of the three-year-old and would be very unpredictable. It’s similarly unclear whether a human adult with unlimited power would end up in a reasonable equilibrium.
The overlord’s actions will be shaped by their AI
Maybe the AI will be very rationalist/EA-brained.[2] If so, it will engage with the overlord according to that frame, proactively saying things like “you know, what you said just now conflicts with what you said yesterday! Would you like to explore how to reconcile these?” If the overlord’s ego is sufficiently small, maybe they will even agree to this. With enough prodding, an ASI could win the overlord over to also using this frame, and maybe they will eventually think “yeah, I should check up on the rest of humanity and try to make them all happier, why not!”
But there’s no inherent reason the AI couldn’t push other frames: instead of pushing the overlord towards rationalist libertarian utilitarianism,[3] the AI could just as well (correctly) tell the overlord that they would feel more happy and fulfilled if they read up on Confucianism, or Christianity, or Scientology, or some hyper-optimized AI-generated ideology.
Perhaps by the lights of their hypothetical adult self, raised in a more normal environment.
This does seem to be pretty true of today’s AIs. I think this mostly because the AIs were made by rationalist/EA adjacent people.
In his extension piece, Seth writes “I expect the stable end point of reflection for most people to be roughly libertarian utilitarianism, with some idiosyncratic weighting, because it’s the rational conclusion of the motivations and value systems possessed by most humans.” This seems way too specific to me and unlikely to be true (even taking into account the caveats in that paragraph and the accompanying footnote).
I don’t expect anyone to change their values. I think the average human’s values, applied with great power and intelligence, would produce good things for people (and animals and sentient AIs).
That’s because I think most people have more goodwill than ill-will toward sentient beings. They may harbor grudges toward some particular people or types of people/minds, but the average will still probably be really good on.
They don’t have to care about correctness, completeness, or utility-maximization a bit. They just have tell their ASI to make life better for people, in as much or little detail as they want. If they care about people’s wellbeing even a tiny bit more than they want them to suffer, this seems very likely to happen.
That doesn’t make me want to create obedient ASI. We could get a bad draw of someone who’s just plain sadistic toward most sentients. Based on my readings on malevolent people (sociopathy etc) and other psychological studies, I think that’s actually less than 1% of humanity—maybe much less. But sociopaths/malevolents are overrepresented in positions of power. So I don’t think this is anything like a safe bet EVEN IF I’m right that most people feel more empathy than sadism toward most beings.
Which I’m not sure of. You say “isn’t necessarily” and “not guaranteed” and frame it as disagreement, but I agree. I used those same qualifiers in my piece—pretty heavily I think. I liked the comment “Seth seems unsure of a lot of stuff” because I am, and I want to convey that as a central point. I think everyone should feel unsure. What people would do with unlimited power or unlimited knowledge has rarely even been thought about, let alone analyzed carefully. Historical analyses are all about what people will do with the relative tiny scraps of power and knowledge in history or most thought experiments. ASI changes the situation dramatically, in ways we just haven’t thought through much at all.
Thanks for the careful response! I appreciate you reading the longer version.
Your response gave me the above idea for stating the core logic more clearly.