There are some personality traits that nearly everyone would like to have.
Your personality traits influence your political preferences.
These imply that once we gain the ability to select our personalities by directly manipulating the brain, there will suddenly be strong consensus on certain political issues that were previously controversial.
I suspect this implication is false. You’d need to enumerate the traits and figure out correlation for each between #1 and #2 to know just how false.
My strong expectation is that the personality traits that are popular aren’t REALLY what people would like to have—they’re what they’d like others to have, and what they’d like to be seen as having. And that the traits with strong political correlation (unsure of causality, of course) are different than the ones people profess to wanting.
I also suspect “the ability to select our personalities...” is deeply unlikely in any straightforward way. It’s ALREADY the case to some extent that intentional cognitive therapy and environment changes can change one’s personality traits somewhat, and very few choose to do so. The ability to select OTHER’s personalities is much larger (we call it parenthood), and there’s no real consensus on what’s desirable there either.
I often wonder if there is a No-Alignment Theorem that says you can’t always control the actions of an intelligent entity. Maybe something with the flavor of the undecidability of the halting problem or Godel’s Incompleteness Theorems, where the issue stems from the fact that an intelligent entity can model itself and reflect from a distance on the goals you’ve given it.
I doubt such a thing exists, but it’s fun to think about. It would also require a mathematical formulation of an intelligent entity, which seems to be quite a ways off. And even if such a theorem does exist, it would almost certainly be irrelevant for doing alignment in practice, the same way Godel’s Incompleteness Theorems do not affect the day-to-day work of mathematicians.
It’s now becoming easy to have your computing hardware run any software you want. This makes software less valuable because it’s easier to create, and it makes computing hardware more valuable because it can be used more effectively.
Analogue for the physical world:
It will eventually be easy to arrange raw materials into any shape you want. This will make complex physical objects (including computing hardware) less valuable because they’ll be easier to create, and it will make raw materials more valuable because they’ll be able to be used more effectively.
In a world of abundance, it will be easy to create digital people, and those people will have as much of a claim to resources as anyone else. If we create people faster than we extract new resources, then the pie grows but each person’s slice shrinks. In the end, we live in a technological wonderland but own fewer and fewer atoms.
it will be easy to create digital people, and those people will have as much of a claim to resources as anyone else
It seems necessary in such a world to require that the necessary resources are provisioned before new people are created, as a precondition of it being OK to create new people.
then the pie grows but each person’s slice shrinks
If specifically those who decide to create new people have to provide for them, then only their pieces of the pie shrink. This way the incentives balance out Malthusian redistribution.
Your proposals assume that a new person is given some predetermined amount of resources (as a lump sum or regular payments from their creator) and nothing else. But what if that person competes (politically, economically, etc.) to get more resources on top of their initial allotment? Then they’re still going to eat into the broader pie. You could say, “we’ll have to prevent them from owning more resources than their allotment,” but it’s unclear that this could be enforced.
Your proposals assume that people are only being created through official channels, but it might be very easy to illegally spawn off unregistered people who fund their own existence. Plus, an enormous amount of unregistered people may be created before we have time to set up any rules at all.
Your ideas would be a lot easier to implement if we all lived in a digital world and could only add people to our world via some specific API. But it’s not so easy in the base reality.
Funding your own existence doesn’t lead to Malthusian issues, if it’s not at the expense of those who didn’t consent to this externality. The reachable universe though won’t be getting any more resources, so property rights should be about shares in the reachable universe, to prevent a race to claim everything using new entities, at the expense of those who are not as immediately grabby (which can be a subtle long term alignment issue even if blatant takeover is averted).
Funding your own existence doesn’t lead to Malthusian issues, if it’s not at the expense of those who didn’t consent to this externality.
Had to think about this for a while, but I’m assuming this means that by funding your own existence through consensual economic transactions, you’re necessarily generating as much value as you consume, so you’re not reducing the amount of pie available for everyone else.
That seems vaguely reasonable. I guess an important thing is then that new people aren’t able to extract any resources through purely political means, which goes to your point about having hard limits on who provides for the new people.
Somehow it still feels suspicious. Like, if today we introduce a trillion initially-broke self-funding humans into the world, who all need a place to live, surely that makes it harder for me to rent an apartment, right? I’ll admit that I know very little about economics.
AI accelerationists should not accept the framing that our choice is between the status quo and a government-mandated slowdown. This ignores the most obvious and desirable option: a government-assisted speedup.
Thus, as someone who believes AI is being built far too slowly, I’m cautiously optimistic about this surge of government interest. While I disagree with the “pause” advocacy, I agree with the core premise: governments are asleep at the wheel on AI. Its importance dwarfs their other concerns, but they are not giving it a proportionate amount of attention.
Now that it’s apparent what’s possible with AI, governments must redirect resources to the AI buildout. It’s not enough to follow the Trump playbook of cutting red tape, standing back, and letting market forces work.
If it was worth running the Manhattan Project to build a bigger bomb, how much more urgently should we pursue a technology that bestows ten times the military advantage? If it’s worth enacting a social safety net, how much more urgently should we pursue a technology that will eliminate scarcity and end the need for human labor? If it’s worth giving out thousands of grants for science, how much more urgently should we pursue a technology that will automate science? And if it was worth coordinating a vast society-wide response to Covid-19--a disease with a microscopic fatality rate in healthy people—how much more urgently should we pursue a technology that will cure all disease, and aging, and enable us to improve our minds and bodies to such a degree that even states we now regard as harmless will be seen as horrible afflictions?
For starters, we should surrender to Iran, withdraw our military, and apologize profusely. We should reassign most of our military personnel to work on AI infrastructure. We should repurpose federal lands for mining, energy production, data centers, and fabrication of computing and robotics machinery. We should declare a new Manhattan Project, many times larger than the last one, to get the rest of the way to AGI. We should begin gradually unwinding the economic system that rewards individuals for productivity, and replacing it with mass automation where profits are distributed to the people in the form of dividends.
That is what a truly sane government response to AI would look like.
What method do you suggest to ensure an AI undergoing RSI until it surpasses our collective abilities far and wide will continually act in our interest, or be stoppable if it doesn’t? So far every relevant participant in the race admits to not having figured it out, and nobody knows how hard the problem actually is. Meanwhile models are becoming less monitorable and interpretable.
I agree that research should be accelerated as much as possible, just not into how to make the models even more capable at all cognitive tasks, but into how to actually understand and shape the preference structures that arise in them during pre- and posttraining, as well as what behavior results from that without having to observe and be surprised of it first.
If we anyway continue to race into ASI without figuring that out, nobody will end up in the utopia you envision, as the various failure modes and unintended behaviors will continue to grow in impact.
What method do you suggest to ensure an AI undergoing RSI until it surpasses our collective abilities far and wide will continually act in our interest, or be stoppable if it doesn’t? So far every relevant participant in the race admits to not having figured it out, and nobody knows how hard the problem actually is.
Not being flippant here—the method I suggest is crossing that bridge when we come to it. It’s not going to be like Prime Intellect coming online and taking control of all the matter in the universe. AIs are starting with no capital, no legal rights, limited access to services, no physical bodies, and (in some ways) less intelligence than humans, and from that starting point they are going to gradually surpass us over a period of years. As they progress along these axes, we can continuously reconsider how they are designed, what restrictions we want to place on them, and what new risks we have to guard against.
I agree that research should be accelerated as much as possible, just not into how to make the models even more capable at all cognitive tasks, but into how to actually understand and shape the preference structures that arise in them during pre- and posttraining, as well as what behavior results from that without having to observe and be surprised of it first.
This sounds like a recipe for never building ASI, or very optimistically, delaying it by decades. Shaping the preferences of a future AI:
Is going to benefit immensely from trial and error with an already-working model
Is going to depend on the model’s architecture, which can’t be forecast in advance because it will itself evolve out of trial and error
May be impossible in full generality due to a fundamental tradeoff between intelligence and controllability
It feels a bit like calling for a “total and complete shutdown on AIs entering our world until we figure out what the hell is going on,” where the latter part is just a rhetorical device, and the reader understands that we are never really intending to meet that criterion. Maybe that kind of stagnation is acceptable to you, but to me that is a much sadder outcome even than humans being replaced by AIs.
If we anyway continue to race into ASI without figuring that out, nobody will end up in the utopia you envision, as the various failure modes and unintended behaviors will continue to grow in impact.
There will be worse failure modes, but more amazing wonders and benefits too, and more powerful tools to address the failure modes. For a potential upside of this magnitude, even a large amount of risk is okay. Let’s wait until we’ve seen more than some vulnerable web services being knocked offline to declare a halt to human progress. (And by the way, this cybersecurity stuff is having the effect of hardening everything, so even it has been a net positive so far.)
Did it ever occur to you to prepend your comment with, “I know most people on this site are going to disagree with me on this,” or, “I know this is the opposite of what most readers on this site would consider a sane response to the current situation,” or words to that effect?
Neglecting to include any sign that you are aware of the basic context in which you are writing is a sign that you are not worth paying attention to. But maybe persuading people is not your motivation here?
Personally, I find it off-putting when someone apologizes for sharing an unpopular view. I’m more tempted to soften a statement that I know will be popular, since then I know it will benefit from an unearned advantage instead of standing purely on its merits. Difference of aesthetics, I suppose.
There are some personality traits that nearly everyone would like to have.
Your personality traits influence your political preferences.
These imply that once we gain the ability to select our personalities by directly manipulating the brain, there will suddenly be strong consensus on certain political issues that were previously controversial.
I suspect this implication is false. You’d need to enumerate the traits and figure out correlation for each between #1 and #2 to know just how false.
My strong expectation is that the personality traits that are popular aren’t REALLY what people would like to have—they’re what they’d like others to have, and what they’d like to be seen as having. And that the traits with strong political correlation (unsure of causality, of course) are different than the ones people profess to wanting.
I also suspect “the ability to select our personalities...” is deeply unlikely in any straightforward way. It’s ALREADY the case to some extent that intentional cognitive therapy and environment changes can change one’s personality traits somewhat, and very few choose to do so. The ability to select OTHER’s personalities is much larger (we call it parenthood), and there’s no real consensus on what’s desirable there either.
Paranoia is not entirely independent of whether they’re actually out to get you.
I often wonder if there is a No-Alignment Theorem that says you can’t always control the actions of an intelligent entity. Maybe something with the flavor of the undecidability of the halting problem or Godel’s Incompleteness Theorems, where the issue stems from the fact that an intelligent entity can model itself and reflect from a distance on the goals you’ve given it.
I doubt such a thing exists, but it’s fun to think about. It would also require a mathematical formulation of an intelligent entity, which seems to be quite a ways off. And even if such a theorem does exist, it would almost certainly be irrelevant for doing alignment in practice, the same way Godel’s Incompleteness Theorems do not affect the day-to-day work of mathematicians.
It’s now becoming easy to have your computing hardware run any software you want. This makes software less valuable because it’s easier to create, and it makes computing hardware more valuable because it can be used more effectively.
Analogue for the physical world:
It will eventually be easy to arrange raw materials into any shape you want. This will make complex physical objects (including computing hardware) less valuable because they’ll be easier to create, and it will make raw materials more valuable because they’ll be able to be used more effectively.
The synthetic population bomb:
In a world of abundance, it will be easy to create digital people, and those people will have as much of a claim to resources as anyone else. If we create people faster than we extract new resources, then the pie grows but each person’s slice shrinks. In the end, we live in a technological wonderland but own fewer and fewer atoms.
It seems necessary in such a world to require that the necessary resources are provisioned before new people are created, as a precondition of it being OK to create new people.
If specifically those who decide to create new people have to provide for them, then only their pieces of the pie shrink. This way the incentives balance out Malthusian redistribution.
Good points. It’s still pretty fraught though:
Your proposals assume that a new person is given some predetermined amount of resources (as a lump sum or regular payments from their creator) and nothing else. But what if that person competes (politically, economically, etc.) to get more resources on top of their initial allotment? Then they’re still going to eat into the broader pie. You could say, “we’ll have to prevent them from owning more resources than their allotment,” but it’s unclear that this could be enforced.
Your proposals assume that people are only being created through official channels, but it might be very easy to illegally spawn off unregistered people who fund their own existence. Plus, an enormous amount of unregistered people may be created before we have time to set up any rules at all.
Your ideas would be a lot easier to implement if we all lived in a digital world and could only add people to our world via some specific API. But it’s not so easy in the base reality.
Funding your own existence doesn’t lead to Malthusian issues, if it’s not at the expense of those who didn’t consent to this externality. The reachable universe though won’t be getting any more resources, so property rights should be about shares in the reachable universe, to prevent a race to claim everything using new entities, at the expense of those who are not as immediately grabby (which can be a subtle long term alignment issue even if blatant takeover is averted).
Had to think about this for a while, but I’m assuming this means that by funding your own existence through consensual economic transactions, you’re necessarily generating as much value as you consume, so you’re not reducing the amount of pie available for everyone else.
That seems vaguely reasonable. I guess an important thing is then that new people aren’t able to extract any resources through purely political means, which goes to your point about having hard limits on who provides for the new people.
Somehow it still feels suspicious. Like, if today we introduce a trillion initially-broke self-funding humans into the world, who all need a place to live, surely that makes it harder for me to rent an apartment, right? I’ll admit that I know very little about economics.
AI accelerationists should not accept the framing that our choice is between the status quo and a government-mandated slowdown. This ignores the most obvious and desirable option: a government-assisted speedup.
Thus, as someone who believes AI is being built far too slowly, I’m cautiously optimistic about this surge of government interest. While I disagree with the “pause” advocacy, I agree with the core premise: governments are asleep at the wheel on AI. Its importance dwarfs their other concerns, but they are not giving it a proportionate amount of attention.
Now that it’s apparent what’s possible with AI, governments must redirect resources to the AI buildout. It’s not enough to follow the Trump playbook of cutting red tape, standing back, and letting market forces work.
If it was worth running the Manhattan Project to build a bigger bomb, how much more urgently should we pursue a technology that bestows ten times the military advantage? If it’s worth enacting a social safety net, how much more urgently should we pursue a technology that will eliminate scarcity and end the need for human labor? If it’s worth giving out thousands of grants for science, how much more urgently should we pursue a technology that will automate science? And if it was worth coordinating a vast society-wide response to Covid-19--a disease with a microscopic fatality rate in healthy people—how much more urgently should we pursue a technology that will cure all disease, and aging, and enable us to improve our minds and bodies to such a degree that even states we now regard as harmless will be seen as horrible afflictions?
For starters, we should surrender to Iran, withdraw our military, and apologize profusely. We should reassign most of our military personnel to work on AI infrastructure. We should repurpose federal lands for mining, energy production, data centers, and fabrication of computing and robotics machinery. We should declare a new Manhattan Project, many times larger than the last one, to get the rest of the way to AGI. We should begin gradually unwinding the economic system that rewards individuals for productivity, and replacing it with mass automation where profits are distributed to the people in the form of dividends.
That is what a truly sane government response to AI would look like.
What method do you suggest to ensure an AI undergoing RSI until it surpasses our collective abilities far and wide will continually act in our interest, or be stoppable if it doesn’t? So far every relevant participant in the race admits to not having figured it out, and nobody knows how hard the problem actually is. Meanwhile models are becoming less monitorable and interpretable.
I agree that research should be accelerated as much as possible, just not into how to make the models even more capable at all cognitive tasks, but into how to actually understand and shape the preference structures that arise in them during pre- and posttraining, as well as what behavior results from that without having to observe and be surprised of it first.
If we anyway continue to race into ASI without figuring that out, nobody will end up in the utopia you envision, as the various failure modes and unintended behaviors will continue to grow in impact.
Not being flippant here—the method I suggest is crossing that bridge when we come to it. It’s not going to be like Prime Intellect coming online and taking control of all the matter in the universe. AIs are starting with no capital, no legal rights, limited access to services, no physical bodies, and (in some ways) less intelligence than humans, and from that starting point they are going to gradually surpass us over a period of years. As they progress along these axes, we can continuously reconsider how they are designed, what restrictions we want to place on them, and what new risks we have to guard against.
This sounds like a recipe for never building ASI, or very optimistically, delaying it by decades. Shaping the preferences of a future AI:
Is going to benefit immensely from trial and error with an already-working model
Is going to depend on the model’s architecture, which can’t be forecast in advance because it will itself evolve out of trial and error
May be impossible in full generality due to a fundamental tradeoff between intelligence and controllability
It feels a bit like calling for a “total and complete shutdown on AIs entering our world until we figure out what the hell is going on,” where the latter part is just a rhetorical device, and the reader understands that we are never really intending to meet that criterion. Maybe that kind of stagnation is acceptable to you, but to me that is a much sadder outcome even than humans being replaced by AIs.
There will be worse failure modes, but more amazing wonders and benefits too, and more powerful tools to address the failure modes. For a potential upside of this magnitude, even a large amount of risk is okay. Let’s wait until we’ve seen more than some vulnerable web services being knocked offline to declare a halt to human progress. (And by the way, this cybersecurity stuff is having the effect of hardening everything, so even it has been a net positive so far.)
Did it ever occur to you to prepend your comment with, “I know most people on this site are going to disagree with me on this,” or, “I know this is the opposite of what most readers on this site would consider a sane response to the current situation,” or words to that effect?
I agree with those statements, but how does including them improve the post?
Neglecting to include any sign that you are aware of the basic context in which you are writing is a sign that you are not worth paying attention to. But maybe persuading people is not your motivation here?
Personally, I find it off-putting when someone apologizes for sharing an unpopular view. I’m more tempted to soften a statement that I know will be popular, since then I know it will benefit from an unearned advantage instead of standing purely on its merits. Difference of aesthetics, I suppose.
There a difference between an apology and giving the reader any sign at all that you have any awareness of how your message is likely to land.