Noting that this doesn’t match my experience. Most of my mathematician friends are very much feeling the approaching end of mathematics, and a number of them got quite scared of the Hugging Face incident too.
David Matolcsi
I agree that we should act like if we were in base reality or in an acausal trade sim, but I think the logic here doesn’t really go through.
First of all, there might be universes where there is only on civilization, and that civilization is not interested in acausal trade but interested in running sims for entertainment. Then there is no one to buy data from them.
Second, Claude estimates that approximately 2% of glabal GDP is spent on facilitating international trade, while 0.25% is spent on making movies, tv series and video games, and maybe 10-20% of that is historical stories or fantasy.
Of course it’s dangerous to generalize from these numbers to the far future, but I think it’s not crazy to imagine that entertainment sims are not entirely negligible compared to sims created to facilitate inter-universal trade.And I don’t buy the argument that the acausal trade folks would always want to buy the results of the entertainment sims. Professional historians don’t pay Ridley Scott for the data on how the simulated battle goes down in Gladiator.
I don’t think we would notice if we lived in a world that is created for entertainment purposes and is about as realistic as Gladiator.
My current actions have different amount of effect on different distributions. The policy “get a sword if you have seen a lot of dragons, but don’t get one if you haven’t” is a good policy that maximizes utility if I care both about a distribution of worlds with a lot of dragons, and a distribution without them. So I follow this policy, and updating on my observations (no dragons) means following one branch of this policy.
I guess I got a bit carried away with poetic phrasing in that sentence, but I kind of endorse it. I’m not a pure hedonic utilitarian, I value multiple things, and truth and beauty are part of them. I have less uncertainty about the values in my utility function than about the weights though. Like people feeling happiness in the highly weighted moments certainly feels like an important part of what I’m optimizing for, and I feel there are not that many things as important as happiness that I want to put in my utility function. For example, I wouldn’t put dramaticness very high among the values I want to maximize.
On the other hand, I feel that mathematical simplicity is such an arbitrary concept to use as weighting. Why that? I don’t feel I have strong intuitions on which moments are more important than others (given that “every moment is equally important” is not an option in infinite multiverses), so I have a lot of approximately equally endorsed random ideas on what could be a weighting that makes some sense. Mathematical simplicity is one of them, but so is dramaticness and probably thousands of other things.
I think this was the biggest hole in the Preferences without Existence post (which I maybe consider the previous best attempt at metaphysics from my perspective) that it claims it as a principle of moral philosophy that you should care about moments weighted by mathematical simplicity, even though mathematical simplicity feels like a not very natural thing to be so fundamental in my moral philosophy. So here I attempt to explain that I think it’s more natural to start from many possible weightings, and then you can still use your updates from observing the world to get back to primarily concentrating on the mathematical simplicity flavored weightings.
Caring about happy experiences is very normal as part of your utility function, but not as part of your weighting.
I care about summing up the happiness of of all moments, weighted by how simple places in how simple universes they occur in—that’s pretty normal, you can hope that you will get back something like materialist utilitarianism from this.
I care about summing up the happiness of all moments, weighted by the ranking of the happiness itself—that’s some wacky stuff!
The best way I can make sense of putting happiness in my weightings (and not just in my utility function) is saying “I’m Gottfried Leibniz, and I believe we should make things better conditional on the assumption that we already live in the best of all logically possible worlds.” I think that’s what corresponds to saying that you put happiness in your weighting, ie you care more about the happiness of the already happiest places.
But I don’t think that the Leibniz-style weighting by happiness is a priori crazier than weighting by mathematical simplicity. This is where I disagree with Scott Garrabrant’s Preferences without Existence post—he assumes a priori without explanation that the weights need to be based on mathematical simplicity. I don’t love math that much, so I’m a priori considering some other weightings too, and I only get back to a ~materialist worldview by updating on my observations that the laws of physics look simple.
Yes, I think you are understanding the Boltzmann brain thing correctly. Originally, I planned to write the idea as if each distribution was represented by a person in a moral parliament, and they could bargain, form coalitions, and win and lose bets in influence-coins against each other. But I couldn’t find good thought experiments where the simpler idea of adding up the values in bounded utilities from each distribution fails, so I didn’t need to invoke the parliamentary vision yet.
I somewhat stand behind the “mystical” name: I feel something spooky is going in on those distributions. Significantly affecting those distributions routes through things like “even though the real weighting in the distribution doesn’t have much to do with mathematical simplicity, for some reason being in this wold with seemingly mathematically simple physics is an unusually great place to influence the distribution”. I think the best explanation for that often routes through something mystical, like the existence of gods.
Altogether, I think it’s maybe helpful to think of the “Updates of mathematical and mystical” section as a final step in the journey of trying to answer the question “what’s the probability that God exists?”
In the first post, I needed to discard the simple notion of probabilities for questions like this. But it’s a cop-out to just say “undefined”, I should still try to answer in the spirit of the question. So we get to something like how much influence the possibility of God’s existence has on my actions. I can try to answer that in the UDASSA framing, or in the Preferences without existence style non-realist UDASSA framing, but then that’s cheating, because it presupposes a mathematical prior, and my religious friends would rightfully object that we had no good reason to privilege mathematical weightings. So this third post tries to be my final answer, that even if we don’t pre-suppose a mathematical prior, and we care about start with a broader set of weightings we could possibly care about, we will still update back to a largely materialist world-view and not giving much weight to God after updating on our observations.
Hi! I’m no longer in the loop enough to constructively engage anymore, but thanks for the update in any case!
Note that in most developed countries, 20-30% of the GDP already goes to supporting people who no longer provide economic value but whom we we still love and respect for having raised the current generation.
I think exogneoue xrisk is probably very small, especially that I’m in favor of starting some chill levels od space colonization (e.g. Mars) pretty soon.
It’s not impossible but I would find it very surprising if your TED AI collective was unwise or untrustrworthy enough that it could drift into misalignment over time without stopping itself, but trustworthy and wise enough that we could have handed off the intelligrnt explosion to them, and they would have safely got us to the limits of intelligence.
Yep, the third point is a serious objection, but I feel that if our descendants never feel ready to buikd superintelligence, then maybe they just shouldn’t. I think it’s plausible that TED AI plus maybe narrow superintelligences (e.g. specialized in space-probe design) are enough for a near maximally good future.
Yeah, I agree he was pretty wrong. But I think that a big part of the story is that he saw at the time politicians on both sides talking about nuclear war as a winnable fight (he quotes politicians on this extensively in his 1958 speech on the topic), and this made him think that one side would eventually pull the trigger. In reality, the concept of mutually assured destruction and nuclear wars being unwinnable eventually became common wisdom among politicians, and then nuclear sabre-rattling went down. And I think that’s downstream of the fact that military tech in fact didn’t improve that much, and no one yet managed to build a reliable response to nukes + ICBMs.
I should probably engage more deeply with the basin of good reference post. But I currently thank 5-10 years is not nearly enough that “we” can confirm that the AI is in the basin of good deference. Sure, we can maybe prove with high confidence that the AI is not actively scheming against us, and maybe even the voters will be justified in believing the narrow claim from the scientists that the AI is not actively scheming. But otherwise, I very much sympathize with Richard Ngo’s point that there is no way for most normal people to understand and consent to this process so quickly. 5-10 is like two election cycles! It’s really not very much time! And I’m not just talking about the random voters, I feel that I myself will be very distrustful of the Jupiter-brained creature we are creating for whom the concept of instruction following breaks down.
I don’t have great ideas what to do instead, but I think probably gradualism is the best approach. The one type of alignment that people have experience with is raising children. The one type of value evolution people feel relatively fine with is the new generation of people gradually shaping culture in their own way.
So let’s just have some kids first who can grow up in a more abundant society, built by the limited involvement of TED AIs. Then we have guided humanity through the times of peril, and solved our own small portion of the alignment problem by lovingly raising a generation of somewhat smarter and better-adjusted people, who can then carry on the torch. Solving the rest of the alignment problem is their job, not ours.
Tentatively, I think that after one or two generations, people should start using a moderate amount of human genetic engineering to shape the next generation to be even smarter, wiser and happier. Then at some point, probably one generation should start merging with the machines and becoming cyborgs. I don’t know what happens after that, it’s their choice. And throughout this process, I hope that there always remain some people (perhaps not very many, probably keeping only a small sliver of the stars to themselves), who wish not to progress to the next level of development, and wish to remain in their only more-or-less enhance human form, but who will still love the people who progress one step further in the next generation, and who will still be loved by them in turn.
I don’t think that’s true, and I tried to argue in the Pausing after unipolarity section that different level of alignment and wisdom is needed for the two goals.
Imagine that individual AI instances are about as smart, thoughtful and benevolent as the average capabilities researcher at an AI company right now. I think setting up a large army of these instances to maintain the peace might work pretty well without them fully taking over the world. With some clever checks-and-balances mechanisms alluded to in this post, I think this collective of moderately well-meaning AIs could remember their oaths to remain loyal to human democracy and function well. But I’m much more scared of tasking the collective of these AI instances to decide if it’s safe to scale further and then solve the alignment of the next generations well.
So I think there is a very wide gap between how wise and aligned the AIs need to be to trust them to do a peace-keeping operation with a narrow scope vs to trust them to build the superintelligence whose values will determine what the galaxies get filled with.
Tangentially answering to the more general point from Richard Ngo: I think the long-term arc of history in fact points towards unipolarity. In another, related Richard Ngo post, 1a3orn comments;
I generally think rationalists gesture at [multipolar equilibria not being stable long-term] as inevitable more than they actually demonstrate that it is inevitable; i.e., long term multi-polar equilibria has been quite sticky for Westphalian states or for plankton.
I don’t know much about plankton, but I think it’s notable that at the time of the Peace of Westphalia, there were about 300 sovereign states within the Holy Roman Empire, but by now they consolidated to only ~7 sovereign states on the same territory, almost all being members of the EU and NATO.[1] I don’t think historical trends are that favorable to multipolarity!
- ^
Austria is not in NATO, and Liechtenstein is neither EU nor NATO, but realistically both are close allies of the larger block around them.
- ^
I think Russell’s prediction was not that bad. I don’t know what probability we should translate to his statement that “unless something quite unforeseen occurs”. 90% maybe. And what was the real probability that we would have gotten a nuclear war or effectively a world government by 2000? I think about 50%?
I think Russell in 1951, seeing the ever-increasing pace of technological development during his life since 1872, underestimated the possibility that the world would soon enter the Great Stagnation, and the pace of technological progress would stop increasing and would in many ways slow down. In the 1951 essay we are talking about, he writes:
”The first possibility, the extinction of the human race, is not to be expected in the next world war, unless that war is postponed for a longer time than now seems probable. But if the next world war is indecisive, or if the victors are unwise, and if organized states survive it, a period of feverish technical development may be expected to follow its conclusion.”
I don’t know how much his prediction of feverish technical development depends on this happening after a war, but I think he would be surprised that in 2000 hydrogen bombs were still the most powerful weapons, and that the equilibrium of both sides having nukes and ICBMs but no one managing to develop reliable ICBM interceptors remained stable for so long.
I think AI will likely end the Great Stagnation and bring us back to ever-accelerating feverish technical progress, and Russell’s prediction will come true.
As I tried to stress throughout the post, I agree that pausing sooner is the better option. But I think people should have contingency plans for pausing at least after unipolarity, instead of just scaling to the limits of intelligence as some currently existing plans suggest.
“I was prompted to think about why standard human checks-and-balances against violent takeover would fail for humans given that they work for AIs.”
Did you mean to write “would fail for AIs given that they work for humans”?
Pause, at least after unipolarity
I think it was never plausible that more than 99% of alien civilizations wipe themselves out with engineered pandemics before they start spreading to the stars. As Scott argues, the Great Filter needs to explain why none of the ~10^18 planets in our past lightcone developed an interstellar civilization that could have colonized us, so something like pandemics are not nearly big enough risk to be plausible filter.
I think this has been pretty clear for a while now. See Scott’s arguments from 2014.
How do you have acausal effects on the most dramatic possible universe? I usually imagine acausal effects coming in two types:
a) Once we achieve maturity, we try to engage in acausal trade in different universes—I mostly imagine this as us being in a simulation or something like that and serving the values of the simulators in our world to some degree, in return for being released from the simulation and being given resources in the outside world.
b) some of our current actions have correlations with other actions in different universes, and we want to take that into account. I think ECL talks about this, but I haven’t actually read much on ECL.
Point a), acausal trade, just means that there is an entity in the most dramatic possible universe who for some reason takes an interest in what happens with us, and is willing to give us resources in their universe in return for something. Because I imagine this entity as a simulator, I’m often thinking of it as some kind of deity. And I feel that that there is a common factor of “how weird it is to give special value to mathematically simple physics” that determines both how much I care about mathematically simple worlds, and how much I should expect entities in other types of universes to care about mathematically simple worlds. That’s what I mean by both being 0.1% in the example.
(Note that acausally dealing with entities in the most dramatic universes and so on feels pretty different to me than acausallly dealing with entities in e.g. other quantum branches. When trading with other quantum branches, I think we will probably have some kind of common currency of the trade, and the more galaxies we conquer, the more we can trade with. When dealing with entities in the most dramatic universes who for some reason care about us, I feel like all bets are off, what they want and I wouldn’t prioritize conquering the galaxies that much. That’s why I emphasize acquiring personal wisdom and virtue as the most robust strategy for dealing with these entities from very exotic universes.)
Point b) is admittedly different, and maybe I should have emphasized it more. When I think about which decision-process I should use for my current actions such that generalizing that decision-process across weird regions of the multiverse leads to positive results, I feel that “just do what seems best based on causal considerations” seems almost always the right choice. I think maybe the main deviation from the causally optimal strategy is some amount of extra kindness towards beings with different values. But I think this effect is pretty weak in practice.
I already mentioned this Point b) effect in my previous post (see this section) because I think the same Point b) considerations apply when thinking of the correlationary effects of our actions in a normal quantum multiverse or in the more exotic universes I talk about here. But I agree I should have probably emphasised the Point b) considerations more here too, because even if they only suggest a small deviation from the causally optimal strategy in our current actions, they still plausibly overwhelm the also small effect coming from the preferences of super unpredictable entities in Point a).