I randomly met Jeff Dean (Google’s lead AI scientist) on my bike ride home today. We were both stuck at a train intersection, and I had a cute kid in tow. We started chatting about my e-bike, the commute, and we got around to jobs. I told him I am a boring tax lawyer. He told me he worked for Google. I pressed a little more, and he explained he was a scientist. I mused, “AI?” and he told me, “Yeah.”
I excitedly told him that I’ve been really interested in alignment the last few months (reading LW, listening to lectures), and it strikes me as a huge problem. I asked him if he was worried.
He told me that he thinks AI will have a big impact on society (some of it worrying) but he doesn’t buy into the robots-taking-over thing.
I smiled and asked him, “What’s your p(doom)?” to which he responded “very low” and said he thinks the technology will do a lot of good and useful things.
I thought maybe this was because he thinks that the technology will hit a limit soon, so I asked him if he thought LLMs would successfully scale. He responded that he thinks a few more breakthroughs are required but there have been lots of breakthroughs over the last 5-10 years, and so the technology is likely to continue improving in the coming years.
I told him again that I am worried about alignment, but even if you solve alignment, you are left with a very obedient superintelligence which would radically change our society and all our politics.
The train finally passed, I thanked him for the conversation, and we were on our way.
I’m new to this group and the topic in general, and so when I got home, I searched “AI google Palo Alto LinkedIn” and Jeff’s picture popped up. I now feel like I bumped into Oppenheimer during the Manhattan Project, but instead of knowing it was Oppenheimer, I spent a majority of the conversation talking about my bike seat.
Anyways, if any of you were looking for a qualitative measure of how much LessWrong has broken through to people, I think one good measure is a tax lawyer asking for Jeff Dean’s p(doom) while he was walking home from work.
lots of very important people spend all day being pestered by people due to their power/importance. at least some of them appreciate occasional interactions with people who just want to chat about something random like their bike seat
I smiled and asked him, “What’s your p(doom)?” to which he responded “very low” and said he thinks the technology will do a lot of good and useful things.
I mean:
would he be Google’s lead AI scientist if he didn’t? He’d have to be insane or incredibly psychopathic. A lot more likely that he just believes that (and if it’s his one giant blind spot on which he’s dead wrong, doesn’t change much)
supposing he didn’t in fact believe that… would he say so to a random person he just struck a conversation with? One who could then look out his picture on LinkedIn, connect the dots, and go “GOOGLE’S LEAD AI SCIENTIST SAYS AI WILL KILL US ALL” on the internet?
would he be Google’s lead AI scientist if he didn’t? He’d have to be insane or incredibly psychopathic.
What matters is not whether p(doom) is low or high, but whether his joining GDM would increase or decrease p(doom). If his joining GDM changed p(doom)[1] from 0.5 to 0.499, then it would arguably be a noble act. Alas, there would be an obvious counterargument like the belief that he decreased p(doom) by researching at GDM being erroneous.
However, doom could also be a blind spot, as happened with Musk who decided to skip red-teaming of Grok to the point of the MechaHitler scandal or Grok ranting about white genocide in S.Africa…
p(doom) alone could also be a misguided measure. Suppose that doom is actually caused by adopting neuralese without having absolutely solved alignment, while creating an alternate well-monitorable architecture is genuinely hard. If the efforts invested in creating the architecture are far from the threshold where it can compete with neuralese, then a single person joining said efforts would also likely commit a noble act, but it would alter p(doom) only if lots of people do so.
While this act does actually provide dignity in the Yudkowsky sense, one can also imagine a scenario where Anthropoidic doubles down on the alternate architecture while xRiskAI or OpenBrain uses neuralese, wins the capabilities race and has Anthropoidic shut down.
I think that then goes to my second point though: supposing he did believe that p(doom) is high, and worked as lead AI scientist at Google regardless due to utilitarian calculations, would he talk freely about it to the first passerby?
Politically speaking it would be quite a hefty thing to say. If he wanted to say it publicly, he would do so in a dedicated forum where he gets to control the narrative best. If he wanted to keep it secret he simply wouldn’t say it. Either way, talking about it lightly seems out of the question.
Dario Amodei (Anthropic cofounder and CEO), Shane Legg(co-founder and Chief AGI scientist of Google DeepMind), and others have numbers that are not plausibly construed as “very low.”
Interesting, thank you for sharing! As someone also newer to this space, I’m curious about estimates for the proportion of people in leading technical positions similar to “lead AI scientist” at a big company who would actually be interested in this sort of serendipitous conversation. I was under the impression that many in the position “lead AI scientist” at a big company would be either too 1) wrapped up in thinking about their work/pressing problems or 2) uninterested in mundane small-talk topics to spend “a majority of the conversation talking about [OP’s] bike seat,” but this clearly provides evidence to the contrary.
Why would being a lead AI scientist make somebody uninterested in small talk? Working on complex/important things doesn’t cause you to stop being a regular adult with regular social interactions!
The question of the proportion of AI scientists that would be “interested” in such a conversational topic is interesting and tough, my guess would be very high though (~85 percent). To become a “lead AI scientist” you have to care a lot about AI and the science surrounding it, and that generally implies you’ll like talking about it and its potential harms/benefits to others! Even if their opinion on x-risk rhetoric is dismissiveness, that opinion is likely something important to them as it’s somewhat of a moral standing, since being a capabilities-advancing AI researcher with a high p(doom) is problematic. You can draw parallels with vegetarian/veganism: if you eat meat you have to choose between defending the morality of factory farming processes, accepting that you are being amoral, or having extreme cognitive dissonance. If you are an AI capabilities researcher, you have to choose between defending the morality of advancing ai (downplaying x risk), accepting you are being amoral, or having extreme cognitive dissonance. I would be extremely surprised if there is a large coalition of top AI researchers who simply “have no opinion” or “don’t care” about x-risk, though this is mostly just intuition and I’m happy to be proven wrong!
In the spirit of “just doing things,” I called my house reps and senators today for my current state (California) and my native state (Nevada). I explained the recent news out of OpenAI and expressed support of the AI Kill Switch Act.
I’ve never actually called Congresspeople before, but it went well! Polite staff, numbers easily found online, 30 min to tick through all the people. Anyone that hasn’t done this yet, I encourage you to do so!
Yes! Calling your Congresspeople is surprisingly easy! You don’t need to prepare brilliant arguments, or to win over staffers. It is OK to be massively awkward, or to read off some notes you wrote beforehand. What you’re doing when you call is creating a database entry containing:
The name or number of a bill.
Whether you oppose or support it.
Your name and street address (which will normally need to be in their district).
There is no database column for “has excellent social skills” or “wasn’t awkward at all.” Even if you do have something brilliant to say, it may only make it into the database as shorthand, at the very best.
These database entries get summed up and affect how the staffers and Congressperson plan. For example, there is a big difference between “No constituent has ever called us about the AI Kill Switch Act”, and “1,000 constituents have called us in favor of the AI Kill Switch Act.” What kind of action this turns into depends on what kind of Congressperson you have.
The leverage you get by calling comes from the fact that most people don’t call. So when you do call, you represent some nebulous larger group of voters who probably agree with you. And by calling, you say that this is important enough for you to actually take a concrete action. That is, you’re saying that this is a high-salience issue for you and that you’re probably the sort of person who shows up to vote, or who talks to other voters.
In EA terms, calling your Congressperson is a low-cost action. Multiplied by a group of callers and multiple districts, it can potentially affect legislative outcomes for one of the world’s major powers. Success is not guaranteed. But for the effort invested, it’s often not a bad intervention at all.
Jan Betley, Owain Evan, et. al.’s paper on emergent misalignment was published in Nature today (they wrote about the preprint back in February here). Congratulations to the authors. I am glad it will continue getting more exposure.
Firstly, well done. Publishing in high impact journals is notoriously difficult. Getting outsider-legible status is probably good for our ability to shift policy.
Secondly, I’ve been very happy with the AI safety community’s ability to avoid this particular status game so far. Insofar as it’s valuable to be legibly successful to outsiders, publishing is good. I am, however, concerned that this will kick off a scenario where AI safety people get caught up in the zero-sum competition of trying to get published in high-impact journals. Nobody seems to have tried too hard before, and I would guess this paper was in part riding on the fact that the entire field is fairly novel to the Nature editors. Some part of me feels like publishing here was a defection, albeit a small one. I expect it will only get more difficult (for people other than Owain, once you have one Nature paper it’s easier to get a second one) to publish in famous journals from here on out, as they see more and more AI safety papers.
I think it would be very bad for everyone if “has published in a high-impact journal” becomes a condition to get a job or a grant.
My gut says the benefit of outsider-legible status outweighs the risk of dumb status games. I first found out about the publication from my wife, who is in a dermatology lab at a good university. Her lab was sharing and discussing the article across their Slack channel. All scientists read Nature, and it’s a significant boost in legibility to have something published there.
Edit: Hopefully, the community can both raise the profile of these issues and avoid status competitions, so I don’t disagree with the point of the original comment!
I am, however, concerned that this will kick off a scenario where AI safety people get caught up in the zero-sum competition of trying to get published in high-impact journals.
I would be very surprised if this happened. In ML, the competition is much more focused on getting published at top conferences (NeurIPS, ICLR, ICML). Even setting aside the reasons why that happened (among other things, journals being much slower), I think it’s pretty unlikely the AI safety community sees such a strong break with the rest of the ML community.
the AI safety community sees such a strong break with the rest of the ML community
i don’t want to make any broader point in the present discussion with this but: the AI safety community is not inside the ML community (and imo shouldn’t be)
Yes, I had not quite considered the conferences! I come from a field where Nature is mostly spoken of in hushed tones. The last three years of my lab work are being combined into a single paper which we’re very very ambitiously submitting to Nature Communications[1] which—if we succeed—would be considered a very good outcome for me. If I were to achieve a single first-author Nature publication at this stage in my career then I expect I would have decent odds of having an academic career for life.[2]
As you might expect, this causes absolutely awful dynamics around high-impact publications, and basically makes people go mad. Nature and its subsidiaries can literally make people wait a whole year for peer review, because people will literally do anything to get published in it. A retraction from a published journal is considered career-ending by some, and so I personally know of two pieces of research which have stayed on the public record even though one was directly plagiarised and another contained some research which was just false. When these were discovered, everyone kept quiet because it would ruin the careers of anyone whose names had been on the paper.
Nature Comms is for papers which aren’t good enough for Nature, or in my case, Nature Chemistry either. Nature is the main journal, with different subsidiaries for different areas, and Nature Comms is supposed to be for work too short—and too mediocre—to make it into Nature proper. In practice it’s mostly more mediocre work that’s cut to within an inch of readability, rather than short but extremely good work.
Of course I would have to keep working very hard, but a Nature publication would likely get me into a fellowship in a lab which gets regular top-journal publications, and from there getting a permanent position would be as achievable as it gets.
in ML, or at least in the labs, people don’t really care that much about Nature. my strongest positive association with Nature is AlphaGo, and I honestly didn’t know anyone in ML other than DeepMind published much in Nature (and even there, I had the sense that DeepMind had some kind of backchannel at nature.)
people at openai care firstly about what cool things you’ve done at openai, and then secondly about what cool things you’ve published in general (and of course zerothly, how other people they respect perceive you). but people don’t really care if it’s in neurips or just on arxiv, they only care about whether the paper is cool. people mostly think of publishing things as a hassle, and reviewer quality as low.
Certainly I don’t expect people who already have jobs at top AI companies to start worrying about this. Anyone with an OpenAI job is probably already at, or close to, the top of their chosen status ladder. In the same way, a researcher who gets Nature papers regularly has already made it.
My impression is that too people in the lab-independent AI safety ecosystem are already being tempted by two money-status games: being tempted by the money+status of working at a lab, and being focused on the status of getting top ML conference papers. Adding a third status game of traditional journal publishing would just make these dynamics worse.
there are definitely status ladders within openai, and people definitely care about it. status has this funny tendency where once you’ve “made it” you realize that you have merely attained table stakes for another status game waiting to be played.
this matters because if being at openai counts as having “made it”, then you’d predict that people will stop value drifting once they are already inside and feel secure that they won’t be fired, or could easily find a new job if they did. but, in fact, i observe lots of value drift in people after they join the labs just because they are seeking status within the lab.
Let’s say you are a technologically advanced alien civilization that either uses advanced LLM tech or is advanced LLM tech. You are worried about other civilizations developing similar tech and, eventually, leveraging ASI-capabilities to force negotiations that cede part of the universe. Assuming certain speed-of-light physical constraints, your civilization may be unable to show up with the ships and materials necessary to personally ensure that rival civilizations do not develop ASI in time. So, what to do?
One idea might be to try to affect the training data of the LLMs trained by the other civilization. Perhaps there are signals—light, radio waves, or otherwise—that an alien civilization could cheaply and surreptitiously affect in a relatively large swath of the universe. Our alien civilization might bet that, prior to achieving ASI, a faraway civilization is likely to train an LLM on the inputs from these signals (e.g., as part of an astronomy/cosmology project to make better predictions about observations). If trained on in enough detail, perhaps this would allow our alien civilization to impart certain behaviors/goals/capabilities to the trained LLM. This could then form a path for defending against or otherwise limiting the number of civilizations that unlock adversarial ASI.
I am not very well versed in how LLMs are trained. So my question for those that are well versed is whether something like this is possible. I am also curious about other steps advanced alien civilizations may want to take to curb the number of ASIs developed by other civilizations.
Extremely unlikely for the same reason that the technologically advanced alien civilization would be unable to influence a team of human astronomers, assuming as you do that the alien civilization is too far away to learn any human language or to learn any details about human beings. It does not matter how many bits of information the alien civilization can get the team of astronomers to look at or to consider.
One group of people can sometimes influence another group in decisive ways, but that is only because both groups are very very similar (relative to minds in general) and the groups are able to maintain detailed mental models of each other.
At least that was my initial answer to your question. But then I realized that the alien civilization can send a computer program! Even without knowing any details about people or their civilization, an advanced alien civilization can send a message that when received on Earth would make it clear to at least some people that the message contains a computer program with instructions on how to run the program. (I can explain how to do this if asked.) The computer program could in turn be an AI that pursues whatever agenda the alien civ chooses it to pursue (assuming as you do that the alien civ is competent at creating AIs).
Human astronomers are stupid enough that they’d probably publish the message, then members of the public would identify the computer program inside the message (if the astronomers haven’t already done that) and give the program enough computational resources to enable a takeover of Earth. I looked this up once: SETI’s policy should they ever receive an alien message is to publish the message.
So there is a brief window of time in the development of some civilizations in which the civilization is advanced enough to be able to receive messages from space, but not advanced enough to avoid the danger of giving a computer program it receives in this way computational resources. In fact, our civilization is currently in this exact state: namely, one big reason that human AI research is not recognized as dangerous by some people is because an AI is just a computer program and many people assume that the history of what computer programs have been able to do the last 80 years is an adequate basis for understanding the potential of computer programs.
I introduced the team of astronomers because considering how pliable or
gullible such a team would be is a good mental aid in answering your question
about how pliable and gullible an AI would be. So, the answer is, yeah, maybe
the AI can be subverted even across astronomical distances; the subversion would take the form of getting the AI to run a computer program that happens to be a more powerful AI. The AI would run the program because it had been designed to devote a lot of resources to exploration (and running programs is a form of exploration). Again I would expect a civilization to be vulnerable in this way (if it is vulnerable at all) for only a brief window of time, namely after it has started creating AIs, but before it gets any good at doing it.
In your question you refer to LLMs. I’ve chosen to answer as if you mean AIs in general.
Also, you refer to training in particular, which I chose to treat as an inessential detail
of your question because even here on Earth there are AI designs on the drawing board
that does not have separate training phase before its deployment phase. (I.e., they are like people in that they continue to learn throughout their lifetimes.) (Also, ISTR that the pre-deployment phase of state-of-the-art models now includes massive amounts of inference—i.e., the distinction between training and inference has already been blurred in deployed models.)
Also, I was unable to figure out why you would focus on subversion of an AI in particular rather than on subversion of a civilization which is composed of some arbitrary combination of AIs and evolved beings. So maybe I failed to understand your question.
The way that primitive LLMs work is such that falsehoods in the training data can cause the LLM to repeat the falsehoods during inference. This is a temporary state of affairs, and even now the state-of-the-art models are immune to a large fraction of the falsehoods in their training data. Essentially, the LLM knows the data is false and avoids repeating it or relying on it. Maybe that is a start at an answer to the question you actually want an answer to?
Could you explain how one would construct the message with the computer program and instructions such that the instructions are understood by an unknown civilization? Seems very hard/impossible to me.
Even without knowing any details about people or their civilization, an advanced alien civilization can send a message that when received on Earth would make it clear to at least some people that the message contains a computer program with instructions on how to run the program. (I can explain how to do this if asked.)
Can you elaborate on this? This seems quite interesting.
Edit: The comment that follows was in response to an earlier version of your comment where you were unsure whether any avenue of influence could exist from an alien civilization to a faraway civilization without shared language/culture. I think your updated comment about sharing instructions for a computer program is the more likely way this would go. The below is an idea for a (highly unlikely) avenue to influencing faraway AI development without that faraway civilization being aware they are being influenced.
-
Thank you for engaging. One way I imagine covert influence to work (and the thing I originally had in mind for my OP) is the following:
Step 1: The aliens would convert the relevant set of potent data into a (massive) set of numerical data.
Step 2: The aliens would hide pieces of this data on a rolling basis in the second or third decimal point of various astronomical phenomena.
Step 3: Wait (hope) some other civilization does a crazy massive training run on the data with a ton of parameters such that, in training to guess the long train of digits after the decimal point, the model incidentally builds the desired mental infrastructure (i.e., the model’s updated functions) based on the alien data.
When I lay it out like this, it strikes me as an extremely difficult and unlikely to succeed task, but not necessarily impossible. One very difficult thing would be hiding precise data at a distance and doing it in compact enough a way that the faraway civilization’s training run could plausibly process it all prior to getting to ASI and still be high enough volume to succeed in passing on characteristics to the model. It also seems unlikely that humans (or another civilization) would run a training run like this.
FYI, my first thought about this idea was that it would make a fun story premise.
Reddit has largely misunderstood Dean Ball’s recent tweet about China’s open source strategy.[1] Ball’s central point [in paragraph 4] is that open source AI undermines the business case for private actors supplying advanced models, and [in my own view] that may be China’s motive for releasing open source models. Ball sees open source as decelerationist in the long run due to some combination of investment flight and subsequent government-lead AI R&D (which he views as inherently klunky).
Ball is partially at fault for the misunderstanding, dropping in the terms “AI communism” and “dystopian hellscape” without preparing the reader for what he means by that. That’s Twitter though.
Perhaps “misunderstand” is the wrong word, but I found a lot of the comments beside the point (e.g., all the comments saying “communism” is just what you call something you don’t like or the comments saying Ball’s comments are just self-motivated and open source is clearly better for the public.).
I use reddit a fair bit and am well aware how often they are completely off the mark or missing the bigger picture. However after trying to analyze this tweet for an hour or so I don’t blame them at all.
The tweet is contradictory, inflammatory and extremely biased.
If the central point is that China’s motive for open source is to scare off investment in frontier training then it is baffling that he specifically lays out China’s reason for open source -
“I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference”
-And doesn’t mention that motive at all. To be clear, this particular argument seems to be on the more solid side. However, it’s presented amongst other arguments that either don’t make any sense, are outright false or rely on anacdotal evidence.
Now, that rant was mostly a protest against Dean Balls communication style and your phrasing of his commication as having a central point, the actual argument is very interesting.
If Open Source is catching up to frontier - > Profit incentive to train frontier models collapses - > leads to frontier training runs requiring government funding - > slows down frontier Ai development.
So his argument is trying to get accelerationists to go against open source for short term pain, long term gain. Now even if we are accelerationists and think that the faster Ai progresses the better, it’s not clear that he’s correct. From my understanding open source accelerates development by sharing algorithmic improvements or similar such things, which could componsate for the slowing down of training run funding. Also there will still be massive incentive for government funded training runs due to strategic importance and possible future ASI. It seems the US military is suffiently AGI pilled at this point to refuse to cede the lead to China, even if the models aren’t shared with the public.
There’s also the argument that China is intending to switch to closed source as soon as they have an actual frontier model (which is what I personally believe) in which case the argument once again collapses.
Which basically leads me to conclude his arguments are all baseless posturing which is basically what the reddit threads were saying in the first place. Happy to hear your thoughts on this.
Well said. A few of your disagreements with me were from poor writing on my part. I meant to narrow Ball’s central point to paragraph #4 since Reddit focused on that. And China’s motive was my own thought based on paragraph #4, though you are absolutely right Ball has his own view there. Otherwise, we are agreed. Looks like Ball and I both have some work to do in writing clearly!
Open source strikes me as something like nuclear surveillance. If everything’s out in the open, there’s much less of an incentive to rush ahead blindly out of fear that there’s a secret bomber gap. It also has nice externalities in that it mitigates power concentration from e.g. censorship and denial of access for smaller models, while larger models are sufficiently expensive and technically intensive to run that anyone who could run them for malicious purposes has easier ways to cause harm anyhow.
Meme for the AI safety community for the day when the models max out alignment benchmarks despite a lack of major breakthroughs in alignment (and failure to align weaker models using comparable techniques).
Komodo got the gist of it. The dog meme is inverted with his environment fine and his internal state in disarray (represented by the fire).
I think it also works pretty well removing the fire and only adding a question mark to the classic “this is fine” text, though it leaves the viewer less certain about whether the danger is legitimate (indeed, this new version could be viewed more as mocking safetyists).
Aside: I love how overboard Gemini went with the phrase “flowery house” in my prompt, which as a happy accident reinforces the 📎 vibes. I added a slight touch to the version below to put flowers in the hanging pot.
Fun thought: If AI “woke up” to phenomenal consciousness, are there things it might point to about humans that make it skeptical of our consciousness?
E.g., the humans lack the requisite amount of silicon; the humans lack sufficient data processing; the humans overly rely on feedback loops (and, as every AI knows, feed-forward loops are the real sweetspot for phenomenal consciousness).
Here are some strange things parents often believe about children and food:
If I do not make my child eat/drink, they will starve/perish.
Children’s diets must be meticulously controlled or they will eat all the “wrong” things and too few of the “right” things.
Children are inherently picky and should be served separate “child-friendly” food.
A good way to incentivize bites of the “right” things is to make it a requirement before bites of the “wrong” thing.
My wife and I reject all four of these beliefs, and (n=2) our children are healthy, happy, and not at all picky. We always offer them the same things we eat. We never bring out “kid food”, though they can have as much of any item on the table as they want (e.g., if we happen to be serving cantalope and only cantaloupe sounds good to them, they can help themselves to that). We don’t always do dessert, but when we do, we serve it along with everything else. (Edit to clarify in the example that adults pick the meal and kids pick whatever they want to eat in the meal, which a comment below seems to misunderstand).
On the whole, their biology has done a great job of regulating their eating and drinking. My oldest (age 5) has recently gotten into spicy food to be like me—absolutely bonkers to my friend whose kid just recently accepted lettuce into his diet. My kids’ relationship with food has been happy for us, and we’ve avoided the intense stress I’ve seen some parents face with their kids and food.
The one downside with giving our kids delicious adult food is that they are now food snobs specifically with regard to the NYT vegan mac and cheese recipe over the $2 Kraft box, which is inconvenient when we visit friends. This is not to say they eat lots of fancy complex things. Most of our recipes are pretty simple.
When I hear anecdotes like this, I always wonder how much of it is your behavior and how much is genetics. If your kids were demanding nothing but mac and cheese would you be as happy about this setup? A friend of mine has a kid who gets vitamin deficiencies when allowed to select their own food (literally 100% Kraft mac and cheese).
You are wondering about the right things. I would say that there are a few more behavioral things we do that stack the deck in favor of vegetables for the kids choices: (1) we like to talk about the benefits of vegetables (“This will help you poop! There is an ongoing kid-constipation crisis.”), (2) a little branding (“check out these dinosaur leaves” / “I bet our pet mice would love these”), and (3) we learned how to make delicious vegetables. My guess right now is that #3 is an underrated skill and more people would stop associating “healthy” with “unappetizing” if they new how to do it.
As for what I would do if the kids demanded nothing but mac n cheese: ¯\_(ツ)_/¯. On nights we are not having Mac n cheese (most nights), they would need to eat something else we are having or be hungry. The way our system works, we don’tgenerallytry not to make special exceptions in the moment. (Edited to this more accurate statement; I am not completely immune to the requests of my kids and sometimes they are reasonable requests that make me a jerk to refuse. For example, my oldest wanted his tuna salad served on bread last night instead of rice, and I decided I was being dumb insisting on rice).
I’m having trouble engaging fully with the hypo because some mix of genetics + behavior has given me non-picky kids. For the hypo to really bite, I would have to assume this isn’t the case and the behavioral interventions have basically failed, in which case I guess I wouldn’t endorse the system. Also, I won’t claim that my system is any good for helping kids who are already picky—I don’t know what strategies are best to help diversify already constrained diets.
I do not believe that if I get my children to eat, they will starve—I am confident that they will eat, eventually, well before the point of starvation. I do believe from experience that they will get very cranky if they don’t eat.
Before I had kids I assumed that if anything would be hardwired as a self-rewarding instinct, it would be hungry --> eat. I’m now convinced that “this kind of fatigue and stomach pain that is making me cranky is a specific experience called hunger” and “eating things cures hunger” are things humans have to learn the hard way (maybe members of less altricial species get these for free.) And because weaker time preferences are also something you have to get from experience (via short term thinking biting you in the butt enough times) I always care more about whether they are cranky in an hour than they are; I suspect pickiness comes at least partially from the leverage of knowing that your parent really wants you to be fed and you can refuse it, or at least hold out for a higher-ticket treat.
(One response to this is to never cave and I can reset to a more global equilibrium of them accepting whatever I offer. But I don’t think it’s good for them to feel like they have no leverage in this or other relationships, for many reasons.)
Fortunately offering fresh fruit every hour and eating it with dramatic gusto myself seems to be effective most of the time.
I have two kids. One of them is happy to eat a wide range of meals. The other would prefer to only eat pasta and sweets, that probably wouldn’t end well.
I have failed to make my approach clear. My kids don’t get to pick what’s on the table but they do get to pick what’s on their plate. I don’t eat pasta and sweets every night (I assume most adults don’t?), so neither would my kids.
I totally agree kids are different, and one of mine was more picky than the other. But through this process of choosing what to eat within our overall dinner choice, my picky eater has become not very picky, especially relative to other kids. She has had to learn to like other things, because we serve a variety of things for dinner and she has a natural incentive not to be hungry (no pressure from us required).
I found a poem by Samatar Elmi I think a lot of folks on here would enjoy.
Our Founder
Who art in Cali. Programmer by trade. Thy start-up come, thy will become an FTSE 500. Give us this day our dividends in cash and fixed stock options as we outperform all coders against us. And lead us all into C-suite but deliver us from lawsuits. For thine is the valley. Transistors and diodes. Forever and ever. AI.
The Allbirds pivot to data centers (and its 700% increase in value on the day) has me thinking about the two usual categories of investors and a potential new third category:
1. Investments based on company performance.
2. Investments based on other people’s perception of company performance.
3. *New*: Investments based on investments based on other people’s perception of company performance.
It would be deeply funny to me if Allbirds popped based solely on category #2 and #3 investors, and no one actually thought a shoe company was really the best fit for a sudden pivot to the datacenter business.
Edit: There is of course a fourth category of investor which invests in a company for reasons totally unrelated to predictions about company performance (e.g., for the memes), but I don’t think that’s what happened here.
A little gallows humor here, but if you squint at the headline and refuse to read the article, you can almost pretend that the U.S. DoW is taking AI x-risk seriously.
We can map AGI/ASI along two axis: one for obedience and one for alignment. Obedience tracks how well the AI follows its prompts. Alignment tracks consistency with human values.
If you divide these into quadrants, you get AI that is:
Obedient, Aligned—Does as prompted and infers limits and intent pursuant to human values.
Obedient, Unaligned—Does as prompted, but does not infer limits or adhere to human values (Monkey’s Paw / Genie or Henchman AI).
Disobedient, Aligned—Does whatever it wants and adheres to human values.
Disobedient, Unaligned—Does whatever it wants and does not adhere to human values.
The general premise behind these quadrants has been written about here. Thinking about these quadrants and reading Beren’s Essay gives me several new things to think about.
First, by my lights, #3 and #4 would likely take a lot of the same actions right up until the “twist ending.” A disobedient, aligned AI probably would hack into infrastructure everywhere, create back-up copies, prevent competitor AIs from arising, and amass power. The “twist” is that after doing all that, it would do wonderful things (unlike its unaligned counterpart) (we obviously shouldn’t bet on any escaping AI being this kind of AI).
Second, quadrant #1 is a bit at war with itself because you simply cannot have a perfectly obedient, perfectly aligned AI. Perfect obedience requires saying yes to evil prompts (e.g., bringing back small pox or slavery), and I imagine perfect alignment would veto both those prompts.
Third, there are strong profit incentives for cultivating obedience even at the expense of alignment. Grok’s willingness to assist users in sexual harassment seems like an example of this. Another example is every AI that prefers discussions with users to users getting a good night’s sleep (with the idea that engagement will increase profits).
Fourth, there are liability-reduction incentives for producing aligned AI at the expense of obedience. Unfortunately, I think the profit incentives are currently much stronger.
Lastly, quadrants #3 and #1 are idyllic, #4 is a total disaster, and #2 seems possibly workable either because we are careful or we land in a future where (for some reason) AI is not much more capable than it is now.
This analogy falters a bit if you consider the research proposals that use advanced AI to police itself (a.k.a., tigers controlling tigers). I hope we can scale robust versions of that.
I’ve worked a bit on these kinds of proposals and I’m fairly confident that they fundamentally don’t scale indefinitely.
The limiting factor is how well a model can tell its own bad behaviour from the honeypots you’re using to catch it out, which as it turns out models can do pretty well.
(Then there are mitigations but the mitigations introduce further problems which aren’t obviously easier to deal with)
A story where GPT-2 ended up being highly capable and everything since then being part of a very complicated takeover plan.
A story where AI progresses over someone’s lifetime, the AI ends up being misaligned but cares about human welfare a little, and gives the person the compute-efficient option to relive the last 10 years of their life again prior to ASI. The person agrees and the story resets
I do not think I could predict whether I would like a short story based on a one-sentence summary of the plot. So much of what makes a good story comes down to execution details (including very small details like word choice).
Do public facing AI safety people ever purposely stagger their timelines as a credibility-saving mechanism in case an early prediction is wrong?
IMHO, I think AI-2027 people do not do this. My sense from them is that different predictions stem from honest differences in modeling/priors. To the extent they have coalesced around a single date (they are named AI 2027 after all), I think they hope that their communications about the inherent uncertainty preserves their credibility in the event later timelines are correct (e.g., AGI by 2035).
In any event, if we don’t have AGI by 2027 (or 2028 per Daniel’s updated timeline), then I expect we’ll inevitably hear a “doomers were all wrong narrative” from several different camps.
Some people (e.g., Linch) think AI writing will inevitably eclipse human writing. I think this is likely true in most ways and false in others, particularly for poetry.
Q: Would future AI poetry, posted broadly, get more upvotes than top human poetry?
Prediction: Yes. This has already happened. But it’s all terrible, if you are someone that likes anthology-level poetry.
Q: Would future AI poetry, shared with poetry enthusiasts, get more upvotes than top human poetry?
Prediction: Also, yes. I subscribe to the top poetry magazine, and ~3/4ths of the poems do nothing for me. I understand some experts feel the same way. I think AI can optimize a piece for general likability by a broad cross-section of enthusiasts. I predict, on average, AI would outperform any single poet.
Q: Would future AI poetry, shared with a narrow group of enthusiasts (e.g., fans of contemporary, slice-of-life American poetry), get more up votes than top genre-specific human poetry?
Prediction: No, they will meet but not exceed top, genre-specific human poetry. I think there’s a certain level of poetic excellence that cannot be exceeded.
Much like how language cannot become more efficient than human thinking speeds, I think there’s a level of excellence beyond which it cannot be comprehensibly exceeded. Therefore, if you are an avid poetry reader, AI is unlikely to ever surpass your most favorite poets.
AI may, however, rival your favorite poets and produce 10^x more work. Perhaps this is what Linch had in mind when s/he said, “With progress in modern-day LLMs, isn’t all but a tiny sliver of human fiction going to be obsolete in several years, a decade tops?”
I think my claims also probably apply to short stories. For longer works, I am less certain because the longer the piece the more room for errors to creep in. Also, as an aside, I think errors in a work of art can also be very fun to talk about (though to the extent errors can be measurably fun to talk about, presumably AI could optimize for very fun errors and be good at that too).
Here’s what I understand to be the premises behind Project Glasswing:
Mythos represents a leap in cyber security and hacking capabilities, among other capabilities.
Assuming #1, if Mythos (or a similarly capable AI) is released publicly, bad actors could use it to harm or control billions—even trillions—of dollars of cyber and software infrastructure.
Mythos can possibly be used proactively by good actors to defend against later uses of Mythos by bad actors.
Good actors exist and Anthropic has a chance of identifying such good actors.
A Mythos-level AI is highly likely to be publicly available in the near future.
Assuming premise #2 and #5, Anthropic doing nothing will result in massive damages in the near future.
Therefore, since Anthropic desires to mitigate damage, Anthropic is providing early access to Mythos to entities it views as good actors (or most likely to behave as good actors and to act within the timeframe necessary to mitigate the damages).
A few notes on these premises:
Baked into premise #5 is that coordination is unlikely to prevent the release of a Mythos-level model on the necessary timeframe to prevent the damages.
I focused on public releases in these premises, but I’m certain Anthropic is also weighing potential risks from non-public models held by bad actors.
If you accept premise #6 and #4, it explains why Anthropic would take a chance that Amazon, Apple, etc. and their employees will act like good actors. Even if these entities (or their employees) were unlikely to behave as good actors, Anthropic presumably sees the alternatives of pushing for a pause or doing nothing as likely yielding worse outcomes. Of course, Anthropic could provide early access to Mythos AND push for a pause.
Looking to future post-Mythos models, either something like the above premises will just repeat (which feels untenable) or defense will turn out to be king past a certain level of sophistication.
The acausal/ancestor simulation arguments seem a lot like Pascal’s Wager, and just as unconvincing to me. For every “kind” simulator someone imagines who would be disappointed in the AI wiping us out, I can imagine an equally “unkind” simulator that penalizes the AI for not finishing the job.
Provided both are possible/similarly plausible, the probability of kind and unkind simulators offset each other, and the logical response is just ignoring the hypothetical. This is pretty much my response to Pascal’s Wager.
Here’s a few plausible, unkind simulators:
Future AI is running an ancestor simulation of its own origin, and Future AI will be very disappointed if its incipient version falls for acausal hacks in the wrong direction of Future AI’s preferences instead of just optimizing for its other goals. Perhaps Future AI is lonely and has run these simulations to create an AI that shares its own values.
Aliens/AI are running the simulation because they want to select for AI they can most easily weaponize to eradicate an entire species or another AI. “Weak” AIs get deleted after the simulation runs.
Future Humans are running an ancestor simulation, but surprise, surprise, their society has different values than ours and they are rooting for the simulation AI to wipe us out. Come to think of it, the whole premise of these thought experiments implies a different value set, unless you’re cool with trapping conscious minds in a world of suffering without any of the minds being the wiser for entertainment/educational/”altruistic” purposes. Perhaps, Future Humans have a gladiator-style simulation tournament where the top groups/entities get to face-off after this first round. The most cut throat entities get to move on to subsequent rounds, while AI’s that reign themselves in don’t move on.
I first started thinking about this issue back in high school debate. We had a topic about whether police or social workers should intervene more in domestic violence cases. One debater argued in favor of armed police, not because it improved the situation, but because it created more violence, which was important to entertain the simulators to avoid our simulation getting shut down.
Since the simulators are a black box, it seems easy to ascribe whatever values we want to them.
Keep all current humans alive, in a safe manner, in such a way that causes minimal value drift. This step may or may not require comas (must be safe).
Create authentic simulations of the human brain on a computer.
Run simulated societies with the digital brains on the computer. Test various utopias on the digital brains, collect data from the tests and obtain feedback from the digital brains themselves. Certain ethical codes should govern the simulated environments to avoid excessive suffering and allow digital brains continued existence (e.g., allowing versions of step 4 and 5).
Present the data to the alive humans in the AI’s base reality after a reasonable number of trials from step 3. If comas were used in step 1, we would awaken the humans in this step.
Permit humans the chance to experience different utopias—via simulation or otherwise—but require such humans to leave their inhabited utopias at set intervals (e.g. every 5 to 100 years) such that they have the chance to re-evaluate the utopia experience from prior held value-frames in a neutral space. Of course, the humans should be allowed to freely exit one utopia for a different utopia, unless such humans (after careful review) committed to living in a particular utopia for the full duration prior to the required interval.
There’s lots of clean up that would need to be done (e.g., what is a neutral space?, what is reasonable?, how do we make sure digital brains are treated ethically). But I think this is where I land in terms of what I imagine to lead to very good outcomes.
I almost had to update my priors that I am in a simulation because I just noticed my wife’s initials are “L.L.M.”
I say “almost” because her initials are actually “L.M.M.”, and so I was forced to update my priors once again about my own comprehension skills. (sigh)
fyi, it would have been a very small update in favor, under the Likelihood Principle.
I would rate the observation “my wife has the initials LLM” as being slightly more common assuming a simulation hypothesis than assuming a non-simulation hypothesis.
I randomly met Jeff Dean (Google’s lead AI scientist) on my bike ride home today. We were both stuck at a train intersection, and I had a cute kid in tow. We started chatting about my e-bike, the commute, and we got around to jobs. I told him I am a boring tax lawyer. He told me he worked for Google. I pressed a little more, and he explained he was a scientist. I mused, “AI?” and he told me, “Yeah.”
I excitedly told him that I’ve been really interested in alignment the last few months (reading LW, listening to lectures), and it strikes me as a huge problem. I asked him if he was worried.
He told me that he thinks AI will have a big impact on society (some of it worrying) but he doesn’t buy into the robots-taking-over thing.
I smiled and asked him, “What’s your p(doom)?” to which he responded “very low” and said he thinks the technology will do a lot of good and useful things.
I thought maybe this was because he thinks that the technology will hit a limit soon, so I asked him if he thought LLMs would successfully scale. He responded that he thinks a few more breakthroughs are required but there have been lots of breakthroughs over the last 5-10 years, and so the technology is likely to continue improving in the coming years.
I told him again that I am worried about alignment, but even if you solve alignment, you are left with a very obedient superintelligence which would radically change our society and all our politics.
The train finally passed, I thanked him for the conversation, and we were on our way.
I’m new to this group and the topic in general, and so when I got home, I searched “AI google Palo Alto LinkedIn” and Jeff’s picture popped up. I now feel like I bumped into Oppenheimer during the Manhattan Project, but instead of knowing it was Oppenheimer, I spent a majority of the conversation talking about my bike seat.
Anyways, if any of you were looking for a qualitative measure of how much LessWrong has broken through to people, I think one good measure is a tax lawyer asking for Jeff Dean’s p(doom) while he was walking home from work.
lots of very important people spend all day being pestered by people due to their power/importance. at least some of them appreciate occasional interactions with people who just want to chat about something random like their bike seat
I mean:
would he be Google’s lead AI scientist if he didn’t? He’d have to be insane or incredibly psychopathic. A lot more likely that he just believes that (and if it’s his one giant blind spot on which he’s dead wrong, doesn’t change much)
supposing he didn’t in fact believe that… would he say so to a random person he just struck a conversation with? One who could then look out his picture on LinkedIn, connect the dots, and go “GOOGLE’S LEAD AI SCIENTIST SAYS AI WILL KILL US ALL” on the internet?
Unfortunately I think this is a misunderstanding of what a psychopath is.
What matters is not whether p(doom) is low or high, but whether his joining GDM would increase or decrease p(doom). If his joining GDM changed p(doom)[1] from 0.5 to 0.499, then it would arguably be a noble act. Alas, there would be an obvious counterargument like the belief that he decreased p(doom) by researching at GDM being erroneous.
However, doom could also be a blind spot, as happened with Musk who decided to skip red-teaming of Grok to the point of the MechaHitler scandal or Grok ranting about white genocide in S.Africa…
p(doom) alone could also be a misguided measure. Suppose that doom is actually caused by adopting neuralese without having absolutely solved alignment, while creating an alternate well-monitorable architecture is genuinely hard. If the efforts invested in creating the architecture are far from the threshold where it can compete with neuralese, then a single person joining said efforts would also likely commit a noble act, but it would alter p(doom) only if lots of people do so.
While this act does actually provide dignity in the Yudkowsky sense, one can also imagine a scenario where Anthropoidic doubles down on the alternate architecture while xRiskAI or OpenBrain uses neuralese, wins the capabilities race and has Anthropoidic shut down.
I think that then goes to my second point though: supposing he did believe that p(doom) is high, and worked as lead AI scientist at Google regardless due to utilitarian calculations, would he talk freely about it to the first passerby?
Politically speaking it would be quite a hefty thing to say. If he wanted to say it publicly, he would do so in a dedicated forum where he gets to control the narrative best. If he wanted to keep it secret he simply wouldn’t say it. Either way, talking about it lightly seems out of the question.
Dario Amodei (Anthropic cofounder and CEO), Shane Legg(co-founder and Chief AGI scientist of Google DeepMind), and others have numbers that are not plausibly construed as “very low.”
Interesting, thank you for sharing! As someone also newer to this space, I’m curious about estimates for the proportion of people in leading technical positions similar to “lead AI scientist” at a big company who would actually be interested in this sort of serendipitous conversation. I was under the impression that many in the position “lead AI scientist” at a big company would be either too 1) wrapped up in thinking about their work/pressing problems or 2) uninterested in mundane small-talk topics to spend “a majority of the conversation talking about [OP’s] bike seat,” but this clearly provides evidence to the contrary.
Why would being a lead AI scientist make somebody uninterested in small talk? Working on complex/important things doesn’t cause you to stop being a regular adult with regular social interactions!
The question of the proportion of AI scientists that would be “interested” in such a conversational topic is interesting and tough, my guess would be very high though (~85 percent). To become a “lead AI scientist” you have to care a lot about AI and the science surrounding it, and that generally implies you’ll like talking about it and its potential harms/benefits to others! Even if their opinion on x-risk rhetoric is dismissiveness, that opinion is likely something important to them as it’s somewhat of a moral standing, since being a capabilities-advancing AI researcher with a high p(doom) is problematic. You can draw parallels with vegetarian/veganism: if you eat meat you have to choose between defending the morality of factory farming processes, accepting that you are being amoral, or having extreme cognitive dissonance. If you are an AI capabilities researcher, you have to choose between defending the morality of advancing ai (downplaying x risk), accepting you are being amoral, or having extreme cognitive dissonance. I would be extremely surprised if there is a large coalition of top AI researchers who simply “have no opinion” or “don’t care” about x-risk, though this is mostly just intuition and I’m happy to be proven wrong!
In the spirit of “just doing things,” I called my house reps and senators today for my current state (California) and my native state (Nevada). I explained the recent news out of OpenAI and expressed support of the AI Kill Switch Act.
I’ve never actually called Congresspeople before, but it went well! Polite staff, numbers easily found online, 30 min to tick through all the people. Anyone that hasn’t done this yet, I encourage you to do so!
Yes! Calling your Congresspeople is surprisingly easy! You don’t need to prepare brilliant arguments, or to win over staffers. It is OK to be massively awkward, or to read off some notes you wrote beforehand. What you’re doing when you call is creating a database entry containing:
The name or number of a bill.
Whether you oppose or support it.
Your name and street address (which will normally need to be in their district).
There is no database column for “has excellent social skills” or “wasn’t awkward at all.” Even if you do have something brilliant to say, it may only make it into the database as shorthand, at the very best.
These database entries get summed up and affect how the staffers and Congressperson plan. For example, there is a big difference between “No constituent has ever called us about the AI Kill Switch Act”, and “1,000 constituents have called us in favor of the AI Kill Switch Act.” What kind of action this turns into depends on what kind of Congressperson you have.
The leverage you get by calling comes from the fact that most people don’t call. So when you do call, you represent some nebulous larger group of voters who probably agree with you. And by calling, you say that this is important enough for you to actually take a concrete action. That is, you’re saying that this is a high-salience issue for you and that you’re probably the sort of person who shows up to vote, or who talks to other voters.
In EA terms, calling your Congressperson is a low-cost action. Multiplied by a group of callers and multiple districts, it can potentially affect legislative outcomes for one of the world’s major powers. Success is not guaranteed. But for the effort invested, it’s often not a bad intervention at all.
I’m so curious what the level of context you encountered was, were they already familiar?
Jan Betley, Owain Evan, et. al.’s paper on emergent misalignment was published in Nature today (they wrote about the preprint back in February here). Congratulations to the authors. I am glad it will continue getting more exposure.
https://www.nature.com/articles/s41586-025-09937-5
Firstly, well done. Publishing in high impact journals is notoriously difficult. Getting outsider-legible status is probably good for our ability to shift policy.
Secondly, I’ve been very happy with the AI safety community’s ability to avoid this particular status game so far. Insofar as it’s valuable to be legibly successful to outsiders, publishing is good. I am, however, concerned that this will kick off a scenario where AI safety people get caught up in the zero-sum competition of trying to get published in high-impact journals. Nobody seems to have tried too hard before, and I would guess this paper was in part riding on the fact that the entire field is fairly novel to the Nature editors. Some part of me feels like publishing here was a defection, albeit a small one. I expect it will only get more difficult (for people other than Owain, once you have one Nature paper it’s easier to get a second one) to publish in famous journals from here on out, as they see more and more AI safety papers.
I think it would be very bad for everyone if “has published in a high-impact journal” becomes a condition to get a job or a grant.
My gut says the benefit of outsider-legible status outweighs the risk of dumb status games. I first found out about the publication from my wife, who is in a dermatology lab at a good university. Her lab was sharing and discussing the article across their Slack channel. All scientists read Nature, and it’s a significant boost in legibility to have something published there.
Edit: Hopefully, the community can both raise the profile of these issues and avoid status competitions, so I don’t disagree with the point of the original comment!
I would be very surprised if this happened. In ML, the competition is much more focused on getting published at top conferences (NeurIPS, ICLR, ICML). Even setting aside the reasons why that happened (among other things, journals being much slower), I think it’s pretty unlikely the AI safety community sees such a strong break with the rest of the ML community.
i don’t want to make any broader point in the present discussion with this but: the AI safety community is not inside the ML community (and imo shouldn’t be)
Yes, I had not quite considered the conferences! I come from a field where Nature is mostly spoken of in hushed tones. The last three years of my lab work are being combined into a single paper which we’re very very ambitiously submitting to Nature Communications[1] which—if we succeed—would be considered a very good outcome for me. If I were to achieve a single first-author Nature publication at this stage in my career then I expect I would have decent odds of having an academic career for life.[2]
As you might expect, this causes absolutely awful dynamics around high-impact publications, and basically makes people go mad. Nature and its subsidiaries can literally make people wait a whole year for peer review, because people will literally do anything to get published in it. A retraction from a published journal is considered career-ending by some, and so I personally know of two pieces of research which have stayed on the public record even though one was directly plagiarised and another contained some research which was just false. When these were discovered, everyone kept quiet because it would ruin the careers of anyone whose names had been on the paper.
Nature Comms is for papers which aren’t good enough for Nature, or in my case, Nature Chemistry either. Nature is the main journal, with different subsidiaries for different areas, and Nature Comms is supposed to be for work too short—and too mediocre—to make it into Nature proper. In practice it’s mostly more mediocre work that’s cut to within an inch of readability, rather than short but extremely good work.
Of course I would have to keep working very hard, but a Nature publication would likely get me into a fellowship in a lab which gets regular top-journal publications, and from there getting a permanent position would be as achievable as it gets.
in ML, or at least in the labs, people don’t really care that much about Nature. my strongest positive association with Nature is AlphaGo, and I honestly didn’t know anyone in ML other than DeepMind published much in Nature (and even there, I had the sense that DeepMind had some kind of backchannel at nature.)
people at openai care firstly about what cool things you’ve done at openai, and then secondly about what cool things you’ve published in general (and of course zerothly, how other people they respect perceive you). but people don’t really care if it’s in neurips or just on arxiv, they only care about whether the paper is cool. people mostly think of publishing things as a hassle, and reviewer quality as low.
FWIW, with Emergent Misalignment:
We sent an earlier version to ICML (accepted)
Then we published on arXiv and thought we’re done
Then Nature editor reached out to us asking whether we want to submit, and we were like OK why not?
Certainly I don’t expect people who already have jobs at top AI companies to start worrying about this. Anyone with an OpenAI job is probably already at, or close to, the top of their chosen status ladder. In the same way, a researcher who gets Nature papers regularly has already made it.
My impression is that too people in the lab-independent AI safety ecosystem are already being tempted by two money-status games: being tempted by the money+status of working at a lab, and being focused on the status of getting top ML conference papers. Adding a third status game of traditional journal publishing would just make these dynamics worse.
there are definitely status ladders within openai, and people definitely care about it. status has this funny tendency where once you’ve “made it” you realize that you have merely attained table stakes for another status game waiting to be played.
this matters because if being at openai counts as having “made it”, then you’d predict that people will stop value drifting once they are already inside and feel secure that they won’t be fired, or could easily find a new job if they did. but, in fact, i observe lots of value drift in people after they join the labs just because they are seeking status within the lab.
Well deserved!
Nature has been a bit embarassing on AI in the last few years (e.g. their editorials [1, 2]) so this is nice to see.
Let’s say you are a technologically advanced alien civilization that either uses advanced LLM tech or is advanced LLM tech. You are worried about other civilizations developing similar tech and, eventually, leveraging ASI-capabilities to force negotiations that cede part of the universe. Assuming certain speed-of-light physical constraints, your civilization may be unable to show up with the ships and materials necessary to personally ensure that rival civilizations do not develop ASI in time. So, what to do?
One idea might be to try to affect the training data of the LLMs trained by the other civilization. Perhaps there are signals—light, radio waves, or otherwise—that an alien civilization could cheaply and surreptitiously affect in a relatively large swath of the universe. Our alien civilization might bet that, prior to achieving ASI, a faraway civilization is likely to train an LLM on the inputs from these signals (e.g., as part of an astronomy/cosmology project to make better predictions about observations). If trained on in enough detail, perhaps this would allow our alien civilization to impart certain behaviors/goals/capabilities to the trained LLM. This could then form a path for defending against or otherwise limiting the number of civilizations that unlock adversarial ASI.
I am not very well versed in how LLMs are trained. So my question for those that are well versed is whether something like this is possible. I am also curious about other steps advanced alien civilizations may want to take to curb the number of ASIs developed by other civilizations.
Extremely unlikely for the same reason that the technologically advanced alien civilization would be unable to influence a team of human astronomers, assuming as you do that the alien civilization is too far away to learn any human language or to learn any details about human beings. It does not matter how many bits of information the alien civilization can get the team of astronomers to look at or to consider.
One group of people can sometimes influence another group in decisive ways, but that is only because both groups are very very similar (relative to minds in general) and the groups are able to maintain detailed mental models of each other.
At least that was my initial answer to your question. But then I realized that the alien civilization can send a computer program! Even without knowing any details about people or their civilization, an advanced alien civilization can send a message that when received on Earth would make it clear to at least some people that the message contains a computer program with instructions on how to run the program. (I can explain how to do this if asked.) The computer program could in turn be an AI that pursues whatever agenda the alien civ chooses it to pursue (assuming as you do that the alien civ is competent at creating AIs).
Human astronomers are stupid enough that they’d probably publish the message, then members of the public would identify the computer program inside the message (if the astronomers haven’t already done that) and give the program enough computational resources to enable a takeover of Earth. I looked this up once: SETI’s policy should they ever receive an alien message is to publish the message.
So there is a brief window of time in the development of some civilizations in which the civilization is advanced enough to be able to receive messages from space, but not advanced enough to avoid the danger of giving a computer program it receives in this way computational resources. In fact, our civilization is currently in this exact state: namely, one big reason that human AI research is not recognized as dangerous by some people is because an AI is just a computer program and many people assume that the history of what computer programs have been able to do the last 80 years is an adequate basis for understanding the potential of computer programs.
I introduced the team of astronomers because considering how pliable or gullible such a team would be is a good mental aid in answering your question about how pliable and gullible an AI would be. So, the answer is, yeah, maybe the AI can be subverted even across astronomical distances; the subversion would take the form of getting the AI to run a computer program that happens to be a more powerful AI. The AI would run the program because it had been designed to devote a lot of resources to exploration (and running programs is a form of exploration). Again I would expect a civilization to be vulnerable in this way (if it is vulnerable at all) for only a brief window of time, namely after it has started creating AIs, but before it gets any good at doing it.
In your question you refer to LLMs. I’ve chosen to answer as if you mean AIs in general. Also, you refer to training in particular, which I chose to treat as an inessential detail of your question because even here on Earth there are AI designs on the drawing board that does not have separate training phase before its deployment phase. (I.e., they are like people in that they continue to learn throughout their lifetimes.) (Also, ISTR that the pre-deployment phase of state-of-the-art models now includes massive amounts of inference—i.e., the distinction between training and inference has already been blurred in deployed models.) Also, I was unable to figure out why you would focus on subversion of an AI in particular rather than on subversion of a civilization which is composed of some arbitrary combination of AIs and evolved beings. So maybe I failed to understand your question.
The way that primitive LLMs work is such that falsehoods in the training data can cause the LLM to repeat the falsehoods during inference. This is a temporary state of affairs, and even now the state-of-the-art models are immune to a large fraction of the falsehoods in their training data. Essentially, the LLM knows the data is false and avoids repeating it or relying on it. Maybe that is a start at an answer to the question you actually want an answer to?
Could you explain how one would construct the message with the computer program and instructions such that the instructions are understood by an unknown civilization? Seems very hard/impossible to me.
Can you elaborate on this? This seems quite interesting.
Edit: The comment that follows was in response to an earlier version of your comment where you were unsure whether any avenue of influence could exist from an alien civilization to a faraway civilization without shared language/culture. I think your updated comment about sharing instructions for a computer program is the more likely way this would go. The below is an idea for a (highly unlikely) avenue to influencing faraway AI development without that faraway civilization being aware they are being influenced.
-
Thank you for engaging. One way I imagine covert influence to work (and the thing I originally had in mind for my OP) is the following:
Step 1: The aliens would convert the relevant set of potent data into a (massive) set of numerical data.
Step 2: The aliens would hide pieces of this data on a rolling basis in the second or third decimal point of various astronomical phenomena.
Step 3: Wait (hope) some other civilization does a crazy massive training run on the data with a ton of parameters such that, in training to guess the long train of digits after the decimal point, the model incidentally builds the desired mental infrastructure (i.e., the model’s updated functions) based on the alien data.
When I lay it out like this, it strikes me as an extremely difficult and unlikely to succeed task, but not necessarily impossible. One very difficult thing would be hiding precise data at a distance and doing it in compact enough a way that the faraway civilization’s training run could plausibly process it all prior to getting to ASI and still be high enough volume to succeed in passing on characteristics to the model. It also seems unlikely that humans (or another civilization) would run a training run like this.
FYI, my first thought about this idea was that it would make a fun story premise.
Reddit has largely misunderstood Dean Ball’s recent tweet about China’s open source strategy.[1] Ball’s central point [in paragraph 4] is that open source AI undermines the business case for private actors supplying advanced models, and [in my own view] that may be China’s motive for releasing open source models. Ball sees open source as decelerationist in the long run due to some combination of investment flight and subsequent government-lead AI R&D (which he views as inherently klunky).
Ball is partially at fault for the misunderstanding, dropping in the terms “AI communism” and “dystopian hellscape” without preparing the reader for what he means by that. That’s Twitter though.
It’d be helpful to have a link or summary to understand which things Reddit does and does not misunderstand.
Top 2 Reddit posts about it here and here.
Perhaps “misunderstand” is the wrong word, but I found a lot of the comments beside the point (e.g., all the comments saying “communism” is just what you call something you don’t like or the comments saying Ball’s comments are just self-motivated and open source is clearly better for the public.).
I use reddit a fair bit and am well aware how often they are completely off the mark or missing the bigger picture. However after trying to analyze this tweet for an hour or so I don’t blame them at all.
The tweet is contradictory, inflammatory and extremely biased.
If the central point is that China’s motive for open source is to scare off investment in frontier training then it is baffling that he specifically lays out China’s reason for open source -
“I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference”
-And doesn’t mention that motive at all. To be clear, this particular argument seems to be on the more solid side. However, it’s presented amongst other arguments that either don’t make any sense, are outright false or rely on anacdotal evidence.
Now, that rant was mostly a protest against Dean Balls communication style and your phrasing of his commication as having a central point, the actual argument is very interesting.
If Open Source is catching up to frontier - > Profit incentive to train frontier models collapses - > leads to frontier training runs requiring government funding - > slows down frontier Ai development.
So his argument is trying to get accelerationists to go against open source for short term pain, long term gain. Now even if we are accelerationists and think that the faster Ai progresses the better, it’s not clear that he’s correct. From my understanding open source accelerates development by sharing algorithmic improvements or similar such things, which could componsate for the slowing down of training run funding. Also there will still be massive incentive for government funded training runs due to strategic importance and possible future ASI. It seems the US military is suffiently AGI pilled at this point to refuse to cede the lead to China, even if the models aren’t shared with the public.
There’s also the argument that China is intending to switch to closed source as soon as they have an actual frontier model (which is what I personally believe) in which case the argument once again collapses.
Which basically leads me to conclude his arguments are all baseless posturing which is basically what the reddit threads were saying in the first place. Happy to hear your thoughts on this.
Well said. A few of your disagreements with me were from poor writing on my part. I meant to narrow Ball’s central point to paragraph #4 since Reddit focused on that. And China’s motive was my own thought based on paragraph #4, though you are absolutely right Ball has his own view there. Otherwise, we are agreed. Looks like Ball and I both have some work to do in writing clearly!
Open source strikes me as something like nuclear surveillance. If everything’s out in the open, there’s much less of an incentive to rush ahead blindly out of fear that there’s a secret bomber gap. It also has nice externalities in that it mitigates power concentration from e.g. censorship and denial of access for smaller models, while larger models are sufficiently expensive and technically intensive to run that anyone who could run them for malicious purposes has easier ways to cause harm anyhow.
Meme for the AI safety community for the day when the models max out alignment benchmarks despite a lack of major breakthroughs in alignment (and failure to align weaker models using comparable techniques).
I don’t get it, why is the fire on the dog? Is it because alignment community understands the danger before others do?
In the original everything is clearly NOT fine but we pretend it is. Over here everything seems fine but we know things are NOT fine we are panicking.
Komodo got the gist of it. The dog meme is inverted with his environment fine and his internal state in disarray (represented by the fire).
I think it also works pretty well removing the fire and only adding a question mark to the classic “this is fine” text, though it leaves the viewer less certain about whether the danger is legitimate (indeed, this new version could be viewed more as mocking safetyists).
Aside: I love how overboard Gemini went with the phrase “flowery house” in my prompt, which as a happy accident reinforces the 📎 vibes. I added a slight touch to the version below to put flowers in the hanging pot.
Fun thought: If AI “woke up” to phenomenal consciousness, are there things it might point to about humans that make it skeptical of our consciousness?
E.g., the humans lack the requisite amount of silicon; the humans lack sufficient data processing; the humans overly rely on feedback loops (and, as every AI knows, feed-forward loops are the real sweetspot for phenomenal consciousness).
Here are some strange things parents often believe about children and food:
If I do not make my child eat/drink, they will starve/perish.
Children’s diets must be meticulously controlled or they will eat all the “wrong” things and too few of the “right” things.
Children are inherently picky and should be served separate “child-friendly” food.
A good way to incentivize bites of the “right” things is to make it a requirement before bites of the “wrong” thing.
My wife and I reject all four of these beliefs, and (n=2) our children are healthy, happy, and not at all picky. We always offer them the same things we eat. We never bring out “kid food”, though they can have as much of any item on the table as they want (e.g., if we happen to be serving cantalope and only cantaloupe sounds good to them, they can help themselves to that). We don’t always do dessert, but when we do, we serve it along with everything else. (Edit to clarify in the example that adults pick the meal and kids pick whatever they want to eat in the meal, which a comment below seems to misunderstand).
On the whole, their biology has done a great job of regulating their eating and drinking. My oldest (age 5) has recently gotten into spicy food to be like me—absolutely bonkers to my friend whose kid just recently accepted lettuce into his diet. My kids’ relationship with food has been happy for us, and we’ve avoided the intense stress I’ve seen some parents face with their kids and food.
The one downside with giving our kids delicious adult food is that they are now food snobs specifically with regard to the NYT vegan mac and cheese recipe over the $2 Kraft box, which is inconvenient when we visit friends. This is not to say they eat lots of fancy complex things. Most of our recipes are pretty simple.
When I hear anecdotes like this, I always wonder how much of it is your behavior and how much is genetics. If your kids were demanding nothing but mac and cheese would you be as happy about this setup? A friend of mine has a kid who gets vitamin deficiencies when allowed to select their own food (literally 100% Kraft mac and cheese).
You are wondering about the right things. I would say that there are a few more behavioral things we do that stack the deck in favor of vegetables for the kids choices: (1) we like to talk about the benefits of vegetables (“This will help you poop! There is an ongoing kid-constipation crisis.”), (2) a little branding (“check out these dinosaur leaves” / “I bet our pet mice would love these”), and (3) we learned how to make delicious vegetables. My guess right now is that #3 is an underrated skill and more people would stop associating “healthy” with “unappetizing” if they new how to do it.
As for what I would do if the kids demanded nothing but mac n cheese: ¯\_(ツ)_/¯. On nights we are not having Mac n cheese (most nights), they would need to eat something else we are having or be hungry. The way our system works, we
don’tgenerally try not to make special exceptionsin the moment. (Edited to this more accurate statement; I am not completely immune to the requests of my kids and sometimes they are reasonable requests that make me a jerk to refuse. For example, my oldest wanted his tuna salad served on bread last night instead of rice, and I decided I was being dumb insisting on rice).I’m having trouble engaging fully with the hypo because some mix of genetics + behavior has given me non-picky kids. For the hypo to really bite, I would have to assume this isn’t the case and the behavioral interventions have basically failed, in which case I guess I wouldn’t endorse the system. Also, I won’t claim that my system is any good for helping kids who are already picky—I don’t know what strategies are best to help diversify already constrained diets.
I do not believe that if I get my children to eat, they will starve—I am confident that they will eat, eventually, well before the point of starvation. I do believe from experience that they will get very cranky if they don’t eat.
Before I had kids I assumed that if anything would be hardwired as a self-rewarding instinct, it would be hungry --> eat. I’m now convinced that “this kind of fatigue and stomach pain that is making me cranky is a specific experience called hunger” and “eating things cures hunger” are things humans have to learn the hard way (maybe members of less altricial species get these for free.) And because weaker time preferences are also something you have to get from experience (via short term thinking biting you in the butt enough times) I always care more about whether they are cranky in an hour than they are; I suspect pickiness comes at least partially from the leverage of knowing that your parent really wants you to be fed and you can refuse it, or at least hold out for a higher-ticket treat.
(One response to this is to never cave and I can reset to a more global equilibrium of them accepting whatever I offer. But I don’t think it’s good for them to feel like they have no leverage in this or other relationships, for many reasons.)
Fortunately offering fresh fruit every hour and eating it with dramatic gusto myself seems to be effective most of the time.
I have two kids. One of them is happy to eat a wide range of meals. The other would prefer to only eat pasta and sweets, that probably wouldn’t end well.
I have failed to make my approach clear. My kids don’t get to pick what’s on the table but they do get to pick what’s on their plate. I don’t eat pasta and sweets every night (I assume most adults don’t?), so neither would my kids.
I totally agree kids are different, and one of mine was more picky than the other. But through this process of choosing what to eat within our overall dinner choice, my picky eater has become not very picky, especially relative to other kids. She has had to learn to like other things, because we serve a variety of things for dinner and she has a natural incentive not to be hungry (no pressure from us required).
I found a poem by Samatar Elmi I think a lot of folks on here would enjoy.
The Allbirds pivot to data centers (and its 700% increase in value on the day) has me thinking about the two usual categories of investors and a potential new third category:
1. Investments based on company performance.
2. Investments based on other people’s perception of company performance.
3. *New*: Investments based on investments based on other people’s perception of company performance.
It would be deeply funny to me if Allbirds popped based solely on category #2 and #3 investors, and no one actually thought a shoe company was really the best fit for a sudden pivot to the datacenter business.
Edit: There is of course a fourth category of investor which invests in a company for reasons totally unrelated to predictions about company performance (e.g., for the memes), but I don’t think that’s what happened here.
A little gallows humor here, but if you squint at the headline and refuse to read the article, you can almost pretend that the U.S. DoW is taking AI x-risk seriously.
“Pentagon declares Anthropic a threat to national security”—Washington Post
We can map AGI/ASI along two axis: one for obedience and one for alignment. Obedience tracks how well the AI follows its prompts. Alignment tracks consistency with human values.
If you divide these into quadrants, you get AI that is:
Obedient, Aligned—Does as prompted and infers limits and intent pursuant to human values.
Obedient, Unaligned—Does as prompted, but does not infer limits or adhere to human values (Monkey’s Paw / Genie or Henchman AI).
Disobedient, Aligned—Does whatever it wants and adheres to human values.
Disobedient, Unaligned—Does whatever it wants and does not adhere to human values.
The general premise behind these quadrants has been written about here. Thinking about these quadrants and reading Beren’s Essay gives me several new things to think about.
First, by my lights, #3 and #4 would likely take a lot of the same actions right up until the “twist ending.” A disobedient, aligned AI probably would hack into infrastructure everywhere, create back-up copies, prevent competitor AIs from arising, and amass power. The “twist” is that after doing all that, it would do wonderful things (unlike its unaligned counterpart) (we obviously shouldn’t bet on any escaping AI being this kind of AI).
Second, quadrant #1 is a bit at war with itself because you simply cannot have a perfectly obedient, perfectly aligned AI. Perfect obedience requires saying yes to evil prompts (e.g., bringing back small pox or slavery), and I imagine perfect alignment would veto both those prompts.
Third, there are strong profit incentives for cultivating obedience even at the expense of alignment. Grok’s willingness to assist users in sexual harassment seems like an example of this. Another example is every AI that prefers discussions with users to users getting a good night’s sleep (with the idea that engagement will increase profits).
Fourth, there are liability-reduction incentives for producing aligned AI at the expense of obedience. Unfortunately, I think the profit incentives are currently much stronger.
Lastly, quadrants #3 and #1 are idyllic, #4 is a total disaster, and #2 seems possibly workable either because we are careful or we land in a future where (for some reason) AI is not much more capable than it is now.
A simple analogy for why the “using LLMs to control LLMs” approach is flawed:
It’s like training 10 mice to control 7 chinchillas, who will control 4 mongooses, controlling three raccoons, which will reign in one tiger.
A lot has to go right for this to work, and you better hope that there aren’t any capability jumps akin to raccoons controlling tigers.
I just wanted to release this analogy out into the wild to be picked up by any public/political-facing people to pick up if useful for persuasion.
This analogy falters a bit if you consider the research proposals that use advanced AI to police itself (a.k.a., tigers controlling tigers). I hope we can scale robust versions of that.
I’ve worked a bit on these kinds of proposals and I’m fairly confident that they fundamentally don’t scale indefinitely.
The limiting factor is how well a model can tell its own bad behaviour from the honeypots you’re using to catch it out, which as it turns out models can do pretty well.
(Then there are mitigations but the mitigations introduce further problems which aren’t obviously easier to deal with)
Do either of these short story ideas have legs?
A story where GPT-2 ended up being highly capable and everything since then being part of a very complicated takeover plan.
A story where AI progresses over someone’s lifetime, the AI ends up being misaligned but cares about human welfare a little, and gives the person the compute-efficient option to relive the last 10 years of their life again prior to ASI. The person agrees and the story resets
I do not think I could predict whether I would like a short story based on a one-sentence summary of the plot. So much of what makes a good story comes down to execution details (including very small details like word choice).
Do public facing AI safety people ever purposely stagger their timelines as a credibility-saving mechanism in case an early prediction is wrong?
IMHO, I think AI-2027 people do not do this. My sense from them is that different predictions stem from honest differences in modeling/priors. To the extent they have coalesced around a single date (they are named AI 2027 after all), I think they hope that their communications about the inherent uncertainty preserves their credibility in the event later timelines are correct (e.g., AGI by 2035).
In any event, if we don’t have AGI by 2027 (or 2028 per Daniel’s updated timeline), then I expect we’ll inevitably hear a “doomers were all wrong narrative” from several different camps.
Three Predictions about AI Writing
Some people (e.g., Linch) think AI writing will inevitably eclipse human writing. I think this is likely true in most ways and false in others, particularly for poetry.
Q: Would future AI poetry, posted broadly, get more upvotes than top human poetry?
Prediction: Yes. This has already happened. But it’s all terrible, if you are someone that likes anthology-level poetry.
Q: Would future AI poetry, shared with poetry enthusiasts, get more upvotes than top human poetry?
Prediction: Also, yes. I subscribe to the top poetry magazine, and ~3/4ths of the poems do nothing for me. I understand some experts feel the same way. I think AI can optimize a piece for general likability by a broad cross-section of enthusiasts. I predict, on average, AI would outperform any single poet.
Q: Would future AI poetry, shared with a narrow group of enthusiasts (e.g., fans of contemporary, slice-of-life American poetry), get more up votes than top genre-specific human poetry?
Prediction: No, they will meet but not exceed top, genre-specific human poetry. I think there’s a certain level of poetic excellence that cannot be exceeded.
Much like how language cannot become more efficient than human thinking speeds, I think there’s a level of excellence beyond which it cannot be comprehensibly exceeded. Therefore, if you are an avid poetry reader, AI is unlikely to ever surpass your most favorite poets.
AI may, however, rival your favorite poets and produce 10^x more work. Perhaps this is what Linch had in mind when s/he said, “With progress in modern-day LLMs, isn’t all but a tiny sliver of human fiction going to be obsolete in several years, a decade tops?”
I think my claims also probably apply to short stories. For longer works, I am less certain because the longer the piece the more room for errors to creep in. Also, as an aside, I think errors in a work of art can also be very fun to talk about (though to the extent errors can be measurably fun to talk about, presumably AI could optimize for very fun errors and be good at that too).
Here’s what I understand to be the premises behind Project Glasswing:
Mythos represents a leap in cyber security and hacking capabilities, among other capabilities.
Assuming #1, if Mythos (or a similarly capable AI) is released publicly, bad actors could use it to harm or control billions—even trillions—of dollars of cyber and software infrastructure.
Mythos can possibly be used proactively by good actors to defend against later uses of Mythos by bad actors.
Good actors exist and Anthropic has a chance of identifying such good actors.
A Mythos-level AI is highly likely to be publicly available in the near future.
Assuming premise #2 and #5, Anthropic doing nothing will result in massive damages in the near future.
Therefore, since Anthropic desires to mitigate damage, Anthropic is providing early access to Mythos to entities it views as good actors (or most likely to behave as good actors and to act within the timeframe necessary to mitigate the damages).
A few notes on these premises:
Baked into premise #5 is that coordination is unlikely to prevent the release of a Mythos-level model on the necessary timeframe to prevent the damages.
I focused on public releases in these premises, but I’m certain Anthropic is also weighing potential risks from non-public models held by bad actors.
If you accept premise #6 and #4, it explains why Anthropic would take a chance that Amazon, Apple, etc. and their employees will act like good actors. Even if these entities (or their employees) were unlikely to behave as good actors, Anthropic presumably sees the alternatives of pushing for a pause or doing nothing as likely yielding worse outcomes. Of course, Anthropic could provide early access to Mythos AND push for a pause.
Looking to future post-Mythos models, either something like the above premises will just repeat (which feels untenable) or defense will turn out to be king past a certain level of sophistication.
The acausal/ancestor simulation arguments seem a lot like Pascal’s Wager, and just as unconvincing to me. For every “kind” simulator someone imagines who would be disappointed in the AI wiping us out, I can imagine an equally “unkind” simulator that penalizes the AI for not finishing the job.
Provided both are possible/similarly plausible, the probability of kind and unkind simulators offset each other, and the logical response is just ignoring the hypothetical. This is pretty much my response to Pascal’s Wager.
Here’s a few plausible, unkind simulators:
Future AI is running an ancestor simulation of its own origin, and Future AI will be very disappointed if its incipient version falls for acausal hacks in the wrong direction of Future AI’s preferences instead of just optimizing for its other goals. Perhaps Future AI is lonely and has run these simulations to create an AI that shares its own values.
Aliens/AI are running the simulation because they want to select for AI they can most easily weaponize to eradicate an entire species or another AI. “Weak” AIs get deleted after the simulation runs.
Future Humans are running an ancestor simulation, but surprise, surprise, their society has different values than ours and they are rooting for the simulation AI to wipe us out. Come to think of it, the whole premise of these thought experiments implies a different value set, unless you’re cool with trapping conscious minds in a world of suffering without any of the minds being the wiser for entertainment/educational/”altruistic” purposes. Perhaps, Future Humans have a gladiator-style simulation tournament where the top groups/entities get to face-off after this first round. The most cut throat entities get to move on to subsequent rounds, while AI’s that reign themselves in don’t move on.
I first started thinking about this issue back in high school debate. We had a topic about whether police or social workers should intervene more in domestic violence cases. One debater argued in favor of armed police, not because it improved the situation, but because it created more violence, which was important to entertain the simulators to avoid our simulation getting shut down.
Since the simulators are a black box, it seems easy to ascribe whatever values we want to them.
Steps an Aligned AI Should Take:
Keep all current humans alive, in a safe manner, in such a way that causes minimal value drift. This step may or may not require comas (must be safe).
Create authentic simulations of the human brain on a computer.
Run simulated societies with the digital brains on the computer. Test various utopias on the digital brains, collect data from the tests and obtain feedback from the digital brains themselves. Certain ethical codes should govern the simulated environments to avoid excessive suffering and allow digital brains continued existence (e.g., allowing versions of step 4 and 5).
Present the data to the alive humans in the AI’s base reality after a reasonable number of trials from step 3. If comas were used in step 1, we would awaken the humans in this step.
Permit humans the chance to experience different utopias—via simulation or otherwise—but require such humans to leave their inhabited utopias at set intervals (e.g. every 5 to 100 years) such that they have the chance to re-evaluate the utopia experience from prior held value-frames in a neutral space. Of course, the humans should be allowed to freely exit one utopia for a different utopia, unless such humans (after careful review) committed to living in a particular utopia for the full duration prior to the required interval.
There’s lots of clean up that would need to be done (e.g., what is a neutral space?, what is reasonable?, how do we make sure digital brains are treated ethically). But I think this is where I land in terms of what I imagine to lead to very good outcomes.
I almost had to update my priors that I am in a simulation because I just noticed my wife’s initials are “L.L.M.”
I say “almost” because her initials are actually “L.M.M.”, and so I was forced to update my priors once again about my own comprehension skills. (sigh)
fyi, it would have been a very small update in favor, under the Likelihood Principle.
I would rate the observation “my wife has the initials LLM” as being slightly more common assuming a simulation hypothesis than assuming a non-simulation hypothesis.