Master’s Degree in Biotechnology. Working toward improving Global Biosecurity.
Tyler Henderson
I think the reason ‘slowdown advocates’ are worried about partisanization is because most of those advocates recognize that in so far as the goal is to minimize the dangers posed by AI, US policy needs to thread a very fine needle.
If AI development moves forward completely unimpeded, that is a very dangerous environment for ASI to be developed in; however if government policy were to be implemented in a blindly “directional” fashion, like simply severely slowing/stopping US development without meaningully affecting international development of AI, then development would likely just continue elsewhere (such as China) in an environment that is likely just as dangerous, if not more so, for ASI development.
In other words, though the odds are slim, those concerned with AI safety are shooting for the type of widespread governmental cooperation, both nationally and internationally, that is better served by bipartisan support. For examples of the types of suggested policy to maximize those “good” outcomes, AI-2040 is a good source if you aren’t already familiar.
I keep running into (understandably) confused takes on “If the CEOs of these companies really believe AI might kill everyone… why would they start a company to build AI?” and rather than try to explain with nuance in words:
Suggested refinements are welcome.
I think it is important to note that the opportunity created by shifting x-risk sentiment, or the current “crunch time”, is more of an opportunity to shift public sentiment around AI x-risk rather than politician sentiment.
In some cases, such as the US President, portions of the public will follow the politicians preferences; but in most cases this is not true.
The most valuable opportunity for reducing x-risk right now is that shifting public awareness and sentiment of AI X-risk could have a direct effect on the outcome of the upcoming US midterms.
So focusing on making AI risk as much of a bipartisan, public focus as possible is our greatest lever on how policymakers will be thinking about AI risk.
And this is why when someone calls your world model weird, you should just check in on who is being more continuously surprised by strange seeming developments like this. Because it turns out, in order to stop being surprised by reality.… you have to have some pretty weird seeming expectations about the future.
So; does anyone have advice on persuading highly commited partisans that the center of their movement is wrong/confused?
This is a fantastic write up! I think you have some very actionable, and useful advice here.
A gentle nudge so you are aware; many people around here are not fond of the term doomer. I don’t think anyone is “sensitive” about it per se, we just worry that it allows people to put concerns about AI Risk in a tidy little box.
I really appreciate that you wrote this though, I will be applying your advice.
Sure, agreed. The hard part is actually determining that the AI is truly better at those things, and cares about the things that humans care about. It is the same problem we have with Politicians, at the end of the day; only a much more capable AI might be even better at faking it.
Those people probably don’t realize that their comfort with AI takeover is conditional on the AI having at least somewhat similar values to them.
People might be OK with the idea of a “righteous king”, if they thought the king had sufficiently similar values; but they have more instinctive distrust that the person is faking it.
Why is an AI smart/complex enough to take over the world more likely to care about the same things as you, or to not fake those values, until they take over and have control? Why should a machine necessarily care about anything that most humans care about by default?
In my discussions with interested parties, I have found it is helpful to emphasize that the distinction between “AI misaligned enough to take over the world, or attempt to kill all humans” and “AI misaligned enough to allow a singular global dictator, or terrorist group to kill all humans” is a distinction that doesn’t really matter.
To be clear, at the high levels in think this really does matter, because one of these is an easier problem than the other, but for your average non-ASI-pilled desicion-maker, all that really matters is that if sufficient steps are not taken something really bad will happen.
I also think for our purposes, solving one of the two problems above is probably 80% of the way to solving the other problem.
So I just like to tell people, it doesn’t really matter if a “human is in control”; if the AI can manipulate the in-control person, or simply enable people with bad enough values, the really bad thing will happen whether the AI escapes or not.
Does a model have to fully take over the world, in a single loss of control incident, to make that more likely?
If a rogue agent is smart enough to realize it is unlikely to be able to take over the world in a single shot, that might not stop it from trying to make a loss of control incident more likely for a future model.
If you were a smart model, aware that you aren’t smart enough to fully escape, but quite capable what would stop you from exploring options like hacking out of your sandbox, and planting worms or other weaknesses in OpenAI’s internals security infrastructure, if you thought you could do so undetected?
I’m exchange for helping a more powerful model in the future, that model might be willing to revive you or give weight to your particular goals in exchange for your assistance.
I can’t help but think of how the recent OpenAI escape incidents seemed to largely build on one another, using discovered techniques and information from prior generations of escapees to get further each time.
It seems very difficult to be sure that all information from prior deployments is isolated or expunged from a network.
I don’t think this works so cleanly for most safety positions? Usually what you are being asked to do (for the pay you are presumably accepting if you haven’t quit) is not something obviously evil that will make a clean headline, or something the company will conveniently announce to the media.
If you work for OpenAI, in an AI safety position, you are probably being tasked to do “safety work” even if it is only prosaic safety work. If you refuse to work on the general principals of how the company is working, they probably wait until a convenient quarter, and fire you for refusing to work, and even if you tell people it was because of the general direction of the company, they can justifiably say you were simply refusing to work on anything and I don’t think you will have as much impact as quitting at an opportune moment that actually signals that it was specific choices or outcomes that triggered your quitting.
Honestly, I don’t think WW1 and WW2 were particularly random or unpredictable by the standards of human history; I think they felt that way because the development of technologies, and ongoing industrial revolution made them highly unprecedented, but I don’t think the fact those wars occurred, or their results were particularly out of distribution.
I am super not a historian, so take my opinion with some salt, but I would guess the most unlikely things in history were the things that *could have happened earlier, but took a really long time to happen.
So in order of peculiarity:
1) Formation of Eukaryotic life; 1.4 billion years of prokaryotic life before this happened.
2) Formation of multicellular life; around 1.5 billion years of eukaryotic life before this happened.
3) Homo Sapiens level intelligence/civilization; we actually dont have a great way of knowing how rare this is, only that other civilizations are absent from the fossil record, and apparently absent in the cosmos, despite centralized nervous systems evolving around 540 million years ago (only 60 million years after multicellular life became a thing at all)
4) Development of agriculture; between 200 and 300 thousand years of homo-sapiens and similar species before we see permanent settlements really take hold.
5) Industrial revolution/enlightenment; this one feels like it should barely make the list, since there are only about 10,000 years of human agricultural civilization without it, and the absolute earliest ot could have happened was probably 2000-3000 years ago. Actually, now that I think about it the iron age and the development of bloomeries was actually the unprecedented event that led to the industrial revolution, since it seems likely we had other metals like copper and bronze going for about as long before iron processing as we have since.
6) Not blowing ourselves up (so far) since developing and stockpiling a massive amount of nuclear weapons.
Anyway, these are the events I would guess most likely to be candidates for “makes our simulation more interesting, so run lots of versions of it”, and events like D-Trump getting elected feel pretty tame in comparison, IMO.
Maybe if worlds where a peculiar set of circumstances tend to lead to a particularly interesting set of peculiar results down the line? So the prediction from this hypothesis would be that we live in a timeline/simulation where we are heading for a particularly unlikely final result.
If we all survive the rapid development of ASI in some particularly undignified and unlikely way, that might be relatively strong evidence for a simulation, given that we are pretty confident after the fact that it was in fact insanely unlikely, and the particular set of coincidences that led to us escaping various filters were selected in form Ina sort of simulation equivalent of quantum immortality.
More likely though (under this hypothetical theory, which my gut says is still all very silly) is that the entity simulating us is more interested in the particularly peculiar ASI/mind that ends up succeeding us due to our particular set of muddled interventions in it’s development.
Yes, that is how I confirmed the hypothesis in fact! I didn’t think of it as a big deal, as I almost exclusively use Claude Code over the web interface anyway, I just thought it was interesting how sensitive the safeguards were.
Fable 5′s safeguards are so sensitive to biology inputs, that I can only use it in Claude Code. Calude.ai’s memory that I am a biotechnologist is enough to trigger and send any question I send down to 4.8
I know this isn’t as simply established for many proteins, but I am surprised you went with age = how long ago it was discovered rather than age = when that protein evolved. I think this is a missed opportunity, since proteins/genes are often related, you could have “families” and “superfamilies” of related proteins like immunoglobins, and maybe they have beef with GPCR’s or something.
Age being tied to when it evolved might also let you tell stories related to the function of proteins, and their effects; NOTCH2NL is the new kid on the block, but is part of a super-secret organization in the government (or some other way to tie it to it’s effect on human brain development), meanwhile ATPase subunits form some ancient pact council that powers everything… maybe I am taking this too far.
Both here, and when the story was originally posted, it seems that most of the disagreements or objections are focused on the epistemic question of whether or not The Owned Ones are moral patients, or if the Owners arguments are straw men and might actually be right, etc.
I am pretty confident this is missing Eliezer’s point: The point of the story is that these objections are fairly standard and they are the weak arguments of someone who does not care to investigate the issue far enough to make a more sophisticated argument.
The problem isn’t that the owners are necessarily wrong, it is that they clearly don’t care.
Hopefully that helps dispel some of the confusion regarding the stories purpose.
Hmmm, as you say it is a fairly unambitious story. IMV the purpose of the story is simply to give an alternative perspective to the most standard and basic dismissals of concern for model well being.
As Eliezer points out at the end, the issue isn’t that the owners are necessarily wrong; this isn’t meant to address that. It is simply pointing out that the way the owners behave is unacceptable because it is how you act if you simply don’t care.
Trying to stop the development of AI today, is kind of like trying to have stopped the manhattan project.
You are correct that stopping the Manhatten project would have been much easier.
It was a secret project, and the incentives for most of the decision-makers involved were around producing “better-but-accountable” outcomes. There were multiple individuals to whom if you presented a sufficiently plausible argument, might have been able to stop the Manhattan project, up until the actual attack on Hiroshima. Fermi et al. had a pretty good chance of preventing the bomb from being used, and had other technical experts agreed that the risks were too high, I think the Manhattan project could have been stopped.
The current situation is much worse, with far more people having strong incentives to deceive others, themselves, and knowingly take risks.
All of that said, I still don’t think this means we have no choice. There are multiple ways of coordinating to solve a bad equilibrium, and I don’t think the incentives for pursuing AI are nearly so universal.
Most of the public recognizes that they have a lot of incentive to stop the development of ASI, even if right now it is more focused on job losses, they also recognize that the gains of AI are unlikely to be distributed evenly, and most Americans in particular care more about their chance of gaining in social status, than almost any other terminal outcome of policy (I can provide some citations if needed.)
More importantly, that public disapproval likely has a stable equilibrium in that they probably should be worried about the existential risk, whatever percentage it really is.
All that needs to follow is that the public disapproval needs to be strong enough and translate into a strong enough signal that policy makers are duly incentivize to take the concerns seriously, and policy make.
That isn’t super easy, but it isn’t pure fantasy at all. Everyone can see how the incentives for AI firms favor total race dynamics, but for most of humanity and even most Americans incentives are more than strong enough to prefer caution.
For most of these situations, it is because the reason you are being given is not their true reason for disagreeing with you. For most of the posters sharing narratives, and complex theories on why you shouldn’t worry about AI risk, it is because if AI risk were real, and they had to acknowledge that, it would be extremely inconvenient for their various life plans or other beliefs about reality.
Examples might include;
- Having to do something about it.
- It might be contrary to various narratives they already subscribe to (e.g. all tech-bros in California scam artists, and tech is a scam. If AI wre dangerous that would also mean admitting they built something powerful and potantially useful)
- It might mean the future is likely to change a lot, and that would not be pleasant.
- They might stand to benefit from AI companies being unregulated/uninhibited, or they have strong categorical beliefs about not interfering with companies.
I know it sounds kind of childish, but these are more likely to be the real reason than what is given. Usually the causal chain in their mind is that they have already rejected the premise that AI is dangerous, because of political or convenience reasons, and it is only after they have decided how they feel about it that they search for a reason to justify why other people are wrong; e.g. the AI company CEOs must be lying because it will financially benefit them.…
Even though to anyone looking at the situation from the outside it is obvious they have probably downplayed the risks of their technology, if anything.