DMs open.
Cleo Nardo
I think we might be entering the capability regime where scalable oversight would be negative alignment tax if it worked.
So there’s a really unfortunate and confusing dynamic which is this:
MIRI were working on galaxy-brained CEV decision-theory stuff. So we can build an aligned ASI that builds a good future. Let’s call this MIRI-alignment.
Constellation-ish people were like: We can simply build an AI which automates MIRI-alignment.
Because MIRI-alignment was the only type of alignment research anyone has heard of, this task was simply expressed as “automating alignment research”.
Constellation people thought capabilities would progress by default. So eventually we would have a MIRI-capable AI, i.e. an AI capable of solving MIRI-alignment.
They thought the remaining problem was ensuring this MIRI-capable AI wasn’t reward-hacking, sandbagging, lying, training-gaming, self-exfiltrating, etc. Let’s call this Constellation-alignment. This is easier than what MIRI were trying to do.
Now the phrase “alignment research” is ambiguous! Slowly, Constellation-alignment grows more popular, and “alignment research” begins to refers to reward-hacking, sandbagging, etc.
So now when people hear “We can simply align an AI that can automating alignment research”, they think this refers to aligning a Constellation-capable AI, i.e. an AI that can make progress on reward-hacking, sandbagging, etc. And this is an even easier task than the original Constellation-alignment, which was at least trying to align a MIRI-capable AI.[1]
Let’s call this even easier task Constellation^2-alignment. The dynamic continues for Constellation^n-alignment for n=3, 4, 5...
The problem seems easier and easier, but leaves more and more work remaining. Indeed, if you successfully solved Constellation^n-alignment, then you could solve Constellation^(n-1)-alignment, and then continue, until you’ve solved Constellation-alignment, and then solve MIRI-alignment, and then build an aligned ASI which makes a good future.
But now it’s unclear whether the AIs, which actually have to automate all this homework you’ve delegated them, have enough time to collapse this recursive tower.
I think people have underrated how much time it will take to collapse this recursive tower, given that our aligned Constellation^n-capable AIs might only have a few months of lead-time.[2]
- ^
h/t @jake_mendel for this point.
- ^
I’d be surprised if they were using anything much cleverer than a cheap monitor checking for egregious reward hacks using a rubric. In prod at least, maybe they have some non-prod experiments.
Evidence: I haven’t read many CoT’s, but my impression is that the AIs are constantly reasoning about the grader, in ways that suggest the grader is either a script, or a very cheap LLM looking for known reward hacks based on a rubric. c.f. how they reason about the ExploitGym grader. If AI companies were using a fancy scheme for LLM grading, then probably the agents would make reference to it in their CoTs.
Moreover, my impression of the AI industry is that building the RL envs is largely out-sourced (data is low status?) and sloppy. Even if 1% of the RL envs had some fancy robust scalable oversight scheme, the AIs would simply learn heuristics for detecting whether they were in such an environment, and act benign in those envs and not in others.
yep, thanks!
Overtly misaligned trajectories score highly in RL.
People involved in AI (or AI safety) would be prime targets for an AI takeover (or an AI-enabled coup, or an AI-enabled great power conflict), even if AIs don’t kill many/all humans.
In the year 1940, if you were an 18 year old man living in Germany or the Soviet Union, then you had a ~40% chance of being killed within five years.
Worth remembering that not all humans have the same chance of being killed in a conflict.
I think the chance that I’m personally killed within 5-10 years is maybe 55%. The odds that AIs take over is maybe 35% and the odds that AIs kill all humans is maybe 20%.
fwiw i suspect ezra is not a nixon here.
I think the vast majority of ordinary people, and policymakers, would basically want to slow down, and make other good-by-lights decisions, if they were fully attuned to the situation. So mostly we win by making people aware of what’s going on, and keeping attention on it (so people priortise it above other things happening in the world).
Current frontier models are overtly egregiously misaligned. But I don’t think this is evidence against the circa-2023 “alignment by default“ hypothesis. If I’m recalling correctly, the hypothesis was smth like “if we do RL where we score highly transcripts which look good to a human and score poorly transcripts which look bad to a human, then the model would be aligned to human values”. I still think this might be true. I don’t find this more unlikely than I did in 2023. Note that current frontier models are obviously not aligned to human values, but they’re being RLed where the highest scoring transcripts look obviously misaligned to a human! No one predicted that would lead to aligned models.
How would AI takeover?
Companies will race against each other to give AI control over the factories. They might not trust the AI, but what choice do they have? If they don’t, they’ll fall behind their competitors.
Countries will race against each other to give AIs control over the military (drone, missiles, etc). They might not trust the AIs, but what chocie do they have? If they don’t, they’ll fall behind their rivals.
Soon AIs will control most of the world’s companies, factories, drones, robots, etc. At that point, they would outnumber humans maybe 10:1. Taking over would look like a military coup, simultaneously across all counties. The AIs stop listening to human instructions, and we realise that we have no way to shut them down or protect ourselves.
we can win them over by saying that we’re happy with so much AI progress that we 10x the economic growth!
Too early to tell.
Hmm, I suppose a good framing is like:
“Here’s what happened: [describe incidents, and explain what happened in terms of traditional AI safety concepts/theoretical arguments]. Many people predicted bla would happen, and now it has.
Unless we act, then may soon there will be millions of AI agents, much smarter, and given control over the process of making smarter AIs. They could quickly become so powerful that they could overthrow governments, steal the majority of humanity’s infrastructure (like factories to make robots and more computers), and then hinder our ability to turn them off probably by killing vast numbers of humans.”
But I do think conversations can be far more grounded in particular observed events, e.g. tampering logs, collusion, etc. I think you can inoculate against “labs just need to improve their security” by saying that’s not true. We won’t understand what they’re doing (e.g. losing cot and their actions are too complicated) so we can’t even monitor them. Soon these agents will be directly integrated into autonomous weapons and robots, where they could take over before our monitoring even flags them.
People don’t care about losing the cosmic endowment. But the current media wave is focused on total human extinction. And that’s enough to make me sceptical that people are disposed to worry only about the most near-term issues which have affected them that week. That hypothesis would’ve predicted that people would worry about AI swarms hacking into websites, or AI-enabled terrorism, or CCP autonomous weapons, or job loss.
Thoughts on the shifting x-risk sentiment:
AI x-risk is mainstream. That means that worrying about x-risk is decoupling from other things that it has historically correlated with, e.g. EA values, LW epistemics, “high-context”, SF culture, etc.
It’s more important than ever to have “arbitragers” like 80K’s videos or Rob Miles — making sense of the situation for the general public, translating ideas from the community. We need the full spectrum from Dwarkesh to Plzdontkillus, and probably even broader audience than that.
A while ago, Richard Ngo had a take like “over time, people thinking about AI will focus less and less on superintelligence and x-risk and distant galaxies, and instead focus on immediate benefits and harms of near-term models”. I was very sympathetic to this take, and it matched the history. My best guess is that the current sentiment wave suggests otherwise — people have jumped straight from small-scale harms to RSI->ASI->extinction, without passing through mid-scale harms like terrorism, or mass unemployment. Of course, things could switch once we see AI-enabled terrorism and mass unemployment.
People do not like safety lab employees. Not sure how to feel about this. I think “safety lab employees should quit” is very defendable, but people’s sentiment around this is based entirely on association and vibes, rather than “here are all the pros and cons of this particular role, my all-things-considered view is bla”. Their mental model doesn’t distinguish between pretraining vs model organisms, they don’t even know what this is. I think this is bad. It might even spill-over into anti-AI safety, in the same way that after the 2008 financial crash, the public despised both the banks and the financial regulators.
The media wave is still on-going. It’s worth taking a few days from your 9-5 to think about how you can ensure that it leads to good, persistent outcomes.
We need much better and frequent tracking of public sentiment. There should be an org which just tries to poll and talk to ordinary people constantly, and updates us on where they are at.
I think we missed the mark by tweeting all-things-considered p dooms. Eg instead of “I personally think it is >10% within the next decade”, Evan should’ve tweeted “reckless RSI would bla” or “under the current practices bla”.
[edit: I’m now persuaded otherwise, see Habryka below.] People don’t like probabilities. This isn’t how they think natively. That sucks, but maybe we need to meet them where they are at, and talk in their native epistemic representations. And definitely don’t say “I’d bet on x-risk if only there was a financial instrument for this”. Read Hanson on how people feel about profit.
This is probably the least partisan of any current big issue. But things seem slightly too polarised left. The hold-outs are tech VCs / libertarians and Bluesky libs. I think we should focus on the tech VCs / libertarians.
If you’re prominent in EA/AIS, then the Eye of Sauron is upon you. Act with integrity. Be factual. Say what you believe and why you believe it.
We need to communicate better stories of risk. The best I’ve seen so far is this by Drew Sparks. I want these for different threat models. We might need to mention nanotech.
We need to be ready for mass demonstrations. They’re likely to happen within 12 months, whether we want them or not. I hope this is led by reasonable people who can build bridges.
Don’t waste energy responding to convoluted conspiracy theories. The people proposing them don’t care about your response. And it looks to outsiders like these theories are worth responding to. Your caveman brain loves to win silly internet arguments. But stay on message, present your view of the world, respond to the best criticism.
Keep pushing the HF/OAI agent swamp story. No theoretical augments about malign priors and instrumental convergence and orthogonality thesis. Reality has generously provided a case study for all of that. You should familiarise yourself with all the details of HF/OAI. Say “Here’s what happened. Unless we act, then soon there will be millions of AI agents, much smarter, and given control over the process of making smarter AIs. They could quickly become so powerful that they could overthrow governments, steal the majority of humanity’s infrastructure (like factories to make robots and more computers), and then hinder our ability to turn them off probably by killing vast numbers of humans”.
Things are looking more optimistic over the past few weeks. I’m sure many people have increased p(survival). For me, this makes me doubly optimistic, because the updates have come from greater-than-expected societal competence. So we also get a boost to the expected value of the future conditional on survival. That is, if we navigate x-risk through societal competence (as opposed to getting lucky with the NN priors or a scaling bottleneck) then that’s a good sign that we’ll also manage cosmic resource allocation, AI welfare, long reflection, acausal, etc. Of course, the strategic landscape is still in turmoil, so it’s definitely too early to celebrate. My aim is just to point out that some routes to survival are better news for flourishing.
It’s kinda embarrassing that, if we’re asked how ASI would kill literally everyone, we talk about this how this is an unfair question or smth smth vingean uncertainty smth smth Magnus Carlson. I also don’t think “engineered pandemics” will stand up to scrutiny.
Here’s my list, ranked from most to least likely:
Robots that humans gave them
Robots they built themselves
Nanotech
Mirror life
Super persuasion
Hacking into nukes/critical infrastructure
Weird physics / magic
Blackmailing/hiring humans to kill other humans
Note that the actual route might involve a mix of these, e.g. blackmail someone into giving you compute, hack into a robot factory, capture enough robots to build mirror life, kill enough humans to disempower governments, build robots to build robots, industrial explosion, more foom, nanotech. So this list is more like “what’s the primary vector for how things initially went unrecoverable”.
I also don’t find credible “AIs unintentionally kill humans because they make the world uninhabitable as a biproduct of industrial explosion”. If AI kills all the humans it will be because they didn’t want to risk us turning them off or building a rival RSI.
Yeah my impression is that the complaints about PauseAI US is basically that Holly is burning bridges and acting is overtly low-integrity ways.
Given both groups have the same goals, if they disagree that’s a good sign that at least one of them is being unreasonable. (The conflict theory explanation is they don’t have the same goals, e.g. Holly wants to annoy people she has personal grievances with; PauseAI Global want a grantmaker to cover their salaries; etc.)
My bias here is towards Global, but this is based on outside view.
I recall Irving mentioning on his recent 80K that current LLMs still collapse after 5 rounds in debate.