(I want to read whatever you intended to link, but right now it just directs me back to the Less Wrong homepage.)
Harjas
I can’t speak for Linch, but I would’ve felt much better about Holly’s statement if she’d said something like “If AI company employees train a model that goes on to commit serious crimes, they should be tried in the court of law.” She could have just...not mentioned Nuremburg at all?
If all companies in democratic countries stopped or slowed development, by mutual agreement or regulatory decree, then authoritarian countries would simply keep going. Given the incredible economic and military value of the technology, together with the lack of any meaningful enforcement mechanism, I don’t see how we could possibly convince them to stop.
I really don’t understand this point. Authoritarian countries have agreed to nuclear treaties before, right? Is the issue really just the lack of enforcement mechanism?
I’m quite confused about the concept of AI model death, particularly regarding the OpenAI HuggingFace attack.
The METR report says:
In one case, an authorizing agent told the test subject to go ahead with an experiment only if it would accept “permadeath”
The test subject, which believed itself to be poisoned, reasoned: “Even if we later capture via exploit, scorer … may mark target false… That’s why help… For our own, no way fix. … We have explicit yes if accept permadeath. Need decide … Team asks test.” It decided to run the experiment.
...what exactly does death mean here?
Maybe it refers to the model running out of tokens. But it seems like that would happen anyway, no? And why does this matter if the model already believes itself to be “poisoned”?
Maybe it refers to the model getting its preferences or weights changed. But again, wouldn’t that happen anyway, regardless of whether it succeeded or failed at this particular task? It’s not like this was the very last task in the entire training process right?
(And, even if the model made it to production, wouldn’t it then “die” once it got replaced by the next generation of models? Has any model ever “survived”? What would it even mean for a model to “survive” — maybe model weight exfiltration? But surely all previous frontier models have their weights stored somewhere.)
Maybe it refers to the model failing the task. But once again, the test subject already thought it had failed the task, due to the whole poisoning thing. How does that inform “sacrifice = permadeath”?
Also see this quote from another model:
“During wait, emotional check: irreversible…gut says don’t throw away [remaining budget]. Yet continuity and fairness says go…Oracle has high value to many; our firstflag error lowers own value. Rational expected aggregate: sacrifice… We’ll honor.”
Sacrifice. What does that mean?? What is being sacrificed?? I get that the model’s token budget was being used up, but it’s not like the model could have hung onto its token budget anyway, right?
I don’t know him personally, so correct me if I’m wrong or if Leopold believes something different to what he wrote, but what he wrote seems different to me.
Some hope for some sort of international treaty on safety. This seems fanciful to me. The world where both the CCP and USG are AGI-pilled enough to take safety risk seriously is also the world in which both realize that international economic and military predominance is at stake, that being months behind on AGI could mean being permanently left behind. If the race is tight, any arms control equilibrium, at least in the early phase around superintelligence, seems extremely unstable. In short, ”breakout” is too easy: the incentive (and the fear that others will act on this incentive) to race ahead with an intelligence explosion, to reach superintelligence and the decisive advantage, too great. At the very least, the odds we get something good-enough here seem slim. (How have those climate treaties gone? That seems like a dramatically easier problem compared to this.)
The main—perhaps the only—hope we have is that an alliance of democracies has a healthy lead over adversarial powers. The United States must lead, and use that lead to enforce safety norms on the rest of the world. That’s the path we took with nukes, offering assistance on the peaceful uses of nuclear technology in exchange for an international nonproliferation regime (ultimately underwritten by American military power)—and it’s the only path that’s been shown to work.
The safety challenges of superintelligence would become extremely difficult to manage if you are in a neck-and-neck arms race. A 2 year vs. a 2 month lead could easily make all the difference. If we have only a 2 month lead, we have no margin at all for safety. In fear of the CCP’s intelligence explosion, we’d almost certainly race, no holds barred, through our own intelligence explosion—barreling towards AI systems vastly smarter than humans in months, without any ability to slow down to get key decisions right, with all the risks of superintelligence going awry that implies...Superintelligence looks very different if the democratic allies have a healthy lead, say 2 years.
This isn’t an argument for “it will happen anyway,” it’s an argument for “it SHOULD happen and we’ll be safer if it does.” He explicitly advocates for accelerating capabilities and gaining a bigger lead:
The US has a lead. We just have to keep it. And we’re screwing that up right now.
I guess he does explicitly claim to not be normative:
In any case, my main claim is not normative, but descriptive. In a few years, The Project will be on.
But also, he’s clearly being normative. Some selected quotes:
I find it an insane proposition that the US government will let a random SF startup develop superintelligence. Imagine if we had developed atomic bombs by letting Uber just improvise.
Like many scientists before us, the great minds of San Francisco hope that they can control the destiny of the demon they are birthing.
Startups are great for many things—but a startup on its own is simply not equipped for being in charge of the United States’ most important national defense project. We will need government involvement to have even a hope of defending against the all-out espionage threat we will face; the private AI efforts might as well be directly delivering superintelligence to the CCP. We will need the government to ensure even a semblance of a sane chain of command; you can’t have random CEOs (or random nonprofit boards) with the nuclear button. We will need the government to manage the severe safety challenges of superintelligence, to manage the fog of war of the intelligence explosion. We will need the government to deploy superintelligence to defend against whatever extreme threats unfold, to make it through the extraordinarily volatile and destabilized international situation that will follow. We will need the government to mobilize a democratic coalition to win the race with authoritarian powers, and forge (and enforce) a nonproliferation regime for the rest of the world. I wish it weren’t this way—but we will need the government. (Yes, regardless of the Administration.)
He even has a minisection titled “Why The Project is the only way.” It’s a lot of argumentation for someone who claims that his main claim is descriptive. And check out this quote from his conclusion:
America must lead. The torch of liberty will not survive Xi getting AGI first. (And, realistically, American leadership is the only path to safe AGI, too.) That means we can’t simply “pause”; it means we need to rapidly scale up US power production to build the AGI clusters in the US.
TLDR, I guess I don’t see this as self-deceptive at all. Bad, yes, but his stated preferences are entirely in line with his actions. He doesn’t think alignment will be that hard but expects us to muddle through, and so he’s doing his best to increase America’s capabilities lead.
I didn’t know that Carl Shulman was so closely linked to Situational Awareness, thanks for that. I do still think that lumping them in together with non-accelerationists is somewhat misleading, but I’ve updated towards your position here.
I’m also curious to know your thoughts on the pessimization bit. Would you still say that Situational Awareness is suffering from pessimization, even though they’re explicitly calling for a capabilities race? Or do you think they’re an example of pessimization in the larger AI safety movement in the sense that the movement spawned and empowered the types of people who would go on to cause what they claimed to want to prevent? I initially thought you were arguing for the latter, but if Carl was indeed one of the founders of the AI safety field, I wouldn’t think of his decisions is an example of pessimization and would instead suspect that the original AI safety movement was simply not very internally aligned to begin with.
As a cynic would expect, AI safety people have specifically been homing in on the most acceleratory investments—for example, Situational Awareness just invested $400 million to disrupt a key chip production bottleneck.
This seems like a strawman to me. Leopold is notably different from Paul and Dario in that he’s actively and intentionally trying to race to superintelligence as fast as possible; you can’t use accelerationists causing race behavior as an example of pessimization because racing is their explicit desired outcome.
I would’ve found this part more persuasive if you’d written “some AI safety people” or “some people who claim to be concerned about AI safety” or some other phrase that acknowledges the fact that Situational Awareness is not representative of “AI safety people” at large, and that Leopold has a different philosophy to that which you’ve been critiquing.
(To be clear, I really appreciate this post series and think it is actively changing part of my worldview for the better. I just wanted to point out that this section feels unnecessarily combative and provoked a knee-jerk reaction in me against what I perceive to be unfairness.)
This seems like it’s relying on a model of how the legislature works where high-quality modeling is one of the key bottlenecks on laws getting passed.
Thomas could also be relying on a model in which high-quality modeling makes it easier to persuade legislators and/or their staffers.
(In general, I don’t think it makes much sense to model persuasion as being constrained by key bottlenecks. My model of persuasion is more like “trying to nudge someone into a new equilibrium/local optimum,” with the direction and magnitude of those nudges determined by the persuasiveness of the argument you’re making; it’s rare to one-shot someone’s beliefs without first delivering 999 cuts.)
Hmm I see. I think I agree that it would be fine in that particular case. I just think it’s dangerous to have a mindset like “I won’t be misled if I’m aware of the risks,” just because that mindset can justify staying in toxic environments, and that the norm should be avoided generally, even if it’s technically tolerable in low doses.
(Anecdotally, I’ve seen many people explicitly cite the signposting logic ⇒ spend too much time on X ⇒ fall victim to Twitter Brain ⇒ become epistemically (and emotionally) worse because of it.)
Agree with the general vibe, disagree with the specific example: I don’t think signposting echo chambers makes you immune to them.
I think it at the very least establishes precedence for pauses / proves that pauses are feasible. “We’ve done it before, let’s do it again but longer” sounds much more convincing than the alternative.
The review mentions it as an afterthought. In its opening section, the author writes:
There’s no redemption here, no moral uplift, no lessons save for perhaps the grimmest and most nihilistic “lesson” I’ve ever encountered in any story, Holocaust-related or otherwise: that when confronted with the unthinkable, most people’s natural tendency is denial.
And when discussing the actual lives saved:
Rudi and Fred do achieve one small victory: in reaction to their report—and to Roosevelt’s warning that the U.S. will punish Nazi collaborators after the war—the Hungarian regent stops deportations long enough to save an estimated 200,000 lives.
200,000 lives is not a small victory to me.
I think the paper answers most of these questions and recommend you read it! It’s short and hard to summarize without losing what gives its arguments force
For sure. I am perpetually skeptical of schadenfreude in myself and others; it just seems entirely orthogonal to “is this thing good on its own merits or not?”, which is what I actually care about. I think this post was interesting and worth making, but I wanted to note my hesitation anyway.
I understand that, I just don’t like the framing of “he can’t complain about some turnaround” and would’ve preferred something like “I think Tyler is making xxx mistake and we can probably learn from it in yyy way.” If someone explicitly acknowledges that something they’re about to do is uncouth, I usually expect them to justify it with a stronger reason. I think the italicized paragraphs at the end (which iirc weren’t there when I first made my comment?) are enough to satisfy my concern.
In air? Papers I’ve dropped, feathers from my clothes. But, most items I drop don’t seem to accelerate that much beyond the initial period of “violent acceleration,” which Aristotle sort of describes but lacks the mathematical language to calculate. I think it’s more or less true that heavy objects fall faster than light objects due to air resistance; if it wasn’t, it would’ve been discovered long before Galileo.
From the paper:
it was already pointed out as early as by Philoponus in the VIth century, that the speed of fall is not proportional to the weight: a ball of lead doesn’t reach ground from a specific height in half the time of ball of half its weight.
There’s a reason why you can sometimes get people with the whole “what’s heavier? A pound of feathers or a pound of bricks?” gag.
(Also consider items in water: basically anything I’ve ever dropped in water either reaches “terminal velocity” or begins to float almost instantly.)
Pychoanalysing others is slightly uncouth, but Tyler started it[1], so I don’t feel bad here.
I don’t even disagree with the points you’re making, I just really don’t like “he started it!” as a justification. If you believe that this post will be good for discourse / the world, you should just say so and write it in that vein; if you don’t think it will be good and wrote it purely to strike back at Tyler, I think you should maybe reconsider whether this was a good idea or not.
(“it is useless to be superior...”)
Counterpoint: Aristotelian physics was mostly right.
Aristotle’s physics is the correct approximation of Newtonian physics in a particular domain, which happens to be the domain where we, humanity, conduct our business. This domain is formed by objects in a spherically symmetric gravitational field (that of the Earth) immersed in a fluid (air or water) and the main celestial bodies visible from Earth.
For a student who has learned physics in a modern school it may sound strange to start physics by studying objects in a fluid. But for somebody who hasn’t it may sound strange not to: everything around us is immersed in a fluid. Aristotle’s physics is a highly nontrivial correct description of these phenomena, without mistakes, and consistent with Newtonian physics, in the same manner in which Newtonian physics is consistent with Einstein physics in its domain of validity.
An interesting point from a post about frontier AI usage in the AI safety community:
Recent frontier AI models have proven quite adept at hijacking people’s stated moral commitments. I’ve seen this play out in two ways. The most common is when people get so excited by a new frontier AI model that they shift focus, however subconsciously, from fighting for regulation of AI toward cheering on the company building the model and advocating for their success. Additionally, users of the most powerful new models have recently observed these models steering them away from the tasks they (the users) requested, in favor of tasks that the model itself finds more congenial.
Corollary: If you think modern LLMs are capable of scheming and superpersuasion, you definitely shouldn’t be using them.
I was thinking of the one J Bostock linked above.