See also David Rein’s AI Companies as ‘hot springs’ for growing AIs or Buck’s AI catastrophes and rogue deployments from 2024.
Vaniver
Model Weight Exfiltration Seems Overrated
FWIW, I don’t think Paul served a large role as an optimistic counterweight to those sounding the alarm, at least in public.
Which time period are you thinking of?
I definitely never perceived Paul as an e/acc, if that’s what you’re imagining as an optimistic counterweight.[1] But I’ve been following this stuff since ~2010, and working in the Bay on it since 2016, and my experience of much of that time was there were two main thought leaders, Eliezer and Paul, with Eliezer characterizing the pessimistic view that thought alignment was difficult enough that you needed agent foundations to have a good scaling story, and Paul championing the view of relative alignment ease and thought we had a reasonable shot of aligning prosaic AI.
And like, if you read that post (originally from 2016, published on LW in 2018), he definitely doesn’t seem happy about the situation:
I feel pessimistic about human prospects in such a world.
So I am extremely unhappy with “a significant risk of trouble.”
We might hope that this situation will change automatically as we build more sophisticated AI systems. But I don’t think that’s necessarily the case.
Like I said, I think this would probably be OK, but it opens up an unreasonably high chance of really bad outcomes.
And in his 2022 post on his disagreements with Eliezer, he clearly thinks humanity can lose if it just doesn’t try:
humanity is fully capable of messing up even very basic challenges, especially if they are novel.
But on my read of the alignment landscape in the 2016-2021 era, I think the main conclusion of Paul’s position was “let’s try and see how it goes.” I think a lot of the funding in the AI safety landscape at the time went into that strategy, and I think it ran into many of the difficulties predicted by the MIRI camp.
[And, to be clear, I don’t pretend to know the counterfactual, here. If Paul had instead been a doomer in 2016, would a different person have emerged as ‘thought leader’ of the ‘optimists’, because the controversy was demand-driven instead of supply-driven?]
First, I don’t think it’s reasonable to have a confident position on [whether aligning prosaic AI is infeasible]. Claims of the form “problem X can’t be solved” are really hard to get right, because you are fighting against the universal quantifier of all possible ways that someone could solve this problem.
Like, in 2017 Paul was at OpenAI working on RLHF while I was at MIRI trying to sound the alarm. (One of my early tasks at MIRI was trying to explain our strategic worldview in some detail, and spent a lot of time thinking about Paul’s arguments as things I needed to address.)
Since ~2021, my read has been that Anthropic is the center of alignment research of the “try and see how it goes” camp, with Dario as the primary thought leader (who, unfortunately, does not write for the public).
- ^
Techcrunch, today, called him a “prominent AI doomer”.
- ^
I’m pretty uncertain, and I expect this to look a lot like ‘flicking a switch’ in retrospect, even tho the leadup to flicking that switch will probably look like smoothly increasing capabilities on ‘toy’ problems.
I think I feel pretty good about this prediction. But also I can see a bit of the case for smoothness—in the two months before April, we got about 10% of the increase that we got from Jan-April.
Was anyone from the Constellation dumb enough to work in any AI-creating lab besides OAI, GDM and Anthropic?
Sorry, I don’t understand how this functions as an objection.
For years I’ve thought people should consider the strategy of joining labs to only work on things-good-for-the-world, and people at labs staying to work on those things also seems promising. But I think I mostly don’t buy this at present margins. Sodom and Gomorrah aren’t destroyed until Lot leaves, and every respectable person who sticks around not committing any crimes themselves is acting as cover for the organization that is happily committing crimes. Resign today.
Happy Labor Day, everyone.
Machine Organizations
A poster child here could be, perhaps, the Scopes Monkey Trial.
Actually, the example I was thinking of while reading this post was v-structures in causal discovery. When I first came across it, I was surprised that there was something that you can’t do with two elements but can do with three. [Of course, someone who understands that doesn’t run into the cat-belling problem with explaining it. They do have the convenience that the base case is still simple enough to easily hold in mind, but even in cases where the enabler is obscure or subtle it seems like they should be able to acknowledge the difficulty.]
A hope for debate might be that, if a person has a good-enough hidden stat to construct an adequate alignment plan in the traditional way (reading, thinking, corresponding with other humans, etc.) given enough time, then that same person will also be good enough to judge a debate between AIs about alignment plans.
For what it’s worth, I think this surfaces a different difficulty. One of the intuitions powering debate is the idea that some things are easier to check than generate—imagine the problem of finding a counterexample to a conjecture, where you might have to search thru a huge set of examples to find one, but then the verifier only needs a few steps to confirm that it’s a counterexample. Debate then represents a way to get exponential effort that is verified in linear time, allowing humans to oversee the work of much more capable machines.
But some things are easier to generate than check. Joel Spolsky, back in 2000, wrote “It’s harder to read code than write it.” I expect it will be easier to be confident in a program that I write than one that Claude Code generates (both that it will be correct and that it doesn’t contain adversarially chosen side effects). For many years, human translators preferred translating a passage from scratch to editing a machine-translated passage. What makes us suspect that the philosophy or science of alignment is fertile ground for debate?
The commitment to stay together permanently is a limited resource. What happens if your wife and your boyfriend get into an argument?
It’s not incompatible with poly—you might have a hierarchical poly structure, where the marriages are durable and take precedence, but this makes the secondary relationships precarious and some poly authors dislike this structure accordingly.
Also, if everyone is together on the basis of “making the decision that’s best for me right now”, monogamy doesn’t encourage building alternatives and poly does. Maybe you’ll end up liking your boyfriend more than your wife; maybe your wife will end up liking her girlfriend more than you. In mono, the choice is instead “do I start an affair, which will maybe blow everything up, or do I stick with the thing that’s good and work to improve it?”
An important fact is that many of us are not, or have divested from investments made in rosier times.
Personally, I’m astounded that they apparently kept their leverage at 3X? As the stocks went up, I simply would not have taken out more loans.
A way that I characterized this (in my feedback) is that ‘high control’ is, from the perspective of the users, often the point! Lots of people will do things like join the military because the military will seize control of their life and turn them into a different sort of person, and they would rather be that kind of person than the person they are now (or that they expect to turn into under their own control).
And so the core questions are: ok, but do you want to be shaped by this person or by this system? It is often hard to tell this from inside the system—if it weren’t self-recommending it would have already collapsed—but I think this is nevertheless the core question. From my (somewhat distant) vantage point, MAPLE doesn’t look great on this front.
This is why I like Carroll Quigley’s argument, that war is for resolving uncertainty in power relations. If Alice and Bob agree on their relative strengths, they enter some settlement where the stronger party dominates the weaker party (but not so much that it’s worth it for the weaker party to resist). But if Alice and Bob disagree on their relative strengths, then they fight it out.
Considering the war in Ukraine, if the Russians expected it to be a quick adventure and the Ukrainians expected that they could hold out, then it makes sense for there to be a war instead of a negotiated settlement, because the two parties have very different expectations of what will happen if there is a war.
I think this lens also helps explain the war in Israel and Palestine, and the war in Iran. (All three feel to me like this explanation is incomplete, and domestic disagreements seem like important factors. 10⁄7 being kicked off by the parts of the government that most benefit from continued conflict feels important—but also uncertainty over the outcome must have played into their decision-making.)
Does The War Trap explain the wars that are happening now?
But almost the same thing is true for causal dependency graphs. Pearl’s definitions of causality are ultimately fairly circular.
Note that BB approvingly quotes Schwarz, who thinks that other formulations of CDT by Lewis, Joyce, or Skyrms are better than Pearl’s.
But why? I just pointed you to what I claim is a very clear example of Achmiz pushing people towards the right sort of sportsmanship, by means of arguing for it at length. How else is one supposed to do it? Reply!
I think there are obviously means for encouraging other people to do things besides arguing for it at length. For example, one could reward people for doing it, or one can model it in one’s own behavior.
I think “require it and let them sort out how it’s produced” is a perfectly effacious plan. Do you think it wouldn’t work? Why not? Reply!
Perfectly? Is such exaggeration proper, here?
I think people differ, and so there’s some value in systems that simply judge the external interface and let the individual sort out how they comply with the interface. But I think people are also similar, and there’s some value in us peeling back the layers and comparing how we’re put together. Part of the study of rationality is figuring out the mechanisms of internal techniques; like looking closely enough at meditation to figure out which biological systems it interacts with, or looking closely enough at therapy to distill out Focusing, and so on.
But maybe you think the decline of Less Wrong 1.0 is definitive evidence that I can’t successfully explain away.
Do you have an explanation for the decline of Less Wrong 1.0? That event feels pretty important to my thinking, here, and my read of you is something like “well we didn’t try to Correct Plan hard enough”. Like, if Yudkowsky had more personal virtue, then he wouldn’t have stopped posting on LW 1.0 and there wouldn’t have been the decline.
I mean, it’s not as if Shapin doesn’t have critics, but I doubt the details are a crux.
...why do you doubt the details are a crux?
To express some annoyance, this is 20 pages of detailed back-and-forth, where a historian criticizes the piece, the editor who commissioned the historian’s essay criticizes the essay, then Shapin responds, then the historian responds to Shapin. It is worth reading if you think the historian makes good points; it is not worth reading if you think the counterpoints hold. (I, reading thru it, found myself agreeing with Dear and Shapin more than Feingold.) I am trying to understand what is going on; I am looking for evidence that bears on the question. For recreational debating, it doesn’t matter whether or not arguments are good—in fact, often bad arguments are more fun to tear apart!--but that is not the game I’m trying to play, at the moment.
I’m actually a little confused about what’s going on here and you might be in a better position to explain it than me. Is it just that my high-effort style makes it look bad to purge me after I put so much work in, whereas Achmiz’s questions put the interpretive labor burden on the author?
I also don’t have a complete picture here. I don’t think the effort story resonates that well because I think the outputs are more meaningful than the inputs. I do think we had more basic disagreements with Said about what made for good conversational approaches.
Academic decision theorists don’t like the theory.
Their loss.
I’m wondering if AIS advocates unnecessarily burned a bridge here, given that Lasher was sympathetic enough to co-sponsor, thus he might have been pulled away from LTF later if the race hadn’t heated up.
From his victory comments, I think he was never close to LTF:
“I have some news for the two big AI companies who are taking such an unusual interest,” he said. “I won’t be taking my cues from either of you when it comes to protecting our kids, our jobs and our families.”
But now he knows who had his back when the chips were down, and conversely who was making him work his ass off in a tight race.
My understanding is that many races involve the winning candidates making some attempts to incorporate losing candidates or their supporters; it’s not obvious this matters for heavily partisan districts like NY-12 (rather than large, close, ideologically diverse races like for the presidency). It’s not obvious whether it’s worth reaching out to both candidates before a race starts, or after the other guy wins and has a real sense of how much strength you were dealing with.

Just fixed, thanks!