The first: on an intuitive level, you should think of many arguments about differential impact on the margin as analogous to arguments for timing a stock market bubble. If someone argues that the market as a whole is in a bubble, but that they’ll invest your money while it’s still going up and sell before it drops, you should probably be very skeptical. I think this analogy is actually quite deep, because the core difficulty in both cases is accounting for other people making decisions which are tightly entangled with yours. It seems possible to account for this in principle, but in practice it’s very easy to fool yourself (especially when you’re used to doing econ-style reasoning about marginal effects)—and when you do so, you’re making the bubble bigger. So I don’t trust myself (or basically anyone else in the field) to think clearly about such cases; it seems far better to focus on more robust strategies.
[The below is verbose, sorry—I feel like I’m struggling to understand what just happened. I may have a blindspot around thinking about what mentality would even lead to “let’s make capability X because that helps with alignment somehow”. I can kinda scan through the logic step by step, e.g. “we need somewhat capable systems in order to study alignment”, but something about it doesn’t make sense. Like, aren’t we worried about AGI? Why would I make the thing I’m worried about? I’ll leave my faffing about here, because maybe it’s related to communal blindspots, though maybe it’s just me being thick at the moment.]
I’m probably being dumb, but I think I don’t understand this paragraph, so maybe my confusion will provoke clarification. Or maybe I disagree or at least am not convinced. I don’t understand how [the core difficulty about evaluating strategies based on differential impact] is accounting for other people making decisions that are entangled with yours. From talking with Fable, it sounds like you’re saying something like, “A bunch of differential impact justifications assume a fixed background of AI progress against which to differentially accelerate some things; but actually, the background isn’t fixed because people like you / people using justifications like yours form some crucial element of the background.”. Is that close to the mark?
I totally buy that “AI safety / alignment” stuff accelerated some key bottlenecks, e.g. serious scaling, and that that presumably pushed timelines forward. But it doesn’t seem like that depends on the thing about entangled decisions? I’m maybe missing something really simple and obvious, like one sentence. …. Ok Fable is saying that your point is that the justifications that were used invoked the premise that someone else will do it anyway, but the someone else is other “AI safety / alignment” people. Is this right?
Now I think that you’re correct that this reasoning is severely flawed in the way you say. But I want to say that this is a subtle kind of flaw, which is less important than another bigger kind of flaw. The bigger flaw, in a word, is that …. uh, it’s bad to make the dangerous thing? See next bullet point:
From talking with Rafe, I have a different and rather more simplistic / blunt analogy: you should think of differential acceleration as being like throwing stuff into a bonfire and hoping that will decrease its long-term growth. It’s conceivable, because for example you could throw in a rock or an ice cube, or you could clear out some nearby brush that could have caused a spread, or people will be impressed with how you made the fire bigger and trust you enough to turn their backs on you while you covertly stamp out the fire, or something. But on priors, if you just throw stuff into a fire, that’s going to make there be more fire. In the analogy, the fire is the ecosystem of AI research (researchers, companies, investors, products, etc.), and catching on fire is that ecosystem absorbing [whatever research you did in the name of differential acceleration] into the progress engine.
Or to invert Bojack Horseman, when you put your glasses back on, all the differential acceleration just looks like acceleration.
Anyway, overall I feel pretty curious about what just happened, and don’t really know how to understand it. My old decrepit frame is that some people feel a very intense pull to work on “the Thing that’s going on”, and considerations that go against that get sidelined. But overall I’m confused.
I can kinda scan through the logic step by step, e.g. “we need somewhat capable systems in order to study alignment”, but something about it doesn’t make sense. Like, aren’t we worried about AGI?
Most people just don’t have the doom-by-default outlook that’s prevalent in the LW/MIRI sphere. Sure, they are worried about AGI, but they also want cancer cured, poverty eradicated, catgirl volcano lairs, etc. Building AGI seems to be the only way we’re getting that stuff any time soon, and in general one of the few remaining avenues for techno-optimism amid the bleak degrowth-and-despair cultural landscape, in the face of which accepting some substantial-but-not-overwhelming level of risk is a no-brainer.
Even they likely think that the counterfactual impact of any (ex ante reasonable) decision in their power could contribute only a minor fraction of that number. Of course, a case could be made that this is pernicious motivated reasoning, I’m only saying that it’s consistent with my model of human thinking and behavior.
[The below is verbose, sorry—I feel like I’m struggling to understand what just happened. I may have a blindspot around thinking about what mentality would even lead to “let’s make capability X because that helps with alignment somehow”. I can kinda scan through the logic step by step, e.g. “we need somewhat capable systems in order to study alignment”, but something about it doesn’t make sense. Like, aren’t we worried about AGI? Why would I make the thing I’m worried about? I’ll leave my faffing about here, because maybe it’s related to communal blindspots, though maybe it’s just me being thick at the moment.]
I’m probably being dumb, but I think I don’t understand this paragraph, so maybe my confusion will provoke clarification. Or maybe I disagree or at least am not convinced. I don’t understand how [the core difficulty about evaluating strategies based on differential impact] is accounting for other people making decisions that are entangled with yours. From talking with Fable, it sounds like you’re saying something like, “A bunch of differential impact justifications assume a fixed background of AI progress against which to differentially accelerate some things; but actually, the background isn’t fixed because people like you / people using justifications like yours form some crucial element of the background.”. Is that close to the mark?
I totally buy that “AI safety / alignment” stuff accelerated some key bottlenecks, e.g. serious scaling, and that that presumably pushed timelines forward. But it doesn’t seem like that depends on the thing about entangled decisions? I’m maybe missing something really simple and obvious, like one sentence. …. Ok Fable is saying that your point is that the justifications that were used invoked the premise that someone else will do it anyway, but the someone else is other “AI safety / alignment” people. Is this right?
Now I think that you’re correct that this reasoning is severely flawed in the way you say. But I want to say that this is a subtle kind of flaw, which is less important than another bigger kind of flaw. The bigger flaw, in a word, is that …. uh, it’s bad to make the dangerous thing? See next bullet point:
From talking with Rafe, I have a different and rather more simplistic / blunt analogy: you should think of differential acceleration as being like throwing stuff into a bonfire and hoping that will decrease its long-term growth. It’s conceivable, because for example you could throw in a rock or an ice cube, or you could clear out some nearby brush that could have caused a spread, or people will be impressed with how you made the fire bigger and trust you enough to turn their backs on you while you covertly stamp out the fire, or something. But on priors, if you just throw stuff into a fire, that’s going to make there be more fire. In the analogy, the fire is the ecosystem of AI research (researchers, companies, investors, products, etc.), and catching on fire is that ecosystem absorbing [whatever research you did in the name of differential acceleration] into the progress engine.
Or to invert Bojack Horseman, when you put your glasses back on, all the differential acceleration just looks like acceleration.
Anyway, overall I feel pretty curious about what just happened, and don’t really know how to understand it. My old decrepit frame is that some people feel a very intense pull to work on “the Thing that’s going on”, and considerations that go against that get sidelined. But overall I’m confused.
Most people just don’t have the doom-by-default outlook that’s prevalent in the LW/MIRI sphere. Sure, they are worried about AGI, but they also want cancer cured, poverty eradicated,
catgirl volcano lairs, etc. Building AGI seems to be the only way we’re getting that stuff any time soon, and in general one of the few remaining avenues for techno-optimism amid the bleak degrowth-and-despair cultural landscape, in the face of which accepting some substantial-but-not-overwhelming level of risk is a no-brainer.The part that doesn’t go through is when the people are people discussed the OP, who do thing there’s some huge (>10%) chance of extinction from AGI.
Even they likely think that the counterfactual impact of any (ex ante reasonable) decision in their power could contribute only a minor fraction of that number. Of course, a case could be made that this is pernicious motivated reasoning, I’m only saying that it’s consistent with my model of human thinking and behavior.