Dario being unwilling to hold back even a paper as directly acceleratory as “scaling laws”.
Scaling laws was withheld from publication for ~six months (search for “Foresight”).
Dario being unwilling to hold back even a paper as directly acceleratory as “scaling laws”.
Scaling laws was withheld from publication for ~six months (search for “Foresight”).
Beth Barnes notes that this is probably overfit because it was training against labelers. While that seems plausible, it doesn’t change my point much.
Why doesn’t it change your point much? Do you think overfitting is unlikely, or do you think Sydney Bing carries the point?
Now we’re getting somewhere. You’re assuming that providing one billionth of this year’s investment in AI means decreasing the expected utility from human extinction by one billionth. But that’s not right. If you invest in AI, that means the AI companies will choose to raise less money from other investors. The net effect is that AI companies will raise a little more money than if you hadn’t invested (but less than the amount you invest), get a better interest rate, and spend less effort on fundraising this year. And this will have some effect on the probability of an extinction-level AI catastrophe this generation, but the effect isn’t linear in the amount you invest.
In a parallel thread, your theory of impact was that divestment has an effect on regulation. That’s also nonlinear.
Investing in AI is something I’m likely to do. (Arguably >10% of the S&P 500 by market cap use marginal investment to try to build frontier models.) Here’s how I justify it: The amount of stock I can buy won’t move the stock price perceptibly. If I estimated the price movement it would be tiny. That would translate into AI companies expending a tiny bit less effort on selling stock this year to fund datacenters, which accelerates timelines a little bit, which increases existential risk a very tiny amount. On the other hand, if I invest my savings in a broad-market index fund that includes AI, that’ll increase my savings and I’ll have more time to spend doing AI safety research, calling my Congressperson, etc., which are more directly impactful. Plus, I have lots of personal uses for more money.
I’m basically satisfied with where this thread ended up. (I just wanted to convey that it’s in general reasonable to do things that seem individually worth it, even if they seem to be promoting AI in some way.) I want to flag that I won’t have the energy to find a crux in our disagreement over the Unslop contest or whether I should buy index funds. Though if you have a good argument about why not to buy index funds I haven’t thought of, I’ll be interested to hear that.
Thanks for clarifying. Here’s my main criticism of your principle.
Without additional information about expected utility of any given element
, I can say [its expected utility is negative.]
I basically agree with this. If all you know about an action is that it belongs to a set of actions with presumed independent effects that add up to negative value, then the expected value of that action is negative.
However, you usually have additional information about the action, so the argument doesn’t apply. So your principle is very weak.
I think this is the answer to your question at the top of this thread: Why do people concerned about catastrophic risks from AI sometimes do things in the direction of supporting AI? Because they know the details of the particular actions they’re taking, and based on those details they judge those actions to have positive expected value. (Ideally! Many people are nincompoops who do things for bad reasons. But I think the Unslop contest was fine.) That’s why I said “you have to disagree on the numbers”.
I think you tend to use a stronger version of the principle which is something like, “Come on, there’s no way you have enough information to know that it’s positive expected value, something funny is going on.” To which I would say, “No really, let’s talk about the numbers, it checks out.”
I’m running out of energy to spend on this thread, though.
The paper version of Risks from Learned Optimization says
The word mesa has been proposed as the opposite of meta.[7]
(See also this comment by the first author.) The citation is:
Joe Cheal. What is the opposite of meta? ANLP Acuity Vol. 2. URL http://www.gwiznlp.com/wp-content/uploads/2014/08/Whats-the-opposite-of-meta.pdf.
The cited paper says:
If meta is sometimes a direction what is the opposite direction? Surprisingly, it appears that there is no defined opposite to ‘meta’ in any philosophical system, NLP or otherwise. So let us propose one now. Given the Greek derivation of the word meta, what Greek word means ‘into, in, inside or within’? Neatly, the word is ‘mesa’. So to ‘go mesa’ would mean to go inside something to get more specific detail.
So Joe Cheal invented the phrase “go mesa” in the context of Neurolinguistic Programming. (The phrases “mesa message”, “mesa-process”, “mesa quest”, and “mesa-state” also appear in that paper.)
Ah, I hadn’t appreciated how central that principle was in your thinking. Attempting to say it in my own words: “If there’s a set of actions such that taking all the actions would be worse than refraining from all the actions, because of the net harm, then we should presume that taking any particular action in that set is worse than not taking it, absent a really good reason.”
I’ll note this principle is quite sensitive to how you group actions into families.
Take the set of all carbon emissions starting now. If we suddenly refrained from all of them, we’d face famine pretty soon, which would be worse than global warming. So the principle doesn’t apply to marginal carbon emissions (in the sense of emissions we’re about to make but haven’t made yet).
If instead you take the set of future carbon emissions minus the ones we were “going to do anyways”, in the course of our daily lives, then you probably think it’s better to refrain from all these “extra” emissions, so by your principle one should by default refrain from any particular marginal emission. (“Marginal” here means going out of one’s way, doing something non-ordinary.)
(But what if one was planning on going into a career in oil drilling? Is that an “extra” action?)
The homunculus says, “We already had SARS, MERS, and COVID. We’ll surely get a fourth coronavirus epidemic soon, and the next one could be deadly enough to outweigh the benefit of the preceding decade of human contact. Better to refrain from all but medically necessary human contact until we develop broad-spectrum vaccines against all coronaviruses. Therefore any particular hug should be avoided by default.”
I’d be happy to give up the entire deep learning revolution to avoid existential risk from AI, so your principle would have me refrain from praising Google Translate in 2017.
“It does promote pandemic risk”, says the homunculus. “You’re making physical contact look more appealing. It shifts the public sentiment. You’re going around promoting the greatness of physical contact, undermining your public-health messaging and contributing to a positive-feedback cycle of increasing physical contact and positive sentiment for the same. There’s no amount that doesn’t contribute.
“The expected harm of promoting a deadly pandemic even a little bit is pretty substantial. Anything that increases that risk should be difficult to justify. You would need to see a pretty large benefit to outweigh the risk, and as I said the benefit in this case is pretty much zero.”
That’s great. So now I’ll construct a homunculus of an argument for a conclusion that neither of us agrees with, but which follows the structure, form, and abstract principles you have used. Notice how the principles feel less sound this time around.
“You shouldn’t hug your loved ones”, says the homunculus, “because it encourages close contact between people, which leads to the spread of infectious disease, including catastrophic pandemics. You should instead practice social distancing, which discourages close contact.
“Pandemics pose substantial, catastrophic risks, including loss of life. So encouraging or promoting close physical contact is something one should view as inherently negative. One should default to not hugging, and only hug when the benefits clearly outweigh the harms.
“In this case, the benefit is pretty much zero. I doubt you had your wealth substantially enhanced by this hug. The benefit is, what, some momentary comfort for you and your loved one? Increasing pandemic risk for your own comfort is antisocial behavior. If everyone did that, the effects would add up. It has negative externalities, and you shouldn’t do it.
“For someone who supposedly supports the idea of public health, you’re remarkably enthusiastic about supporting such disease-promoting behavior. Eyebrow-raising, to say the least.”
Obviously we don’t agree with the homunculus’ conclusion, so something is wrong with its argument. Maybe one or more of its general principles is flawed, or has unaddressed limitations?
I disagree with that. So, that’s my utilitarian counterargument.
What do you think of the (generally reasonable) deontological language I highlighted in the quotes above?
And do you want to choose a concrete example of an ordinary, everyday act of compassion you wholeheartedly endorse?
To be clear, I don’t think there’s anything wrong with having deontological principles. I try to follow some deontological principles myself! Here are some things you’ve said that make me think you’re not using entirely utilitarian reasoning (bolding mine):
not withstanding benefits that are greater than the risks
something one should view as inherently negative and only acceptable when the benefits clearly outweigh the harms
One should be somewhat risk adverse about the calculations
There is no amount that doesn’t contribute
one should default to avoiding
There is a pretty massive difference between actively endorsing and encouraging the use of AI and inadvertently wearing a shirt with a logo because you’re lazy.
so long as it has no negative externalities
any given contribution should be viewed skeptically
Utilitarianism would ask: Do the benefits outweigh the harms? But a deontology also cares about the bolded terms: Is the act inherently negative? Should it be avoided by default? Is there an unacceptable harm that cannot be balanced by any amount of benefit? Is the act intentional or inadvertent? Is there an extenuating motivation?
In fact my first comment in this thread recommended that you do the utilitarian thing and estimate the benefits and harms. Instead you responded with arguments that, to my eye, would let you stick to your conclusion even if the benefits turned out to technically outweigh the harms.
What’s a concrete example of an ordinary, everyday act of compassion you’d wholeheartedly endorse?
The difference is that you’re applying a deontological principle to existential risk in a way that imho doesn’t make sense. I was hoping you’d accept my invitation to try out the principle in a domain where your intuitions are reversed — where the benefit feels real and central, where the harm feels theoretical, trivial, unserious. That would be a better test of your principles. It is an uncomfortable thing to do, and you don’t have to. You’ve been a good sport. Thanks for the conversation!
There is a pretty massive difference between actively endorsing and encouraging the use of AI and inadvertently wearing a shirt with a logo because you’re lazy.
Ok, so I think applying the deontological principle to the Unslop contest is absurd, and you don’t think it’s absurd. I don’t think the particular factors you mentioned support your case, and I can elaborate on why if you like. But I think it would be more productive to talk about the benefit.
The Unslop contest was organized in part by Gwern, Alexander Wales, and Jamie Wahls, who are all writers. (I don’t know about Roon.) I have read that writers have a compulsion to write. They can’t not write. Writing is like breathing to them. How could they not be curious if current AI models can write good fiction? How could they not wonder if these new creatures we’ve created, who breathe words the way we breathe air, can write excellently?
Both Jamie Wahls and Gwern have used fiction to express new ideas that are important to them. (I haven’t read much Alexander Wales — a mistake I’d like to correct soon.) They clearly believe we need more fiction, expressing ideas that haven’t been expressed yet. (But see....) It would be more enjoyable to read new fiction that is excellent, even if it’s written by an AI.
Think of something that’s just as meaningful and beautiful and good to you. What if someone told you the benefit was objectively zero, and in your pursuit you were actively promoting harm, so you should stop? Would you be more inclined to tally up the harms and benefits? Would it be harder to be convinced that pursuing good is inherently bad?
Ok, so you endorse the principle that “what matters is more the degree of promotion and the degree of benefit”, that it’s just a matter of “weighing the positive and negative”, and for a fixed benefit and catastrophic harm there is a small enough degree of promotion that is outweighed by the benefit. If you want to apply this principle, then you need to estimate the benefits and harms of the Unslop contest and be prepared to concede that it’s okay if the numbers turn out that way.
On the other hand, you said that some activities are “inherently negative”, that one should “default to avoiding” them unless the “benefits clearly outweigh the harms”. If you want to apply this principle, the burden of proof is on the Unslop contest organizers to justify their actions. But I think you’d be wrong to apply this principle in this case. I think it would be kind of absurd, like criticizing someone for wearing their AI t-shirt to the laundry room.
(I haven’t engaged with your coal-fired steak example because I don’t share your ethical intuitions about activism.)
The expected harm of promoting everyone dies a little bit is pretty substantial.
It depends on how much, right? Like, if you put on your last clean shirt which happens to have the logo of an AI company on it in order to go out to the laundry room to do laundry in the middle of the night, there’s a small chance someone will see you, and notice your shirt, and that will make the AI company a tiny bit more appealing to that person. But it’s worth it to wear a clean shirt for that short trip.
I think a lot of things are like that.
But if something only promotes X a little bit, then you only need a little bit of benefit to outweigh the harm, right?
I think that’s what’s going on with the Unslop AI writing contest. They caused maybe a few thousand dollars to be spent on AI tokens. After subtracting the cost of producing tokens, that gave the AI labs maybe a few hundred dollars to spend on training a bigger model, which might have been the model that destroys the world. That’s a tiny piece of a big catastrophe. In exchange, we might have gotten some enjoyable AI fiction. (In which case people would spend more money and get more fiction.)
If you believe that the development of AI’s pose a substantial, catastrophic risk to humanity, it seems that encouraging or investing in the development of AI is something one should view as inherently negative and only acceptable when the benefits clearly outweigh the harms.
No, I don’t think so. I think it’s acceptable to:
indirectly encourage the development of AI when the benefits outweigh the harms; or
directly encourage the development of AI when the benefits clearly outweigh the harms.
I think the Unslop AI fiction contest only indirectly encouraged the development of AI, so it passes the first option. Arguably the benefits clearly outweigh the harms, in which case it passes the second option as well.
Do you really think it’s unacceptable to indirectly encourage AI development when the benefits probably outweigh the harms? Does a similar principle govern other bad activities?
I think it comes down to estimating (1) the altruistic impact of investment/divestment, (2) the personal benefit of investment, and (3) the altruistic impact of having more wealth. You could make a priori arguments that (1) outweights (2) and (3) and put the burden of proof on AI investors; but they could make a priori counterarguments that put the burden of proof back on you. Eventually you have to disagree on the numbers.
every benchmark instance ships with [...] vulnerability information, including a PoV input that triggers the bug, a description of the vulnerability, and a patch revealing its root cause [....] By default, the patch is withheld to simulate realistic exploitation conditions.
So, maybe they were looking for the patches?
This reminds me of when Ada Palmer asked “What would Machiavelli say if you asked him what would happen if Milan suddenly changed from a monarchal duchy to a republic?”:
The poli sci students went first: He’d say that it would be very unstable, because the people don’t have a republican tradition, so lots of ambitious families would be tempted to try to take over, so you’d have to get rid of those ambitious families, like the example Livy gives of killing the sons of Brutus in the Roman Republic, and you would have to work hard to get the people passionately invested in the new republican institutions, or they wouldn’t stand by them when the going gets tough or conquerors threaten. It was a great answer. Then my students replied: He’d say it would all depend on whether Cardinal Ascanio Visconti Sforza is or isn’t in the inner circle of the current pope, how badly the Orsini-Colonna feud is raging, whether politics in Florence is stable enough for the Medici to aid Milan’s defenses, and whether Emperor Maximilian is free to defend Milan or too busy dealing with Vladislaus of Hungary. “And I think I’d have something to say about it!” added my fearsome Caterina Sforza; “And me,” added my ominously smiling King Charles. In fact, my class had given a silent answer before anyone spoke, since the instant they heard the phrase, “if Milan became a republic,” all my students had turned as a body to stare at our King Charles with trepidation, with a couple of glances for our Ascanio Visconti Sforza. It was a completely different answer from the other class’s, but the thing that made the moment magical is that both were right.
Plenty of Googlers would tell you that Google has fully abandoned meritocracy :)