Earth science PhD student.
Mark Hellwig
If a pause requires the development of some kind of political and technical framework for ensuring all actors (states and companies) comply, and that framework can be (even partially) reused for a subsequent pause, then I would imagine that a second pause, all things equal, would have a lower “cost.” For instance, if a main objection from people in the US to a pause is that China will never comply, then demonstrating (somehow) that a bilateral agreement works even once would make future attempts at it more likely.
I suppose the counter to that is that if compliance can’t be monitored well (or if it’s demonstrated that people aren’t complying), then the opposite happens.
There’s a lot that would shift that “cost” balance one way or another: the relative capabilities gap between labs, the degree to which people think that alignment was already solved the last time a pause happened, any recent warning shots (e.g. OpenAI/HF), but if we ignore these things, I feel somewhat confident that people would get used to the idea of pauses happening as soon as you do it once.
In this model, it’s more advantageous to get started on a pause as soon as possible, I imagine.But also it’s pretty obvious that current models have a ~0% risk of any significant, world-scale harm, so investing effort at this stage is like wasting a bunch of time and resources preparing for a chess match with an opponent who is well below your Elo anyway.
Are you assuming here that present effort is unlikely to be applicable to a future model?
It’s not obvious to me that the public generalizing a particular incident to AI as a whole is a less sophisticated response. There might be situations in which one response is more likely than another (e.g. Mecha-Hitler feels like a particularly xAI-flavored problem but OAI/HF doesn’t feel particularly limited to OAI, at least to me). But if the (lack of) response so far to this current incident is anything to go off of, there seems to be a prior held by many people against the notion that AGI is actually liable to “go rogue.” If that’s true, then I would expect that the less obviously lethal a warning shot is, the less likely people are to generally view it as a warning shot for AI in general, and the more likely they are to view it as the specific fault of a particular company.
Mark Hellwig’s Shortform
I’ve had lengthy one-on-one conversations with a few people (5, IIRC) since the Hugging Face incident about AI alignment (all of whom were either unaware of or unconcerned by the issue), and only one resulted in the person I was talking to being convinced that alignment is difficult, misalignment in highly intelligent AI is dangerous, and something ought to be done about it.
All other conversations had us moving back and forth between a number of different objections to misalignment risk, but in every case we eventually concluded that the source of our disagreement was that they did not believe that ASI was likely to ever exist. That belief, I think, meant that the possibility of agents autonomously developing goals in any meaningful sense was hard to believe, as well as the possibility that they might ever want or be able to deliberately circumvent guardrails against ‘bad’ behavior. I’m certainly not the first to notice this, but it was a striking thing for me to observe anyway, because it took getting through a number of seemingly unrelated topics to get to that conclusion.
I’m in the earth sciences, so a natural connection for me is to my conversations about climate change. Even though most people in my life believe that it is happening and nominally support action against it, I have met somewhat fewer people that have stated support for policies that would reduce greenhouse gas emissions if they also hurt the economy or would require meaningful lifestyle changes (e.g. drastically reducing the number of livestock raised for meat production). It seems that either some of these people are not as convinced about the likelihood or magnitude of harm as they say, or conviction in the likelihood of harm is insufficient reason to prioritize solutions to it over other things they value. My guess is that some combination of the two of these underlies a lot of collective complacency toward it.
I say this to express my concern that is is unlikely that the general public will be convinced to support a pause on capabilities research because of xrisk. They might support it for some other reason (jobs, datacenters, etc.), but the existential risk of misaligned ASI seems diffuse enough to give policies like this some substantial inertia. Until we have a particularly egregious example of a mustache-twirling rogue agent evading guardrails and cybersecurity defenders and committing some close-to-home harmful action, it will be easy to rationalize why we don’t need to do anything “drastic” about it, and (though this is just my hunch) people are and will continue to be motivated to find those rationalizations.
This is mostly anecdotal, so I might be underestimating people’s general willingness to update their beliefs.
This is a helpful way to reframe things, especially from the perspective of “I don’t know enough about [thing] to make meaningful progress on it by myself” and trying to not feel hopeless about it.
)) is dominated by strategies relating to recruiting others. If you need, for instance, 51% of 340 million people to firmly support a development pause for it to happen (in the US), then you need 5.1%, or 17 million people to directly convince 10 people in the US to support it. This would be roughly n = 500. This can be upwards of a few orders of magnitude higher if you consider network effects, or if the people you recruit are more directly able to do an n = 1 billion thing.
and copies of me at my current skill level (which is not very high) to appreciably move p(doom) -- so n ranges from 800 to 80,000. If I spend the next few years upskilling as much as I can, maybe that moves to n = 8,000,000 (1,000 copies of me needed). But the question still remains whether I’d have done more just by trying to get somebody smarter or more politically connected than me to work on the problem, and I’d guess the answer is yes.
I imagine a lot of the intermediate range (between n = 1 and n = O(
It’s not obvious to me how I can hill climb on n in a more direct way. I would guess that you would need somewhere between