Great post, excepct, maybe now is the time to be going out and pushing for policies?
Chris_Leong
A quick list of policies that are likely much easier to advocate for in light of the OpenAI-Hugging Face incident:
• More attention for loss-of-control risks
• Mandatory incident reporting (including pre-harm incidents)
• Third-party safety evaluators
• Stronger monitoring systems
• Better sandboxes (perhaps air-gaps)
• Honeypots for certain training runs
• Higher cybersecurity standards
• Government incident investigation capabilities
• AI kill switches
• Emergency procedures (for both AI labs and governments)
• Perhaps even licensing for training frontier models (would need to be politically-neutral)
• The ability of the government to force a company to pause training runs
• AI pause in general (though likely as something that may be necessary in the future)
• More model organisms work (though as we’ve seen, this comes with risks)
• AI liability (I have mixed feelings on whether this is a good idea given the potential for this to lead to a compliance mindset)
Would love to hear any other thoughts.
1) I would recommend against selecting out: it’s up to BlueDot whether to accept you or not. They likely have a better idea of how many facilitators they have and who else could take the slot than you would (I don’t feel like I have a solid idea and I facilitate for them). My opinion, which I expect to be quite different from BlueDot’s, is that if someone is spending a significant amount of time in EA/rationalist/AI safety spaces, then it is valuable for them to have a better understanding of these issues as it raises the standard of intellectual environment.
2) When you’re just reading things yourself, it’s very easy to just end up reading things that are recent or which you find personally interesting; and this can cause you to miss some things that are important.
3) Reading hits a limit at some point. Eventually you need to be discussing these ideas with other people too, as they will point out things that you’ve missed.
4) Are you any good at teaching? It might be worth doing the course if you think you would have a shot at being a good facilitator, even if you don’t think you’d be competitive for other opportunities in the space.
I know there’s some tension between my points 1) and 3), but I still think it is better to post an imperfect comment than to not respond at all.
Have you considered doing a BlueDot course?
I’ve been facilitating for them recently, so I’m biased, but I recommend that everyone does it unless they’re explicitly not doing it as part of a strategy to increase intellectual diversity in the space.
CeSIA has activated its warning-shot protocol
Can you share anything about your warning shot protocol?
In any case, I think that the AI Safety and Governance communities should be building more capacity to act on warning shots.
The first implicit pillar of the Berkeley Model that we want to criticize is the assumption of content indifference: The Berkeley Model assumes we can fully separate the technical problem of aligning an AI to some values or goals and the governance problem of choosing what values or goals to target. While it is logically possible that we’ll discover some fully generic method of pointing to goals or values (e.g. brain-reading), it’s equally plausible that different goals or values will effectively have different ‘type-signatures’: goals or values that are highly unnatural or esoteric given one training method or specification-format may be readily accessible given another training method or specification-format, and vice versa.
I’m glad to see someone articulate this point. On the other hand...
Consider that we humans ourselves manage to be respectful, caring, and helpful to our friends despite not fully knowing what they care about or what their life plans are—thereby providing an informal human proof for the possibility of beneficial and safe behavior without exhaustive learning of the target’s values. And as concerns sufficiency, the recent literature on deceptive alignment vividly demonstrates that value learning by itself can’t guarantee the right relationship to motivation: understanding human value and caring about values are different things.
Perhaps I’m confused, but this feels like what someone who misunderstood the Berkeley Model of Alignment and accidentally strawmanned it would write?
Do you know what your plans are now?
Relevant damage here is reputational, rather than monetary
Okay, that’s interesting, and a good point. That said, given decreasing marginal returns, I expect such reputational interventions to be less effective than you might think (only a proportion of the reputational loss will actually be counterfactual to how things would have played out anyway, maybe it’ll happen slightly faster).
Also, insofar as suing labs for reasons that are absolutely legit — both legally and ethically — is going to “damage the relationships”, this seems like evidence that the lab doesn’t care that much about the Good
Not necessarily. If we’re really doing it to make traction on larger-scale issues distinct from the actual lawsuit, it’s quite understandable why folks at the labs might consider these to be ‘bad faith’. I would even be tempted to make a virtue ethics argument against this.
Alex Turner’s post from yesterday is emblematic of how little influence even “top tier” AI safety people employed at an AI lab can expect to have over the lab’s decisions.
Disagree. It just showed that the US government is even more powerful.
Let’s handicap/slow them down into oblivion
That’s prob. a crux.
I’m not expecting any significant effect here. They can afford better lawyers than you and absorb a significant amount of losses.
Not clear it is worth completing burning the AIS communities relationships with the labs (OpenAI might be an exception as that is arguably already burnt).I find it hard to believe that they would do a decent enough job on their own in the absence of regulation/suing, such that legal action against them would be net-negative in expectation
Trade-off is increase in policy effectiveness from positive pressure (likely low given how poorly understood the issue is) minus increase in defensiveness (likely high if we go hard here).
This balance works out different from other issues.
Responding directly to these questions might take us too into the weeds. Do you understand why my argument isn’t fully generalisable?
We shouldn’t assume that “unfavourable for the AI companies” automatically means that it is good for the world or that they won’t just eat the lawsuits.
It’s not “fully generalisable” because it is extremely contingent on safe AGI development being especially hard to standardise and courts having decades or even centuries less experience handling these issues than they have in other areas.
If we sell the concept that “you’re serving your client the best by telling them to deploy the safest AIs”, then we’re steering money away from AI companies that choose to move away from safety research...
Hmm… whilst I wouldn’t completely rule out that theory of change given short-timelines (given the potential to get lucky), it doesn’t feel very reliable as a short-timelines plan. It feels more like a medium-long timelines play to me.
On the other hand: the lawyers that are representing people suing said companies for AI harms (e.g. Adam Raine’s parents) must and should be assisted by the community.
I’m pretty skeptical of this ToC. Suing labs pushes them towards a defensive, compliance mindset and away from focusing on, “How do we make sure that this AI actually remains safe as we scale up?”.
Lawyers working at Anthropic, OpenAI, Microsoft, Google or Meta, are as compromised / able OR unable to influence things as AI engineers or technical staff at said companies.
This ToC certainly has its limits, but is much more likely to deliver value on short-timelines (not to mention that if folks decide to quit, then their frontier lab experience gives them a lot of policy credibility).
And, if we really anticipate timelines to be short, I think training the people who are already fluent in implementation to understand AI Safety, is worthwhile.
If timelines are short, then it is much more likely that governance at only a few companies will matter[1].
That said, maybe there’s value in getting more practitioners onside, either to assist in writing laws, to work at the companies that matter or just to have a more diverse range of voices supporting governance policies.- ^
Though perhaps not as few as people think insofar as it is possible to have access to frontier bio-capabilities without necessarily being at a frontier lab.
- ^
Best practices already exist… the priority bottleneck is not finding more best practices
I’m pretty skeptical of this framing. It’s not clear to me that current ‘best practises’ are sufficient or, stronger, whether the 80⁄20 framing holds.
That said, I suspect that a decent percentage of non top-tier researchers should pivot more towards advocacy assuming they meet a minimum proficiency bar for advocacy and that they have access to such opportunties.
Total research transparency feels a bit too galaxy-brained for me. It makes non-robust assumptions that newly discovered techniques won’t be usable to enhance already existing open-weights models to excessively dangerous capability levels. I also think the disincentive for research is overstated as it neglects first-mover advantage.
I recently facilitated the AISF course again (technical, governance and strategy courses). If you’ve been in AIS for a while and you want to revisit the basics, there’s certainly worse ways of doing it.
This is a good time to be asking this question.
We’ve recently seen a massive shift from the White House. Honestly, it’s still hard to imagine, but things can change fast.
Gary Tan and Marc Andreessen were boosting it at one point.
“When or if”—what if it is open now