my (quickly-made) snapshot: https://elicit.ought.org/builder/dmtz3sNSY
one conceptual contribution I’d put forward for consideration is whether this question may more about emotions or social equilibria than about reaching a reasoned intellectual consensus. it’s worth considering how a relatively proximate/homogenous group of people tends to change its beliefs. for better or worse, everything from viscerally compelling demonstrations of safety problems to social pressure to coercion or top-down influence to the transition from intellectual to grounded/felt risk should be part of the model of change—alongside rational, lucid, considered debate tied to deeper understanding or the truth of the matter. the demonstration doesn’t actually have to be a compelling demonstration of risks to be a compelling illustration of them (imagine a really compelling VR experience, as a trivial example).
maybe the term I’d use is ‘belief cascades’, and I might point to a rapid shift towards office closures during early COVID as an example of this. the tipping point arrived sooner than some expected, not due to considered updates in beliefs about risk or the utility of closures (the evidence had been there for a while), but rather from a cascade of fear, a noisy consensus that not acting/thinking in alignment with the perceived consensus (‘this is a real concern’) would lead to social censure, etc.
in short, this might happen sooner, more suddenly, and for stranger reasons than I think the prior distribution implies.
NB the point about a newly unveiled population of researchers in my first bin might stretch the definition of ‘top AI researchers’ in the question specification, but I believe it’s in line with the spirit of the question
Thank you for posting this!
The section on DeepMind explains how capabilities & alignment came to be intertwined. Since GDM was also probably the organization with the strongest prestige dynamics & status hierarchies adversarial to safety, I’ll add a couple of subjective anecdotes on what it was like from the inside to hold the view that DeepMind’s core AGI roadmap was both (a) plausible on short timelines and (b) dangerous, such that alignment should be taken seriously.
I worked in the comms and policy org starting in 2018. All external comms were carefully monitored—not only for confidentiality reasons, but also to avoid reputational damage. If there was perceived risk of online controversy (social media, press articles), the team would either deny the request outright or take some actions (like comms training, editing the content) to reduce reputational risk. In my recollection, the operating principle was to earn either positive attention or no attention. Controversy would be met with a stern email or a meeting that appeared on your calendar, and future comms would be monitored more carefully.
If a researcher was asked about their views on existential risk from AI, they were guided to respond with a statement along the lines of: “It’s not useful to engage in that kind of alarmism, which originates in science fiction. Some people confuse AI with movies like Terminator—that’s simply not the reality. The AI we develop will be safe by design. After all, we have a team of expert researchers building it!” (--> Steer conversation towards beneficial applications in health, climate, etc).
It took months of intensive internal advocacy to change this protocol towards one where publicly acknowledging AI risk was less actively discouraged, and the Safety team could start publishing a higher volume of (positively valenced, comms-friendly) content on AI safety without being internally censured. To be clear, the early resistance to commenting on AI safety in particular did not come from the top of the organization, but rather from middle layers that were responsible for policies on external communications, but lacked sufficient technical or strategic context to understand AGI was plausible in the near term and not safe by default.
There was also a default of distrust with other labs, which—whether or not justified—was notable for the status dynamics that rewarded this. My comms training and instincts on risk-aversion run too deep even now for me to say much more on this, but one can imagine it being socially more favorable to echo a sense of vague distrust than attempt to drive forward policy or safety work that involved coordination with other labs.
These dynamics steadily improved as the Overton window shifted and it became a positive status signal to be supportive of safety work. People like Geoffrey, Rohin, and Allan joining were helpful for this, as was Jan, Vika, et al’s early work in this area.