Technical Director at Principles of Intelligence (PrincInt, PIBBSS) keen on narrowing the theory-practice gap in AI safety through research, field building, and metascientific reflections.
Formerly: Theoretical high energy physicist; science, technology, and society research; professor and academic advisor to interdisciplinary undergrads
Thanks for this post. I think the problem it names is a real bottleneck for field building and grant making in AI safety rather than only an epistemics puzzle (which, as I read them, was the focus of many comments).
First, the question: “Is this group an opaque meritocracy or a cabal?” often isn’t well-posed globally. ‘Consensus’ underlying meritocratic transparency is tied to a specific scientific culture and depends on its granularity. For example, the goals, standards of evidence, and scientific values of ‘physics’ as a whole are not the same as for ‘high-energy physics’ as a whole or ‘quantum gravity’. I like string theory as a reference for AI safety because it did involve the explicit collaboration between physicists and mathematicians that led to the formulation of a combined scientific culture (in string theory). Clearly, this led to concerns over the health and direction of physics as a field, which is echoed in the many opinions for what counts as pursuit-worthy new directions in AI safety. Also relevant to AI safety is that the scientific cultures that populate a field tend to get more granular over time as we learn more, and that we have a unique (if less ‘natural’, in an academic sense) field building opportunity to decide how to incorporate new ideas. I hope we can do this in a way that does not result in a collection of mutually opaque meritocracies. Personally, I think an approach that is pluralist about research programs and strict about the artifacts that should come out of those programs strikes a good balance, though defining this merit criteria per subculture in a way that aggregates to a comprehensive ‘good’ collective is hard.
This leads to a second point: this criteria changes over time. In a rough historical sketch, the String Wars came after a first revolution (mid 1980s work fixing anomalies, resulting in a landscape of superstring theories as QG candidates), a period of criticism (for example, in 1986: desperately seeking superstrings) and revolution again (1990s work on dualities and introducing branes, AdS/CFT). As Dmitry pointed out, this had implications on the field’s scientific culture, including its goals and normative standards. I haven’t thought this analogy through, but this may have a parallel in AI safety, with ‘string theory as quantum gravity’ standing in epistemically for agent foundations, and the empirical contact of AdS/CFT with real-world complex systems playing the role of mechanistic interpretability.
Lastly, the string wars example is as much about resource scarcity/allocation as criteria for ‘good science’. A less granular example is the Anderson-Weinberg fight over the SSC, where each subfield was unable to adjudicate the other’s merit while competing for the same funding. It’s also worth noting that resource competition is not only where consensus gets expressed, but also how it solidifies within a scientific culture. This is perhaps another lesson for AI safety. Our scarce resource may not be less about money than grantmaker attention and limits to individual expertise (so that ‘opaque’ defaults to ‘unfundable’). I agree with Jonas’ comment below.