When I read Brian Greene’s “The Elegant Universe”, I kept thinking “huh, this sure reminds me of alignment!” Both in terms of social dynamics, motivations, poor feedback loops, focus on mathematical elegance, and so on. Consuming the book made me much more sympathetic to string theory than before, when my (stringy) diet consisted mainly of Woit’s blog. This, I think, is common in science. The preponderance of rivalrous paradigms are always much closer than you think—if they weren’t, there’d be no rivalry! Like the calorie theory of heat vs the kinetic theory of heat, which both gave quantitative predictions that matched the data of their most rigorous experiments at the time.
Anyway, all of this is to say that I wound thinking that alignment researchers were like string theorists, and vice versa. For which group is this an unflattering comparison? Well, I leave that as an exercise to the reader.
There’s one key difference, and it’s that for the string theory case, we want to go beyond the effective field theories of general relativity and QFT/Standard Model of particle physics, and in particular they’re not aiming to control a system, but to understand how the universe works.
(One could validly argue that even accepting that string theory is promising/likely for our own universe, that the strategy to testing string theory on whether it describes the universe accurately should be radically different than current string theory research.)
But for AI alignment, we are aiming to control/engineer something, with science necessary as a byproduct of unique features of AI risk, and it’s plausible we might only need an effective field theory equivalent (but stated in much vaguer terms than physics does), and humans might not need to understand the limiting cases/the cases where the effective field theory fails, because we can outsource those problems to AIs.
(And this is of course very, very deeply debated, with fractal cruxes like timelines, takeoff speeds, how good AI/humans needs to be at domains where things are hard to verify, how well AI control/AI automated alignment works and more, and notably some of the cruxes can be traded off because you don’t need to get everything right, and wrongness can be tolerated, so long as it’s limited (but how limited we are is yet another deep crux here), so I won’t rehash these debates here)
Put another way, even accepting that the Agent Foundation worldview is correct at the limits of power and efficiency of real-world AIs that we can build in the long-term, it’s possible to build AIs that are far enough from those limits of power and efficiency such that the effective theories/approaches that AI safety people like Redwood Research can work.
This is unlike string theory or other quantum gravity approaches, where we know that the effective theories we have do not work to lead us anywhere close to testing them (with one caveat.)
That’s fair, but when I said “alignment” I was thinking about agent foundations in particular. Which does focus more on understanding most prosaic alignment work.
Yeah, this really does feel like the modern version of the mostly unproductive AGI/ASI/superintelligence debates, where people have implicitly different bars for what counts and doesn’t count for these concepts, and this was due to different empirical beliefs that were only somewhat validated for people thinking about AI seriously, because things that were assumed to be bundled came apart.
When I read Brian Greene’s “The Elegant Universe”, I kept thinking “huh, this sure reminds me of alignment!” Both in terms of social dynamics, motivations, poor feedback loops, focus on mathematical elegance, and so on. Consuming the book made me much more sympathetic to string theory than before, when my (stringy) diet consisted mainly of Woit’s blog. This, I think, is common in science. The preponderance of rivalrous paradigms are always much closer than you think—if they weren’t, there’d be no rivalry! Like the calorie theory of heat vs the kinetic theory of heat, which both gave quantitative predictions that matched the data of their most rigorous experiments at the time.
Anyway, all of this is to say that I wound thinking that alignment researchers were like string theorists, and vice versa. For which group is this an unflattering comparison? Well, I leave that as an exercise to the reader.
There’s one key difference, and it’s that for the string theory case, we want to go beyond the effective field theories of general relativity and QFT/Standard Model of particle physics, and in particular they’re not aiming to control a system, but to understand how the universe works.
(One could validly argue that even accepting that string theory is promising/likely for our own universe, that the strategy to testing string theory on whether it describes the universe accurately should be radically different than current string theory research.)
But for AI alignment, we are aiming to control/engineer something, with science necessary as a byproduct of unique features of AI risk, and it’s plausible we might only need an effective field theory equivalent (but stated in much vaguer terms than physics does), and humans might not need to understand the limiting cases/the cases where the effective field theory fails, because we can outsource those problems to AIs.
(And this is of course very, very deeply debated, with fractal cruxes like timelines, takeoff speeds, how good AI/humans needs to be at domains where things are hard to verify, how well AI control/AI automated alignment works and more, and notably some of the cruxes can be traded off because you don’t need to get everything right, and wrongness can be tolerated, so long as it’s limited (but how limited we are is yet another deep crux here), so I won’t rehash these debates here)
Put another way, even accepting that the Agent Foundation worldview is correct at the limits of power and efficiency of real-world AIs that we can build in the long-term, it’s possible to build AIs that are far enough from those limits of power and efficiency such that the effective theories/approaches that AI safety people like Redwood Research can work.
This is unlike string theory or other quantum gravity approaches, where we know that the effective theories we have do not work to lead us anywhere close to testing them (with one caveat.)
That’s fair, but when I said “alignment” I was thinking about agent foundations in particular. Which does focus more on understanding most prosaic alignment work.
Yeah, this really does feel like the modern version of the mostly unproductive AGI/ASI/superintelligence debates, where people have implicitly different bars for what counts and doesn’t count for these concepts, and this was due to different empirical beliefs that were only somewhat validated for people thinking about AI seriously, because things that were assumed to be bundled came apart.