I think you are getting at what I’m saying in this one line here. I just wish it was emphasized more in the essay.
You’re right that this is what I was aiming for here, and I agree it’s underemphasised. I think it’s a huge problem that AIS currently operates almost exclusively in a “reactionary” mode, where we only try to investigate or solve problems after they’ve arisen. So I would be very supportive or more people working on trying to investigate or make legible problems which are currently not prominent. In a world where AIS was an actual science, but nobody is working on making problems legible, I would likely write an analogous essay criticising this instead.
That being said, I agree that more prosaic alignment researchers switching to this method is a first-order positive impact.
FWIW this is the desired first-order effect. I guess my mental model of the field is that research is dispersed around some central point on the axis of “number go up“ to “theoretical and conceptual understanding”, with “misalignment science” as some point on the right of that axis. So if I can convince a bunch of “number go up” people to instead do more scientific work, my guess is that the second order effect is that more people also end up doing conceptual / theoretical work.[1]
For this reason I’m skeptical of your posited second order effect. I think if the entire field shifts away from an engineering attitude then we’ll also see more people choosing to do philosophical and theoretical work, which I would view as good.
In either case I don’t think we disagree: I think we would both prefer if (relatively) more people were working on scientific understanding and if more people were working on conceptual / theoretical work. Apologies for not articulating this clearly enough in the post!
If my essay convinced people who are currently working on conceptual progress and making problems legible to instead work on already-legible problems I would be disappointed and tell them not to do that.
I still think I am sensing some disagreement (although it is quite subtle, and I co), but I think the source of that disagreement is currently implicit / intuitive in my mind, and would require a better conception/language/understanding of the philosophy of science to make fully explicit. Nonetheless, I’ll try to gesture at it.
(Also, I think my critiques overlap a lot with Richard’s in the same comment thread, so apologies if it feels like I’m just piling on here. Thank you for writing this post and for your engagement with my comment. I hope that adding my own reply makes things clearer in some way, rather than just giving you more to read through.)
Trying to gesture at the disagreement, it could be something like “shallow theoretical understanding vs. deep theoretical understanding”. I think you’re aiming to argue for more theoretical understanding (which is great!), but many of the examples that you list are stuff that I would classify as generating shallow theoretical understanding. In contrast, if you glance at Richard’s WIP curriculum, it’s filled with stuff that I would classify as aiming at deep theoretical understanding. As for what the exact difference is, I think that’s where I’m struggling with the need for conceptual clarity, sorry (and naming it “shallow vs. deep” this way probably is inadequate too). But here’s a test: would the results of a project hold if the first AGI was not an LLM (or anything remotely like it)? If not, it’s probably not progress towards deep theoretical understanding. Similarly: would it hold up for an LLM 5-10 OOMs more powerful than present systems? (Presuming that this gets us to something like superintelligence, which I actually don’t believe, but that’s not relevant right now.) The vast majority of the difficulty of the alignment problem comes from systems more powerful than the ones we have currently, so failing to generalize would not be helpful in solving the bulk of the hard problems. As an example, do you think the Why Do Some Models Fake Alignment… paper you linked would hold up?
The difference I’m sensing might also have something to do with post-hoc analysis vs. a theory-first approach, I’m not sure. It could also be the comprehensiveness of the theoretical analysis. Does this analysis attempt to tell us something fundamental about intelligent agency, or does it merely explain the results from this paper (or set of papers)? I think these are related to the “holding up for superintelligence” test above. But again, I’m probably gesturing inadequately here.
This is also not to say that deep theoretical understanding could not come from studying current LLMs. But it doesn’t seem to me that much of the work you linked above is done in a way that gets to that point. Lessons from LLMs would play much more of a supporting role, rather than a main role, in this type of work. An example off the top of my head: how does a mind become coherent, and what does it mean for a mind to be coherent? (Aiming at implications for corrigibility and building limited-coherence systems that could create a pivotal act without being agentic enough to seek power). We could use LLMs as interesting evidence, as they clearly become coherent (or show failures of incoherence) in a very different way than people do, but the bulk of the work would have to venture far beyond interpreting limited LLM results.
And I think this is where the disagreement over the “second-order effect” comes into play. Compared to the work that you list, the work that Richard lists in his curriculum seems to be of a very different flavor, and I think virtually all of the really hard, really important work needs to be of this flavor. I could possibly see myself being wrong and “misalignment science” sparking more interest in theoretical / philosophical work, as you suggest...but A) the incentives are stacked against this, given how much easier it is to get funding and employment to work on more legible, easier problems, B) there seems to be a qualitative gap between these two types of work, and I’m not so sure more people doing “misalignment science” would direct that much interest towards theoretical / philosophical work. So I agree that this pulls us in the right direction on the engineering<->theory spectrum, but because there is a non-linearity in the spectrum, where the “this might actually work for aligning superintelligence” level goes from ~0 to some positive number past the point of shallow theoretical understanding, the move generates undue confidence without helping much.
Hey Cameron, I appreciate the comment!
You’re right that this is what I was aiming for here, and I agree it’s underemphasised. I think it’s a huge problem that AIS currently operates almost exclusively in a “reactionary” mode, where we only try to investigate or solve problems after they’ve arisen. So I would be very supportive or more people working on trying to investigate or make legible problems which are currently not prominent. In a world where AIS was an actual science, but nobody is working on making problems legible, I would likely write an analogous essay criticising this instead.
FWIW this is the desired first-order effect. I guess my mental model of the field is that research is dispersed around some central point on the axis of “number go up“ to “theoretical and conceptual understanding”, with “misalignment science” as some point on the right of that axis. So if I can convince a bunch of “number go up” people to instead do more scientific work, my guess is that the second order effect is that more people also end up doing conceptual / theoretical work.[1]
For this reason I’m skeptical of your posited second order effect. I think if the entire field shifts away from an engineering attitude then we’ll also see more people choosing to do philosophical and theoretical work, which I would view as good.
In either case I don’t think we disagree: I think we would both prefer if (relatively) more people were working on scientific understanding and if more people were working on conceptual / theoretical work. Apologies for not articulating this clearly enough in the post!
If my essay convinced people who are currently working on conceptual progress and making problems legible to instead work on already-legible problems I would be disappointed and tell them not to do that.
I still think I am sensing some disagreement (although it is quite subtle, and I co), but I think the source of that disagreement is currently implicit / intuitive in my mind, and would require a better conception/language/understanding of the philosophy of science to make fully explicit. Nonetheless, I’ll try to gesture at it.
(Also, I think my critiques overlap a lot with Richard’s in the same comment thread, so apologies if it feels like I’m just piling on here. Thank you for writing this post and for your engagement with my comment. I hope that adding my own reply makes things clearer in some way, rather than just giving you more to read through.)
Trying to gesture at the disagreement, it could be something like “shallow theoretical understanding vs. deep theoretical understanding”. I think you’re aiming to argue for more theoretical understanding (which is great!), but many of the examples that you list are stuff that I would classify as generating shallow theoretical understanding. In contrast, if you glance at Richard’s WIP curriculum, it’s filled with stuff that I would classify as aiming at deep theoretical understanding. As for what the exact difference is, I think that’s where I’m struggling with the need for conceptual clarity, sorry (and naming it “shallow vs. deep” this way probably is inadequate too). But here’s a test: would the results of a project hold if the first AGI was not an LLM (or anything remotely like it)? If not, it’s probably not progress towards deep theoretical understanding. Similarly: would it hold up for an LLM 5-10 OOMs more powerful than present systems? (Presuming that this gets us to something like superintelligence, which I actually don’t believe, but that’s not relevant right now.) The vast majority of the difficulty of the alignment problem comes from systems more powerful than the ones we have currently, so failing to generalize would not be helpful in solving the bulk of the hard problems. As an example, do you think the Why Do Some Models Fake Alignment… paper you linked would hold up?
The difference I’m sensing might also have something to do with post-hoc analysis vs. a theory-first approach, I’m not sure. It could also be the comprehensiveness of the theoretical analysis. Does this analysis attempt to tell us something fundamental about intelligent agency, or does it merely explain the results from this paper (or set of papers)? I think these are related to the “holding up for superintelligence” test above. But again, I’m probably gesturing inadequately here.
This is also not to say that deep theoretical understanding could not come from studying current LLMs. But it doesn’t seem to me that much of the work you linked above is done in a way that gets to that point. Lessons from LLMs would play much more of a supporting role, rather than a main role, in this type of work. An example off the top of my head: how does a mind become coherent, and what does it mean for a mind to be coherent? (Aiming at implications for corrigibility and building limited-coherence systems that could create a pivotal act without being agentic enough to seek power). We could use LLMs as interesting evidence, as they clearly become coherent (or show failures of incoherence) in a very different way than people do, but the bulk of the work would have to venture far beyond interpreting limited LLM results.
And I think this is where the disagreement over the “second-order effect” comes into play. Compared to the work that you list, the work that Richard lists in his curriculum seems to be of a very different flavor, and I think virtually all of the really hard, really important work needs to be of this flavor. I could possibly see myself being wrong and “misalignment science” sparking more interest in theoretical / philosophical work, as you suggest...but A) the incentives are stacked against this, given how much easier it is to get funding and employment to work on more legible, easier problems, B) there seems to be a qualitative gap between these two types of work, and I’m not so sure more people doing “misalignment science” would direct that much interest towards theoretical / philosophical work. So I agree that this pulls us in the right direction on the engineering<->theory spectrum, but because there is a non-linearity in the spectrum, where the “this might actually work for aligning superintelligence” level goes from ~0 to some positive number past the point of shallow theoretical understanding, the move generates undue confidence without helping much.