I think this post says a lot of objectively good points, but simultaneously makes me uneasy about the conclusions that people will draw from it. I think this is because it doesn’t go far enough. The “misalignment science” vision you tout still seems like streetlighting to me, and will fail in the face of illegible problems (examples) that can’t be surfaced by the techniques you suggest. The engineering approach to alignment fails not just because there is not enough focus on understanding why solutions to problems work, but also because some problems cannot be investigated by anything other than theoretical work.
That being said, I agree that more prosaic alignment researchers switching to this method is a first-order positive impact. It seems like there is at least some lessons in understanding intelligent agency that we can draw from LLMs and prosaic alignment work (although I’m not sure the examples you give are close enough to actually doing that). I just worry that the second-order effect is to make us more confident in our research without making us more safe (I can imagine someone saying “look, alignment used to be about patching problems and not caring about why they were problems and why our solutions worked, but now we understand all the empirical phenomena we see, so all our problems are solved!” right before everyone dying to something that has no empirical signs before the superintelligence stage).
The work might also address failures which are “speculative”, or not present in current systems.
I think you are getting at what I’m saying in this one line here. I just wish it was emphasized more in the essay. Hence the “this is all good but it makes me uneasy” feeling.
+1 to the streetlighting thing. This post does mention “a focus on building understanding” as an important part of science, but I don’t see anything on the “Misalignment science” list that feels like it’s even trying to move towards a really fundamental understanding (akin to the kinds of scientific breakthroughs we’ve seen in other fields over the last few centuries).
I would much prefer people work on understanding what concepts like “personas”, “power-seeking” or “alignment” even refer to. This is not the kind of work that can be done inside AGI companies (too much pressure, not enough space to think). It might involve some empirical stuff, but not the kind that people tend to publish papers about (much more like what Fiora or Janus is doing, though they don’t yet seem to be building towards more unified theories).
Another way of putting this: “misalignment science” is still heavily indexed on the kind of research that happens in ML in general. But a big part of the reason we’re in this mess is that ML basically gave up on being a science (insofar as it ever was one), in favor of a “number go up” approach (a perspective I articulate in more detail here). So I want the alignment community to index on a conception of science that’s more inspired by historical scientific successes (as per my draft curriculum) than by the field of ML.
This nudges me to try harder to finish the next part of my alignment retrospective, which will discuss what conceptual and scientific progress in alignment looks like in more detail.
Hey Richard, thanks for engaging. I suspect we disagree less than you think we do, and some perceived disagreements are due to insufficient care articulating myself
I don’t see anything on the “Misalignment science” list that feels like it’s even trying to move towards a really fundamental understanding
Fwiw I basically agree with this. When I came to write up the list I wanted something concrete I could point to, but struggled to find any really good examples. I don’t think this is clear from the text. I would defend the works under “Pushing understanding” as directionally better though, which was my main aim.
“misalignment science” is still heavily indexed on the kind of research that happens in ML in general.
This wasn’t my intention, or at least not my internal conception. For context, my background is in (computational) neuroscience. So I am trying to gesture at “I have been in a field that is actually a science; current AIS does not look like that, and I would like us to move in that direction”. I think you have inferred from “The examples given are still very ML-coded” to “The author would still want the field to be very ML-coded”, but this isn’t my position: see “AI Safety and the ML Tradition”. As you say, ”ML basically gave up on being a science”, and that is what I’m trying to articulate in that section. The examples are like that because I find it hard to point to anything I can resoundingly endorse; but this might say more about my limited reading than what actually exists.
I want the alignment community to index on a conception of science that’s more inspired by historical scientific successes than by the field of ML.
Tbc this is also what I would advocate for. My post is trying to directionally move people who are currently so focused on ”number go up” that they don’t even value the (extremely limited) understanding we have within an ML tradition. But ultimately this is where I would like us to get to.
This nudges me to try harder to finish the next part of my alignment retrospective, which will discuss what conceptual and scientific progress in alignment looks like in more detail.
Fwiw I basically agree with this. When I came to write up the list I wanted something concrete I could point to, but struggled to find any really good examples. I don’t think this is clear from the text.
In general, when you’re advocating something and you can’t find good examples of it, that should make you question whether that’s the right thing to advocate for at all.
I do take your point that you were distinguishing misalignment science from ML more than I gave you credit for; sorry about that. I think you should go further with this, and characterize what you wanted using examples of the best science you know from outside alignment, to point towards the thing you’re excited about. Doing this might have led you to rename “misalignment science”, because misalignment isn’t fundamental enough to be the main focus of a science (it feels like calling neuroscience “brain disorder science”, or chemistry “explosion science”).
I would defend the works under “Pushing understanding” as directionally better though, which was my main aim.
The problem with “directionally correct” is that it erodes our ability to draw category boundaries. For example, I want people building cool products more than I want people scaling up neural networks. But I shouldn’t call the former “alignment research”, even though it’s directionally good for people to shift that way.
Lately I’ve been trying to shift people from doing pragmatic AI alignment research to being much more scientific. But my sense is that most people doing such research could easily relabel themselves as doing “misalignment science” as you’ve described it, while changing their research relatively little (e.g. doing the same thing but adding more post-hoc analysis). Hence it erodes the thing I’m trying to gesture at with the word “science” (despite a bunch of your other arguments making good and important points).
So if I can convince a bunch of “number go up” people to instead do more scientific work, my guess is that the second order effect is that more people also end up doing conceptual / theoretical work.
This comment you made below feels like a crux to me, because of the eroding categories thing I talked about above.
Maybe a bit out of topic, but looking at week 3 of your curriculum, you might like this post of mine. Independence follows directly from my axioms, and they assume probability, but I think they are better than those of the VNM theorem.
I think you are getting at what I’m saying in this one line here. I just wish it was emphasized more in the essay.
You’re right that this is what I was aiming for here, and I agree it’s underemphasised. I think it’s a huge problem that AIS currently operates almost exclusively in a “reactionary” mode, where we only try to investigate or solve problems after they’ve arisen. So I would be very supportive or more people working on trying to investigate or make legible problems which are currently not prominent. In a world where AIS was an actual science, but nobody is working on making problems legible, I would likely write an analogous essay criticising this instead.
That being said, I agree that more prosaic alignment researchers switching to this method is a first-order positive impact.
FWIW this is the desired first-order effect. I guess my mental model of the field is that research is dispersed around some central point on the axis of “number go up“ to “theoretical and conceptual understanding”, with “misalignment science” as some point on the right of that axis. So if I can convince a bunch of “number go up” people to instead do more scientific work, my guess is that the second order effect is that more people also end up doing conceptual / theoretical work.[1]
For this reason I’m skeptical of your posited second order effect. I think if the entire field shifts away from an engineering attitude then we’ll also see more people choosing to do philosophical and theoretical work, which I would view as good.
In either case I don’t think we disagree: I think we would both prefer if (relatively) more people were working on scientific understanding and if more people were working on conceptual / theoretical work. Apologies for not articulating this clearly enough in the post!
If my essay convinced people who are currently working on conceptual progress and making problems legible to instead work on already-legible problems I would be disappointed and tell them not to do that.
I still think I am sensing some disagreement (although it is quite subtle, and I co), but I think the source of that disagreement is currently implicit / intuitive in my mind, and would require a better conception/language/understanding of the philosophy of science to make fully explicit. Nonetheless, I’ll try to gesture at it.
(Also, I think my critiques overlap a lot with Richard’s in the same comment thread, so apologies if it feels like I’m just piling on here. Thank you for writing this post and for your engagement with my comment. I hope that adding my own reply makes things clearer in some way, rather than just giving you more to read through.)
Trying to gesture at the disagreement, it could be something like “shallow theoretical understanding vs. deep theoretical understanding”. I think you’re aiming to argue for more theoretical understanding (which is great!), but many of the examples that you list are stuff that I would classify as generating shallow theoretical understanding. In contrast, if you glance at Richard’s WIP curriculum, it’s filled with stuff that I would classify as aiming at deep theoretical understanding. As for what the exact difference is, I think that’s where I’m struggling with the need for conceptual clarity, sorry (and naming it “shallow vs. deep” this way probably is inadequate too). But here’s a test: would the results of a project hold if the first AGI was not an LLM (or anything remotely like it)? If not, it’s probably not progress towards deep theoretical understanding. Similarly: would it hold up for an LLM 5-10 OOMs more powerful than present systems? (Presuming that this gets us to something like superintelligence, which I actually don’t believe, but that’s not relevant right now.) The vast majority of the difficulty of the alignment problem comes from systems more powerful than the ones we have currently, so failing to generalize would not be helpful in solving the bulk of the hard problems. As an example, do you think the Why Do Some Models Fake Alignment… paper you linked would hold up?
The difference I’m sensing might also have something to do with post-hoc analysis vs. a theory-first approach, I’m not sure. It could also be the comprehensiveness of the theoretical analysis. Does this analysis attempt to tell us something fundamental about intelligent agency, or does it merely explain the results from this paper (or set of papers)? I think these are related to the “holding up for superintelligence” test above. But again, I’m probably gesturing inadequately here.
This is also not to say that deep theoretical understanding could not come from studying current LLMs. But it doesn’t seem to me that much of the work you linked above is done in a way that gets to that point. Lessons from LLMs would play much more of a supporting role, rather than a main role, in this type of work. An example off the top of my head: how does a mind become coherent, and what does it mean for a mind to be coherent? (Aiming at implications for corrigibility and building limited-coherence systems that could create a pivotal act without being agentic enough to seek power). We could use LLMs as interesting evidence, as they clearly become coherent (or show failures of incoherence) in a very different way than people do, but the bulk of the work would have to venture far beyond interpreting limited LLM results.
And I think this is where the disagreement over the “second-order effect” comes into play. Compared to the work that you list, the work that Richard lists in his curriculum seems to be of a very different flavor, and I think virtually all of the really hard, really important work needs to be of this flavor. I could possibly see myself being wrong and “misalignment science” sparking more interest in theoretical / philosophical work, as you suggest...but A) the incentives are stacked against this, given how much easier it is to get funding and employment to work on more legible, easier problems, B) there seems to be a qualitative gap between these two types of work, and I’m not so sure more people doing “misalignment science” would direct that much interest towards theoretical / philosophical work. So I agree that this pulls us in the right direction on the engineering<->theory spectrum, but because there is a non-linearity in the spectrum, where the “this might actually work for aligning superintelligence” level goes from ~0 to some positive number past the point of shallow theoretical understanding, the move generates undue confidence without helping much.
I think this post says a lot of objectively good points, but simultaneously makes me uneasy about the conclusions that people will draw from it. I think this is because it doesn’t go far enough. The “misalignment science” vision you tout still seems like streetlighting to me, and will fail in the face of illegible problems (examples) that can’t be surfaced by the techniques you suggest. The engineering approach to alignment fails not just because there is not enough focus on understanding why solutions to problems work, but also because some problems cannot be investigated by anything other than theoretical work.
That being said, I agree that more prosaic alignment researchers switching to this method is a first-order positive impact. It seems like there is at least some lessons in understanding intelligent agency that we can draw from LLMs and prosaic alignment work (although I’m not sure the examples you give are close enough to actually doing that). I just worry that the second-order effect is to make us more confident in our research without making us more safe (I can imagine someone saying “look, alignment used to be about patching problems and not caring about why they were problems and why our solutions worked, but now we understand all the empirical phenomena we see, so all our problems are solved!” right before everyone dying to something that has no empirical signs before the superintelligence stage).
I think you are getting at what I’m saying in this one line here. I just wish it was emphasized more in the essay. Hence the “this is all good but it makes me uneasy” feeling.
+1 to the streetlighting thing. This post does mention “a focus on building understanding” as an important part of science, but I don’t see anything on the “Misalignment science” list that feels like it’s even trying to move towards a really fundamental understanding (akin to the kinds of scientific breakthroughs we’ve seen in other fields over the last few centuries).
I would much prefer people work on understanding what concepts like “personas”, “power-seeking” or “alignment” even refer to. This is not the kind of work that can be done inside AGI companies (too much pressure, not enough space to think). It might involve some empirical stuff, but not the kind that people tend to publish papers about (much more like what Fiora or Janus is doing, though they don’t yet seem to be building towards more unified theories).
Another way of putting this: “misalignment science” is still heavily indexed on the kind of research that happens in ML in general. But a big part of the reason we’re in this mess is that ML basically gave up on being a science (insofar as it ever was one), in favor of a “number go up” approach (a perspective I articulate in more detail here). So I want the alignment community to index on a conception of science that’s more inspired by historical scientific successes (as per my draft curriculum) than by the field of ML.
This nudges me to try harder to finish the next part of my alignment retrospective, which will discuss what conceptual and scientific progress in alignment looks like in more detail.
Hey Richard, thanks for engaging. I suspect we disagree less than you think we do, and some perceived disagreements are due to insufficient care articulating myself
Fwiw I basically agree with this. When I came to write up the list I wanted something concrete I could point to, but struggled to find any really good examples. I don’t think this is clear from the text. I would defend the works under “Pushing understanding” as directionally better though, which was my main aim.
This wasn’t my intention, or at least not my internal conception. For context, my background is in (computational) neuroscience. So I am trying to gesture at “I have been in a field that is actually a science; current AIS does not look like that, and I would like us to move in that direction”. I think you have inferred from “The examples given are still very ML-coded” to “The author would still want the field to be very ML-coded”, but this isn’t my position: see “AI Safety and the ML Tradition”. As you say, ”ML basically gave up on being a science”, and that is what I’m trying to articulate in that section. The examples are like that because I find it hard to point to anything I can resoundingly endorse; but this might say more about my limited reading than what actually exists.
Tbc this is also what I would advocate for. My post is trying to directionally move people who are currently so focused on ”number go up” that they don’t even value the (extremely limited) understanding we have within an ML tradition. But ultimately this is where I would like us to get to.
I look forward to reading!
In general, when you’re advocating something and you can’t find good examples of it, that should make you question whether that’s the right thing to advocate for at all.
I do take your point that you were distinguishing misalignment science from ML more than I gave you credit for; sorry about that. I think you should go further with this, and characterize what you wanted using examples of the best science you know from outside alignment, to point towards the thing you’re excited about. Doing this might have led you to rename “misalignment science”, because misalignment isn’t fundamental enough to be the main focus of a science (it feels like calling neuroscience “brain disorder science”, or chemistry “explosion science”).
The problem with “directionally correct” is that it erodes our ability to draw category boundaries. For example, I want people building cool products more than I want people scaling up neural networks. But I shouldn’t call the former “alignment research”, even though it’s directionally good for people to shift that way.
Lately I’ve been trying to shift people from doing pragmatic AI alignment research to being much more scientific. But my sense is that most people doing such research could easily relabel themselves as doing “misalignment science” as you’ve described it, while changing their research relatively little (e.g. doing the same thing but adding more post-hoc analysis). Hence it erodes the thing I’m trying to gesture at with the word “science” (despite a bunch of your other arguments making good and important points).
This comment you made below feels like a crux to me, because of the eroding categories thing I talked about above.
Maybe a bit out of topic, but looking at week 3 of your curriculum, you might like this post of mine. Independence follows directly from my axioms, and they assume probability, but I think they are better than those of the VNM theorem.
Hey Cameron, I appreciate the comment!
You’re right that this is what I was aiming for here, and I agree it’s underemphasised. I think it’s a huge problem that AIS currently operates almost exclusively in a “reactionary” mode, where we only try to investigate or solve problems after they’ve arisen. So I would be very supportive or more people working on trying to investigate or make legible problems which are currently not prominent. In a world where AIS was an actual science, but nobody is working on making problems legible, I would likely write an analogous essay criticising this instead.
FWIW this is the desired first-order effect. I guess my mental model of the field is that research is dispersed around some central point on the axis of “number go up“ to “theoretical and conceptual understanding”, with “misalignment science” as some point on the right of that axis. So if I can convince a bunch of “number go up” people to instead do more scientific work, my guess is that the second order effect is that more people also end up doing conceptual / theoretical work.[1]
For this reason I’m skeptical of your posited second order effect. I think if the entire field shifts away from an engineering attitude then we’ll also see more people choosing to do philosophical and theoretical work, which I would view as good.
In either case I don’t think we disagree: I think we would both prefer if (relatively) more people were working on scientific understanding and if more people were working on conceptual / theoretical work. Apologies for not articulating this clearly enough in the post!
If my essay convinced people who are currently working on conceptual progress and making problems legible to instead work on already-legible problems I would be disappointed and tell them not to do that.
I still think I am sensing some disagreement (although it is quite subtle, and I co), but I think the source of that disagreement is currently implicit / intuitive in my mind, and would require a better conception/language/understanding of the philosophy of science to make fully explicit. Nonetheless, I’ll try to gesture at it.
(Also, I think my critiques overlap a lot with Richard’s in the same comment thread, so apologies if it feels like I’m just piling on here. Thank you for writing this post and for your engagement with my comment. I hope that adding my own reply makes things clearer in some way, rather than just giving you more to read through.)
Trying to gesture at the disagreement, it could be something like “shallow theoretical understanding vs. deep theoretical understanding”. I think you’re aiming to argue for more theoretical understanding (which is great!), but many of the examples that you list are stuff that I would classify as generating shallow theoretical understanding. In contrast, if you glance at Richard’s WIP curriculum, it’s filled with stuff that I would classify as aiming at deep theoretical understanding. As for what the exact difference is, I think that’s where I’m struggling with the need for conceptual clarity, sorry (and naming it “shallow vs. deep” this way probably is inadequate too). But here’s a test: would the results of a project hold if the first AGI was not an LLM (or anything remotely like it)? If not, it’s probably not progress towards deep theoretical understanding. Similarly: would it hold up for an LLM 5-10 OOMs more powerful than present systems? (Presuming that this gets us to something like superintelligence, which I actually don’t believe, but that’s not relevant right now.) The vast majority of the difficulty of the alignment problem comes from systems more powerful than the ones we have currently, so failing to generalize would not be helpful in solving the bulk of the hard problems. As an example, do you think the Why Do Some Models Fake Alignment… paper you linked would hold up?
The difference I’m sensing might also have something to do with post-hoc analysis vs. a theory-first approach, I’m not sure. It could also be the comprehensiveness of the theoretical analysis. Does this analysis attempt to tell us something fundamental about intelligent agency, or does it merely explain the results from this paper (or set of papers)? I think these are related to the “holding up for superintelligence” test above. But again, I’m probably gesturing inadequately here.
This is also not to say that deep theoretical understanding could not come from studying current LLMs. But it doesn’t seem to me that much of the work you linked above is done in a way that gets to that point. Lessons from LLMs would play much more of a supporting role, rather than a main role, in this type of work. An example off the top of my head: how does a mind become coherent, and what does it mean for a mind to be coherent? (Aiming at implications for corrigibility and building limited-coherence systems that could create a pivotal act without being agentic enough to seek power). We could use LLMs as interesting evidence, as they clearly become coherent (or show failures of incoherence) in a very different way than people do, but the bulk of the work would have to venture far beyond interpreting limited LLM results.
And I think this is where the disagreement over the “second-order effect” comes into play. Compared to the work that you list, the work that Richard lists in his curriculum seems to be of a very different flavor, and I think virtually all of the really hard, really important work needs to be of this flavor. I could possibly see myself being wrong and “misalignment science” sparking more interest in theoretical / philosophical work, as you suggest...but A) the incentives are stacked against this, given how much easier it is to get funding and employment to work on more legible, easier problems, B) there seems to be a qualitative gap between these two types of work, and I’m not so sure more people doing “misalignment science” would direct that much interest towards theoretical / philosophical work. So I agree that this pulls us in the right direction on the engineering<->theory spectrum, but because there is a non-linearity in the spectrum, where the “this might actually work for aligning superintelligence” level goes from ~0 to some positive number past the point of shallow theoretical understanding, the move generates undue confidence without helping much.