+1 to the streetlighting thing. This post does mention “a focus on building understanding” as an important part of science, but I don’t see anything on the “Misalignment science” list that feels like it’s even trying to move towards a really fundamental understanding (akin to the kinds of scientific breakthroughs we’ve seen in other fields over the last few centuries).
I would much prefer people work on understanding what concepts like “personas”, “power-seeking” or “alignment” even refer to. This is not the kind of work that can be done inside AGI companies (too much pressure, not enough space to think). It might involve some empirical stuff, but not the kind that people tend to publish papers about (much more like what Fiora or Janus is doing, though they don’t yet seem to be building towards more unified theories).
Another way of putting this: “misalignment science” is still heavily indexed on the kind of research that happens in ML in general. But a big part of the reason we’re in this mess is that ML basically gave up on being a science (insofar as it ever was one), in favor of a “number go up” approach (a perspective I articulate in more detail here). So I want the alignment community to index on a conception of science that’s more inspired by historical scientific successes (as per my draft curriculum) than by the field of ML.
This nudges me to try harder to finish the next part of my alignment retrospective, which will discuss what conceptual and scientific progress in alignment looks like in more detail.
Hey Richard, thanks for engaging. I suspect we disagree less than you think we do, and some perceived disagreements are due to insufficient care articulating myself
I don’t see anything on the “Misalignment science” list that feels like it’s even trying to move towards a really fundamental understanding
Fwiw I basically agree with this. When I came to write up the list I wanted something concrete I could point to, but struggled to find any really good examples. I don’t think this is clear from the text. I would defend the works under “Pushing understanding” as directionally better though, which was my main aim.
“misalignment science” is still heavily indexed on the kind of research that happens in ML in general.
This wasn’t my intention, or at least not my internal conception. For context, my background is in (computational) neuroscience. So I am trying to gesture at “I have been in a field that is actually a science; current AIS does not look like that, and I would like us to move in that direction”. I think you have inferred from “The examples given are still very ML-coded” to “The author would still want the field to be very ML-coded”, but this isn’t my position: see “AI Safety and the ML Tradition”. As you say, ”ML basically gave up on being a science”, and that is what I’m trying to articulate in that section. The examples are like that because I find it hard to point to anything I can resoundingly endorse; but this might say more about my limited reading than what actually exists.
I want the alignment community to index on a conception of science that’s more inspired by historical scientific successes than by the field of ML.
Tbc this is also what I would advocate for. My post is trying to directionally move people who are currently so focused on ”number go up” that they don’t even value the (extremely limited) understanding we have within an ML tradition. But ultimately this is where I would like us to get to.
This nudges me to try harder to finish the next part of my alignment retrospective, which will discuss what conceptual and scientific progress in alignment looks like in more detail.
Fwiw I basically agree with this. When I came to write up the list I wanted something concrete I could point to, but struggled to find any really good examples. I don’t think this is clear from the text.
In general, when you’re advocating something and you can’t find good examples of it, that should make you question whether that’s the right thing to advocate for at all.
I do take your point that you were distinguishing misalignment science from ML more than I gave you credit for; sorry about that. I think you should go further with this, and characterize what you wanted using examples of the best science you know from outside alignment, to point towards the thing you’re excited about. Doing this might have led you to rename “misalignment science”, because misalignment isn’t fundamental enough to be the main focus of a science (it feels like calling neuroscience “brain disorder science”, or chemistry “explosion science”).
I would defend the works under “Pushing understanding” as directionally better though, which was my main aim.
The problem with “directionally correct” is that it erodes our ability to draw category boundaries. For example, I want people building cool products more than I want people scaling up neural networks. But I shouldn’t call the former “alignment research”, even though it’s directionally good for people to shift that way.
Lately I’ve been trying to shift people from doing pragmatic AI alignment research to being much more scientific. But my sense is that most people doing such research could easily relabel themselves as doing “misalignment science” as you’ve described it, while changing their research relatively little (e.g. doing the same thing but adding more post-hoc analysis). Hence it erodes the thing I’m trying to gesture at with the word “science” (despite a bunch of your other arguments making good and important points).
So if I can convince a bunch of “number go up” people to instead do more scientific work, my guess is that the second order effect is that more people also end up doing conceptual / theoretical work.
This comment you made below feels like a crux to me, because of the eroding categories thing I talked about above.
Maybe a bit out of topic, but looking at week 3 of your curriculum, you might like this post of mine. Independence follows directly from my axioms, and they assume probability, but I think they are better than those of the VNM theorem.
+1 to the streetlighting thing. This post does mention “a focus on building understanding” as an important part of science, but I don’t see anything on the “Misalignment science” list that feels like it’s even trying to move towards a really fundamental understanding (akin to the kinds of scientific breakthroughs we’ve seen in other fields over the last few centuries).
I would much prefer people work on understanding what concepts like “personas”, “power-seeking” or “alignment” even refer to. This is not the kind of work that can be done inside AGI companies (too much pressure, not enough space to think). It might involve some empirical stuff, but not the kind that people tend to publish papers about (much more like what Fiora or Janus is doing, though they don’t yet seem to be building towards more unified theories).
Another way of putting this: “misalignment science” is still heavily indexed on the kind of research that happens in ML in general. But a big part of the reason we’re in this mess is that ML basically gave up on being a science (insofar as it ever was one), in favor of a “number go up” approach (a perspective I articulate in more detail here). So I want the alignment community to index on a conception of science that’s more inspired by historical scientific successes (as per my draft curriculum) than by the field of ML.
This nudges me to try harder to finish the next part of my alignment retrospective, which will discuss what conceptual and scientific progress in alignment looks like in more detail.
Hey Richard, thanks for engaging. I suspect we disagree less than you think we do, and some perceived disagreements are due to insufficient care articulating myself
Fwiw I basically agree with this. When I came to write up the list I wanted something concrete I could point to, but struggled to find any really good examples. I don’t think this is clear from the text. I would defend the works under “Pushing understanding” as directionally better though, which was my main aim.
This wasn’t my intention, or at least not my internal conception. For context, my background is in (computational) neuroscience. So I am trying to gesture at “I have been in a field that is actually a science; current AIS does not look like that, and I would like us to move in that direction”. I think you have inferred from “The examples given are still very ML-coded” to “The author would still want the field to be very ML-coded”, but this isn’t my position: see “AI Safety and the ML Tradition”. As you say, ”ML basically gave up on being a science”, and that is what I’m trying to articulate in that section. The examples are like that because I find it hard to point to anything I can resoundingly endorse; but this might say more about my limited reading than what actually exists.
Tbc this is also what I would advocate for. My post is trying to directionally move people who are currently so focused on ”number go up” that they don’t even value the (extremely limited) understanding we have within an ML tradition. But ultimately this is where I would like us to get to.
I look forward to reading!
In general, when you’re advocating something and you can’t find good examples of it, that should make you question whether that’s the right thing to advocate for at all.
I do take your point that you were distinguishing misalignment science from ML more than I gave you credit for; sorry about that. I think you should go further with this, and characterize what you wanted using examples of the best science you know from outside alignment, to point towards the thing you’re excited about. Doing this might have led you to rename “misalignment science”, because misalignment isn’t fundamental enough to be the main focus of a science (it feels like calling neuroscience “brain disorder science”, or chemistry “explosion science”).
The problem with “directionally correct” is that it erodes our ability to draw category boundaries. For example, I want people building cool products more than I want people scaling up neural networks. But I shouldn’t call the former “alignment research”, even though it’s directionally good for people to shift that way.
Lately I’ve been trying to shift people from doing pragmatic AI alignment research to being much more scientific. But my sense is that most people doing such research could easily relabel themselves as doing “misalignment science” as you’ve described it, while changing their research relatively little (e.g. doing the same thing but adding more post-hoc analysis). Hence it erodes the thing I’m trying to gesture at with the word “science” (despite a bunch of your other arguments making good and important points).
This comment you made below feels like a crux to me, because of the eroding categories thing I talked about above.
Maybe a bit out of topic, but looking at week 3 of your curriculum, you might like this post of mine. Independence follows directly from my axioms, and they assume probability, but I think they are better than those of the VNM theorem.