Ajeya Cotra comments on The case for aligning narrowly superhuman models

Ajeya Cotra 10 Mar 2021 20:00 UTC
LW: 10 AF: 7
AF
In my head the point of this proposal is very much about practicing what we eventually want to do, and seeing what comes out of that; I wasn’t trying here to make something different sound like it’s about practice. I don’t think that a framing which moved away from that would better get at the point I was making, though I totally think there could be other lines of empirical research under other framings that I’d be similarly excited about or maybe more excited about.

In my mind, the “better than evaluators” part is kind of self-evidently intriguing for the basic reason I said in the post (it’s not obvious how to do it, and it’s analogous to the broad, outside view conception of the long-run challenge which can be described in one sentence/phrase and isn’t strongly tied to a particular theoretical framing):

I’m excited about tackling this particular type of near-term challenge because it feels like a microcosm of the long-term AI alignment problem in a real, non-superficial sense. In the end, we probably want to find ways to meaningfully supervise (or justifiably trust) models that are more capable than ~all humans in ~all domains.[4] So it seems like a promising form of practice to figure out how to get particular humans to oversee models that are more capable than them in specific ways, if this is done with an eye to developing scalable and domain-general techniques.

A lot of people in response to the draft were pushing in the direction that I think you were maybe gesturing at (?) -- to make this more specific to “knowing everything the model knows” or “ascription universality”; the section “Why not focus on testing a long-term solution?” was written in response to Evan Hubinger and others. I think I’m still not convinced that’s the right way to go.