Hmm, yeah, we might disagree about how much reflection(self-reference) is a central part of agency in general.
It seems plausible that it is important to distinguish between the e-coli and the human along a reflection axis (or even more so, distinguish between evolution and a human). Then maybe you are more focused on the general class of agents, and MIRI is more focused on the more specific class of “reflective agents.”
Then, there is the question of whether reflection is going to be a central part of the path to (F/D)OOM.
To operationalize, I claim that MIRI has been directed at a close enough target to yours that you probably should update on MIRI’s lack of progress at least as much as you would if MIRI was doing the same thing as you, but for half as long.
Which isn’t *that* large an update. The average number of agent foundations researchers (That are public facing enough that you can update on their lack of progress) at MIRI over the last decade is like 4.
Figuring out how to factor in researcher quality is hard, but it seems plausible to me that the amount of quality adjusted attention directed at your subgoal over the next decade is significantly larger than the amount of attention directed at your subgoal over the last decade. (Which would not all come from you. I do think that Agent Foundations today is non-trivially closer to John today that Agent Foundations 5 years ago is to John today.)
It seems accurate to me to say that Agent Foundations in 2014 was more focused on reflection, which shifted towards embeddedness, and then shifted towards abstraction, and that these things all flow together in my head, and so Scott thinking about abstraction will have more reflection mixed in than John thinking about abstraction. (Indeed, I think progress on abstraction would have huge consequences on how we think about reflection.)
In case it is not obvious to people reading, I endorse John’s research program. (Which can maybe be inferred by the fact that I am arguing that it is similar to my own). I think we disagree about what is the most likely path after becoming less confused about agency, but that part of both our plans is yet to be written, and I think the subgoal is enough of a simple concept that I don’t think disagreements about what to do next to have a strong impact on how to do the first step.
In particular, for folks reading, I symmetrically agree with this part:
In case it is not obvious to people reading, I endorse John’s research program. (Which can maybe be inferred by the fact that I am arguing that it is similar to my own). I think we disagree about what is the most likely path after becoming less confused about agency, but that part of both our plans is yet to be written, and I think the subgoal is enough of a simple concept that I don’t think disagreements about what to do next to have a strong impact on how to do the first step.
… i.e. I endorse Scott’s research program, mine is indeed similar, I wouldn’t be the least bit surprised if we disagree about what comes next but we’re pretty aligned on what to do now.
Also, I realize now that I didn’t emphasize it in the OP, but a large chunk of my “50/50 chance of success” comes from other peoples’ work playing a central role, and the agent foundations team at MIRI is obviously at the top of the list of people whose work is likely to fit that bill. (There’s also the whole topic of producing more such people, which I didn’t talk about in the OP at all, but I’m tentatively optimistic on that front too.)
I do expect reflection to be a pretty central part of the path to FOOM, but I expect it to be way easier to analyze once the non-reflective foundations of agency are sorted out. There are good reasons to expect otherwise on an outside view—i.e. all the various impossibility results in logic and computing. On the other hand, my inside view says it will make more sense once we understand e.g. how abstraction produces maps smaller than the territory while still allowing robust reasoning, how counterfactuals naturally pop out of such abstractions, how that all leads to something conceptually like a Cartesian boundary, the relationship between abstract “agent” and the physical parts which comprise the agent, etc.
If I imagine what my work would look like if I started out expecting reflection to be the taut constraint, then it does seem like I’d follow a path a lot more like MIRI’s. So yeah, this fits.
If I imagine what my work would look like if I started out expecting reflection to be the taut constraint, then it does seem like I’d follow a path a lot more like MIRI’s. So yeah, this fits.
One thing I’m still not clear about in this thread is whether you (John) would feel that progress has been made for the theory of agency if all the problems on which MIRI were instantaneously solved. Because there’s a difference between saying “this is the obvious first step if you believe reflection is the taut constraint” and “solving this problem would help significantly even if reflection wan’t the taut constraint”.
I expect that progress on the general theory of agency is a necessary component of solving all the problems on which MIRI has worked. So, conditional on those problems being instantly solved, I’d expect that a lot of general theory of agency came along with it. But if a “solution” to something like e.g. the Tiling Problem didn’t come with a bunch of progress on more foundational general theory of agency, then I’d be very suspicious of that supposed solution, and I’d expect lots of problems to crop up when we try to apply the solution in practice.
(And this is not symmetric: I would not necessarily expect such problems in practice for some more foundational piece of general agency theory which did not already have a solution to the Tiling Problem built into it. Roughly speaking, I expect we can understand e-coli agency without fully understanding human agency, but not vice-versa.)
One thing I am confused about is whether to think of the e-coli as qualitatively different from the human. The e-coli is taking actions that can be well modeled by an optimization process searching for actions that would be good if this optimization process output them, which has some reflection in it.
It feels like it can behaviorally be well modeled this way, but is mechanistically not shaped like this, I feel like the mechanistic fact is more important, but I feel like we are much closer to having behavioral definitions of agency than mechanistic ones.
I would say the e-coli’s fitness function has some kind of reflection baked into it, as does a human’s fitness function. The qualitative difference between the two is that a human’s own world model also has an explicit self-model in it, which is separate from the reflection baked into a human’s fitness function.
After that, I’d say that deriving the (probable) mechanistic properties from the fitness functions is the name of the game.
… so yeah, I’m on basically the same page as you here.
Hmm, yeah, we might disagree about how much reflection(self-reference) is a central part of agency in general.
It seems plausible that it is important to distinguish between the e-coli and the human along a reflection axis (or even more so, distinguish between evolution and a human). Then maybe you are more focused on the general class of agents, and MIRI is more focused on the more specific class of “reflective agents.”
Then, there is the question of whether reflection is going to be a central part of the path to (F/D)OOM.
Does this seem right to you?
To operationalize, I claim that MIRI has been directed at a close enough target to yours that you probably should update on MIRI’s lack of progress at least as much as you would if MIRI was doing the same thing as you, but for half as long.
Which isn’t *that* large an update. The average number of agent foundations researchers (That are public facing enough that you can update on their lack of progress) at MIRI over the last decade is like 4.
Figuring out how to factor in researcher quality is hard, but it seems plausible to me that the amount of quality adjusted attention directed at your subgoal over the next decade is significantly larger than the amount of attention directed at your subgoal over the last decade. (Which would not all come from you. I do think that Agent Foundations today is non-trivially closer to John today that Agent Foundations 5 years ago is to John today.)
It seems accurate to me to say that Agent Foundations in 2014 was more focused on reflection, which shifted towards embeddedness, and then shifted towards abstraction, and that these things all flow together in my head, and so Scott thinking about abstraction will have more reflection mixed in than John thinking about abstraction. (Indeed, I think progress on abstraction would have huge consequences on how we think about reflection.)
In case it is not obvious to people reading, I endorse John’s research program. (Which can maybe be inferred by the fact that I am arguing that it is similar to my own). I think we disagree about what is the most likely path after becoming less confused about agency, but that part of both our plans is yet to be written, and I think the subgoal is enough of a simple concept that I don’t think disagreements about what to do next to have a strong impact on how to do the first step.
This all sounds right.
In particular, for folks reading, I symmetrically agree with this part:
… i.e. I endorse Scott’s research program, mine is indeed similar, I wouldn’t be the least bit surprised if we disagree about what comes next but we’re pretty aligned on what to do now.
Also, I realize now that I didn’t emphasize it in the OP, but a large chunk of my “50/50 chance of success” comes from other peoples’ work playing a central role, and the agent foundations team at MIRI is obviously at the top of the list of people whose work is likely to fit that bill. (There’s also the whole topic of producing more such people, which I didn’t talk about in the OP at all, but I’m tentatively optimistic on that front too.)
That does seem right.
I do expect reflection to be a pretty central part of the path to FOOM, but I expect it to be way easier to analyze once the non-reflective foundations of agency are sorted out. There are good reasons to expect otherwise on an outside view—i.e. all the various impossibility results in logic and computing. On the other hand, my inside view says it will make more sense once we understand e.g. how abstraction produces maps smaller than the territory while still allowing robust reasoning, how counterfactuals naturally pop out of such abstractions, how that all leads to something conceptually like a Cartesian boundary, the relationship between abstract “agent” and the physical parts which comprise the agent, etc.
If I imagine what my work would look like if I started out expecting reflection to be the taut constraint, then it does seem like I’d follow a path a lot more like MIRI’s. So yeah, this fits.
One thing I’m still not clear about in this thread is whether you (John) would feel that progress has been made for the theory of agency if all the problems on which MIRI were instantaneously solved. Because there’s a difference between saying “this is the obvious first step if you believe reflection is the taut constraint” and “solving this problem would help significantly even if reflection wan’t the taut constraint”.
I expect that progress on the general theory of agency is a necessary component of solving all the problems on which MIRI has worked. So, conditional on those problems being instantly solved, I’d expect that a lot of general theory of agency came along with it. But if a “solution” to something like e.g. the Tiling Problem didn’t come with a bunch of progress on more foundational general theory of agency, then I’d be very suspicious of that supposed solution, and I’d expect lots of problems to crop up when we try to apply the solution in practice.
(And this is not symmetric: I would not necessarily expect such problems in practice for some more foundational piece of general agency theory which did not already have a solution to the Tiling Problem built into it. Roughly speaking, I expect we can understand e-coli agency without fully understanding human agency, but not vice-versa.)
I agree with this asymmetry.
One thing I am confused about is whether to think of the e-coli as qualitatively different from the human. The e-coli is taking actions that can be well modeled by an optimization process searching for actions that would be good if this optimization process output them, which has some reflection in it.
It feels like it can behaviorally be well modeled this way, but is mechanistically not shaped like this, I feel like the mechanistic fact is more important, but I feel like we are much closer to having behavioral definitions of agency than mechanistic ones.
I would say the e-coli’s fitness function has some kind of reflection baked into it, as does a human’s fitness function. The qualitative difference between the two is that a human’s own world model also has an explicit self-model in it, which is separate from the reflection baked into a human’s fitness function.
After that, I’d say that deriving the (probable) mechanistic properties from the fitness functions is the name of the game.
… so yeah, I’m on basically the same page as you here.