So it’s probably worth distinguishing a few questions:
Are there possible LTP setups where the STP is predicting “what’s the next override”?
Are there any of those that are biologically plausible? (Like, it’s not surprising if an LTP configured like this actually exists in the brain.)
Do any of them actually exist?
Are there possible LTP setups where the STP isn’t predicting “what’s the next override”?
Are there any of those that are biologically plausible?
Do any of them actually exist?
In any specific hypothetical example, is the STP predicting “what’s the next override”?
Is this example biologically plausible?
Does it actually exist?
I think for the (x.3) questions we’d need data we don’t have, so I’ll forget about them for the rest of the comment, they just seemed worth noting.
We both agree the answer to (1.1) is “yes”. In particular, this is what happens when the output of the STP doesn’t affect the next override. Like, if you’re on a rollercoaster, using an LTP to predict which direction to brace, and the overrides are sudden accelerations.
(1.2) I’m actually not sure about… not a confident “no”, but this setup feels to me like it would be kinda unlikely? Whereas I think you think these think are every biologically plausible LTP setup. This is just vague intuition though, and not cruxy for me.
(2.1) and (2.2) I think the answer is yes, and I think you think at least one of them is no.
So let’s look at some specific examples.
For the digestive enzymes, I actually think the binary model is already a biologically plausible example where the STP doesn’t predict the next override. But I’m happy to stick with the quantitative model too.
You’re brushing this aside, but to me it’s load-bearing. The R=1 output leads to more enzymes than R=0, and maybe even that small amount will wind up being too much. Probably not, but you only need that to happen 10% of the time.
I don’t think the numbers work out here. Compare R=1 to R=2. If these are predictions, then R=1 is a higher [predicted probability that the next override is R=0] than R=2. But R=1 produces fewer enzymes, which makes the R=0 override less likely. I think this is the case no matter what the specific curves look like, as long as they’re monotonic in the right direction.
It might be possible to come up with some way to use an LTP to handle digestive enzyme production, and have the STP predict the next override. But I feel like it doesn’t happen by default, and there’s no need for it. The job of the LTP is to handle digestive enzyme production; there’s no pressure towards “it should be possible to use the STP’s output to predict the next override”, so that doesn’t happen.
For the go-karting, I think the question is “how much do I steer?”
If I don’t steer at all in response to my predictions, then the STP can be straightforwardly correct.
If I steer a small amount, then maybe I’m still more likely to hit the side the STP predicts. Like, it thinks I’m 90% likely to hit the left side, and in fact I’m 70% more likely. We can still think of the STP as predicting the next override, but we need to recalibrate the function that reads an STP output and tells us a probability.
Ideally, I steer just the right amount, and I’m now equally likely to hit either side. The STP output is now uncorrelated with “which side do I hit”.
If I steer too much, I’m now more likely to hit the other side, and… ??? I’m not sure about the dynamics here, especially because in the time between “I oversteer” and “I actually hit the other side”, the STP is likely to be freaking out and (correctly) predicting that I’m now about to hit the other side, and I just don’t have time to react to that.
Connecting this to the digestive enzymes… both examples are going to have some kind of approximately-steady-state, like “producing a low-ish level of enzymes” or “keeping the steering wheel at a fixed angle while the track has a fixed curve”. From there, maybe there’s enough randomness that you might make too many enzymes or too few / hit either side of the track, but most likely that’s not going to happen, at least not much. (At least, we can’t assume that’s going to happen much. A simple model might have it not happening at all, and more complicated models can have it happening different amounts depending on parameters, and “it doesn’t happen much” is a reasonable setting for the parameters.)
At some point I’m going to sit down to dinner / the track is going to sharply veer left. (Definitely left, not right.) I know that’s going to happen at some point, and if my steady-state is steady enough, that means the most likely next override is “I’m not making enough enzymes” / “I’m going to hit the right wall”. But the STP isn’t going to be outputting “make more enzymes” or “turn harder left” yet, because it’s not coming up yet.
(I’m hoping to make an executable model we can play with, so if you still disagree then maybe the productive thing is to wait until I have that and that might help us figure out what’s up. Though it’s not definite I’ll succeed.)
That’s very helpful, thanks. I think the key here is this part:
The short-term predictor is a learning algorithm. Given infinite time, we normally expect learning algorithms to settle into some steady-state configuration where they no longer update. We can think of this configuration as “what we are training it to do”. So, in this toy model, what are we training the short term predictor to do?
I’m trying to talk about what will happen given infinite time, in steady state, i.e. when it gets to a fixed-point / self-consistent solution. In steady-state / at a self-consistent fixed point, the STP output approximates the expectation of the next override (and/or the expectation of the STP output at the next change of context data). So we can call that “prediction”, in the sense that “it’s a signal which tells us that a certain thing will happen later”. But it’s not a “prediction” in the sense of “passively predicting an independent, exogenous event”.
Rather, the LTP is one component of a machine (that also includes the stomach or whatever), and we’re narrowly zooming into that one component and seeing that its outputs can be interpreted as predictions, at the fixed point.
Relatedly, in logical induction (cf. here, or Theorem 4.11.2 in the original paper), they formulate a seemingly-paradoxical sentence “this sentence is true if you predict it to be true with probability LESS than 50%”. Does it have a fixed-point / steady-state? Yes, 50%. And that’s what it converges to.
Another pathological case would be to make everything a fixed point: in the case at hand, we could set up a LTP where the override is by definition whatever the STP is outputting at that moment. So then the LTP will stably output whatever random value the STP was spitting out when it was randomly initialized. Nevertheless, we can still call that a kind of “prediction”, I think. Like, we look at the STP output of 1.3 and say “it’s predicting that the next override will be 1.3”, and then the next override comes, and indeed it is 1.3, which validates that point of view.
Anyway, the text didn’t make any of this very clear, because I wasn’t really thinking about this aspect of it, so again I appreciate your comments.
Indeed, I’m now questioning a bit whether “prediction” is the right word here, since the word “prediction” does usually have a connotation of “passively predicting an exogenous event”, and I have previously criticized people for using the word “prediction” in weird situations where that connotation does not apply, so it might be hypocritical if I’m doing that myself. Hmm, I think the word “prediction” is still OK here, but I would want to add some clarifying text for sure.
Compare R=1 to R=2. If these are predictions, then R=1 is a higher [predicted probability that the next override is R=0] than R=2. But R=1 produces fewer enzymes, which makes the R=0 override less likely.
Think of this in terms of what the fixed point / steady-state / self-consistent solution would be.
Let’s say context is always the same, and on day 0, the STP outputs R=2, and that’s almost always too much digestive enzymes, so there’s R=0 overrides 90% of the time (and R=10 the other 10%). Over the next week, the STP weights update to make the predictions incrementally lower, and now R=1.8, and there’s R=0 overrides 85% of the time. Over the next week, the STP weights continue to update in the same direction, until now R=1.7, and there’s R=0 overrides 83% of the time, and R=10 overrides the other 17%. And now we’re at the fixed point! And indeed the STP output is now the expectation value of the next override, so we can (maybe slightly dubiously) use the word “prediction” to describe this output.
So it’s probably worth distinguishing a few questions:
Are there possible LTP setups where the STP is predicting “what’s the next override”?
Are there any of those that are biologically plausible? (Like, it’s not surprising if an LTP configured like this actually exists in the brain.)
Do any of them actually exist?
Are there possible LTP setups where the STP isn’t predicting “what’s the next override”?
Are there any of those that are biologically plausible?
Do any of them actually exist?
In any specific hypothetical example, is the STP predicting “what’s the next override”?
Is this example biologically plausible?
Does it actually exist?
I think for the (x.3) questions we’d need data we don’t have, so I’ll forget about them for the rest of the comment, they just seemed worth noting.
We both agree the answer to (1.1) is “yes”. In particular, this is what happens when the output of the STP doesn’t affect the next override. Like, if you’re on a rollercoaster, using an LTP to predict which direction to brace, and the overrides are sudden accelerations.
(1.2) I’m actually not sure about… not a confident “no”, but this setup feels to me like it would be kinda unlikely? Whereas I think you think these think are every biologically plausible LTP setup. This is just vague intuition though, and not cruxy for me.
(2.1) and (2.2) I think the answer is yes, and I think you think at least one of them is no.
So let’s look at some specific examples.
For the digestive enzymes, I actually think the binary model is already a biologically plausible example where the STP doesn’t predict the next override. But I’m happy to stick with the quantitative model too.
I don’t think the numbers work out here. Compare R=1 to R=2. If these are predictions, then R=1 is a higher [predicted probability that the next override is R=0] than R=2. But R=1 produces fewer enzymes, which makes the R=0 override less likely. I think this is the case no matter what the specific curves look like, as long as they’re monotonic in the right direction.
It might be possible to come up with some way to use an LTP to handle digestive enzyme production, and have the STP predict the next override. But I feel like it doesn’t happen by default, and there’s no need for it. The job of the LTP is to handle digestive enzyme production; there’s no pressure towards “it should be possible to use the STP’s output to predict the next override”, so that doesn’t happen.
For the go-karting, I think the question is “how much do I steer?”
If I don’t steer at all in response to my predictions, then the STP can be straightforwardly correct.
If I steer a small amount, then maybe I’m still more likely to hit the side the STP predicts. Like, it thinks I’m 90% likely to hit the left side, and in fact I’m 70% more likely. We can still think of the STP as predicting the next override, but we need to recalibrate the function that reads an STP output and tells us a probability.
Ideally, I steer just the right amount, and I’m now equally likely to hit either side. The STP output is now uncorrelated with “which side do I hit”.
If I steer too much, I’m now more likely to hit the other side, and… ??? I’m not sure about the dynamics here, especially because in the time between “I oversteer” and “I actually hit the other side”, the STP is likely to be freaking out and (correctly) predicting that I’m now about to hit the other side, and I just don’t have time to react to that.
Connecting this to the digestive enzymes… both examples are going to have some kind of approximately-steady-state, like “producing a low-ish level of enzymes” or “keeping the steering wheel at a fixed angle while the track has a fixed curve”. From there, maybe there’s enough randomness that you might make too many enzymes or too few / hit either side of the track, but most likely that’s not going to happen, at least not much. (At least, we can’t assume that’s going to happen much. A simple model might have it not happening at all, and more complicated models can have it happening different amounts depending on parameters, and “it doesn’t happen much” is a reasonable setting for the parameters.)
At some point I’m going to sit down to dinner / the track is going to sharply veer left. (Definitely left, not right.) I know that’s going to happen at some point, and if my steady-state is steady enough, that means the most likely next override is “I’m not making enough enzymes” / “I’m going to hit the right wall”. But the STP isn’t going to be outputting “make more enzymes” or “turn harder left” yet, because it’s not coming up yet.
(I’m hoping to make an executable model we can play with, so if you still disagree then maybe the productive thing is to wait until I have that and that might help us figure out what’s up. Though it’s not definite I’ll succeed.)
That’s very helpful, thanks. I think the key here is this part:
I’m trying to talk about what will happen given infinite time, in steady state, i.e. when it gets to a fixed-point / self-consistent solution. In steady-state / at a self-consistent fixed point, the STP output approximates the expectation of the next override (and/or the expectation of the STP output at the next change of context data). So we can call that “prediction”, in the sense that “it’s a signal which tells us that a certain thing will happen later”. But it’s not a “prediction” in the sense of “passively predicting an independent, exogenous event”.
Rather, the LTP is one component of a machine (that also includes the stomach or whatever), and we’re narrowly zooming into that one component and seeing that its outputs can be interpreted as predictions, at the fixed point.
Relatedly, in logical induction (cf. here, or Theorem 4.11.2 in the original paper), they formulate a seemingly-paradoxical sentence “this sentence is true if you predict it to be true with probability LESS than 50%”. Does it have a fixed-point / steady-state? Yes, 50%. And that’s what it converges to.
Another pathological case would be to make everything a fixed point: in the case at hand, we could set up a LTP where the override is by definition whatever the STP is outputting at that moment. So then the LTP will stably output whatever random value the STP was spitting out when it was randomly initialized. Nevertheless, we can still call that a kind of “prediction”, I think. Like, we look at the STP output of 1.3 and say “it’s predicting that the next override will be 1.3”, and then the next override comes, and indeed it is 1.3, which validates that point of view.
Anyway, the text didn’t make any of this very clear, because I wasn’t really thinking about this aspect of it, so again I appreciate your comments.
Indeed, I’m now questioning a bit whether “prediction” is the right word here, since the word “prediction” does usually have a connotation of “passively predicting an exogenous event”, and I have previously criticized people for using the word “prediction” in weird situations where that connotation does not apply, so it might be hypocritical if I’m doing that myself. Hmm, I think the word “prediction” is still OK here, but I would want to add some clarifying text for sure.
Think of this in terms of what the fixed point / steady-state / self-consistent solution would be.
Let’s say context is always the same, and on day 0, the STP outputs R=2, and that’s almost always too much digestive enzymes, so there’s R=0 overrides 90% of the time (and R=10 the other 10%). Over the next week, the STP weights update to make the predictions incrementally lower, and now R=1.8, and there’s R=0 overrides 85% of the time. Over the next week, the STP weights continue to update in the same direction, until now R=1.7, and there’s R=0 overrides 83% of the time, and R=10 overrides the other 17%. And now we’re at the fixed point! And indeed the STP output is now the expectation value of the next override, so we can (maybe slightly dubiously) use the word “prediction” to describe this output.