I think it’s kind of weird to try to get probability 0 of catastrophe.
Like, for the continuous case, the PAC-like inequality you take as the starting point is what I’m fine with as the end point—I’m fine with a reasonable guarantee of good behavior with high probability. If there’s a measure ~0 spike of catastrophe out there if we initialize the parameters just wrong, I don’t in practice mind.
So I’m more interested in the story of how the researchers possibly got to the PAC-like bound (and the philosophy of how they formulated it).
One thing to note is that this result does not say anything about the probability of catastrophe.
To do that, you’d need to define a prior over proxies resulting from the training process (which we don’t do in this paper/post). Depending on this prior, catastrophe could be arbitrarily likely.
Because of this, I’m not comfortable with the continuous framework inequality alone as an endpoint; I’d also need an understanding of the priors of our learning algorithms or alternate incentive structures that require agents to adjust their values based on feedback from humanity
I think it’s kind of weird to try to get probability 0 of catastrophe.
Like, for the continuous case, the PAC-like inequality you take as the starting point is what I’m fine with as the end point—I’m fine with a reasonable guarantee of good behavior with high probability. If there’s a measure ~0 spike of catastrophe out there if we initialize the parameters just wrong, I don’t in practice mind.
So I’m more interested in the story of how the researchers possibly got to the PAC-like bound (and the philosophy of how they formulated it).
Thanks for commenting!
One thing to note is that this result does not say anything about the probability of catastrophe. To do that, you’d need to define a prior over proxies resulting from the training process (which we don’t do in this paper/post). Depending on this prior, catastrophe could be arbitrarily likely.
Because of this, I’m not comfortable with the continuous framework inequality alone as an endpoint; I’d also need an understanding of the priors of our learning algorithms or alternate incentive structures that require agents to adjust their values based on feedback from humanity