I still don’t see it, sorry. If I think of deep learning as an approximation of some kind of simplicity prior + updating on empirical evidence, I’m not very surprised that it solves the capacity allocation problem and learns a productive model of the world. [1] The price is that the simplicity prior doesn’t necessarily get rid of scheming. The big extra challenge for heuristic explanations is that you need to do the same capacity allocation in a way that scheming reliably gets explained (even though it’s not relevant for the model’s performance and doesn’t make things classically simpler), while no capacity is spent on explaining other phenomena that are not relevant for the model’s performance. I still don’t see at all how we can get the the non-malign prior that can do that.
I still don’t see it, sorry. If I think of deep learning as an approximation of some kind of simplicity prior + updating on empirical evidence, I’m not very surprised that it solves the capacity allocation problem and learns a productive model of the world. [1] The price is that the simplicity prior doesn’t necessarily get rid of scheming. The big extra challenge for heuristic explanations is that you need to do the same capacity allocation in a way that scheming reliably gets explained (even though it’s not relevant for the model’s performance and doesn’t make things classically simpler), while no capacity is spent on explaining other phenomena that are not relevant for the model’s performance. I still don’t see at all how we can get the the non-malign prior that can do that.
Though I’m still very surprised that it works in practice.