I’d very much appreciate a discussion on the argument.
Some ways I could see pushing back in it:
Humans learn better out of order: there are studies showing this, but as far as I could find the material being shuffled were not compositionally depenedent, meaning they could be learned in any order.
Better architecture, not priors: in a previous article on AlphaFold, I argued that its hard-coded inductive priors lent it the sample-efficiency it needed to learn something generalizable for the relatively sparse protein database. There, the priors were baked in a specialized architecture, but humans are sample-efficient more broadly. So perhaps there is an architecture that is just more sample efficient without needing ordered training or hierarchical priors.
What about replay: our brain does replay experiences, which nominally mitigates catastrophic interference. Although the replay seems to prioritize surprise/novelty rather than already-learned experiences.
I’d very much appreciate a discussion on the argument.
Some ways I could see pushing back in it:
Humans learn better out of order: there are studies showing this, but as far as I could find the material being shuffled were not compositionally depenedent, meaning they could be learned in any order.
Better architecture, not priors: in a previous article on AlphaFold, I argued that its hard-coded inductive priors lent it the sample-efficiency it needed to learn something generalizable for the relatively sparse protein database. There, the priors were baked in a specialized architecture, but humans are sample-efficient more broadly. So perhaps there is an architecture that is just more sample efficient without needing ordered training or hierarchical priors.
What about replay: our brain does replay experiences, which nominally mitigates catastrophic interference. Although the replay seems to prioritize surprise/novelty rather than already-learned experiences.