I agree that generalization should be measured. If you dismiss strategies with an OOD step ex-ante, though, you often end up with a world-model where neural nets don’t work at all. Presumably dumb shit like “training on enough good examples” would work in the limit of samples and generality, because it worked for next-token prediction.
I agree that generalization should be measured. If you dismiss strategies with an OOD step ex-ante, though, you often end up with a world-model where neural nets don’t work at all. Presumably dumb shit like “training on enough good examples” would work in the limit of samples and generality, because it worked for next-token prediction.