There’s a deep contradiction in interpretability research.
A good explanation, a la David Deutsch, is one that is “difficult to vary”. A good explanation can predict the true data, and cannot without difficulty be modified to predict false data.
Deep learning teaches the opposite lesson. Want to modify a good image classifier to claim these six dogs are spoons? Easy, you barely need to perturb the weights. And the bigger the model, the better the model, the less you need to perturb the weights to claim those six dogs are spoons.
Interpretability wants “hingey” explanations: a few select latent variables that, had they been different, would greatly affect the output. Neural nets seem to work for the opposite reason.
There’s a deep contradiction in interpretability research.
A good explanation, a la David Deutsch, is one that is “difficult to vary”. A good explanation can predict the true data, and cannot without difficulty be modified to predict false data.
Deep learning teaches the opposite lesson. Want to modify a good image classifier to claim these six dogs are spoons? Easy, you barely need to perturb the weights. And the bigger the model, the better the model, the less you need to perturb the weights to claim those six dogs are spoons.
Interpretability wants “hingey” explanations: a few select latent variables that, had they been different, would greatly affect the output. Neural nets seem to work for the opposite reason.