I will make a prediction before I perform a quick experiment to test my reasoning.
Pre-experimental reasoning: After training, G’s outputs are normalized while P’s outputs are unnormalized. This means that the gradient of for pulls in various directions; only remains stationary after training because all of the directions cancel out. Before are fully trained, the gradient of pushes away from . But after training, since and is normalized, the gradient of the parameters of will be approximately zero. If I remember correctly, I think that I set up with a higher learning rate so that it can maneuver around , but a better experimental setup would be to train more than so that has more opportunity (but less ability due to a smaller parameter size) to learn.
Post-experimental reasoning:
In my experiments, has very high loss levels (about 1 for mean loss which means does not learn much) when is constantly trained to evade from . When is initially trained to evade the probe but later continues to be trained while is frozen, obtains a loss of about 0.77. If is not trained to evade at all, then obtains a loss of 0.63. But if we freeze after training, reinitialize and then retrain to probe , then attains a loss of about 0.73. But if we freeze after training, and downgrade to a linear network, then won’t be able to learn anything about . Of course, the results depend greatly on the architecture. For example, if we have loss (an unusual loss function that I only use for not too deep networks), then the network has a great ability to evade probes and the probe still won’t be able to learn after retraining while is frozen.
It therefore appears as if probe evasion training works best for polynomial probes or inherently interpretable probes, so prove evasion training might still be used to hamper interpretability.
Probe evasion training might also work best for polynomial probes because polynomial probes are linearly dependent, so the network only needs to be able to evade finitely many probes. The linear dependence of polynomials is probably why it is easier to evade polynomial probes.
I will make a prediction before I perform a quick experiment to test my reasoning.
Pre-experimental reasoning: After training, G’s outputs are normalized while P’s outputs are unnormalized. This means that the gradient of for pulls in various directions; only remains stationary after training because all of the directions cancel out. Before are fully trained, the gradient of pushes away from . But after training, since and is normalized, the gradient of the parameters of will be approximately zero. If I remember correctly, I think that I set up with a higher learning rate so that it can maneuver around , but a better experimental setup would be to train more than so that has more opportunity (but less ability due to a smaller parameter size) to learn.
Post-experimental reasoning:
In my experiments, has very high loss levels (about 1 for mean loss which means does not learn much) when is constantly trained to evade from . When is initially trained to evade the probe but later continues to be trained while is frozen, obtains a loss of about 0.77. If is not trained to evade at all, then obtains a loss of 0.63. But if we freeze after training, reinitialize and then retrain to probe , then attains a loss of about 0.73. But if we freeze after training, and downgrade to a linear network, then won’t be able to learn anything about . Of course, the results depend greatly on the architecture. For example, if we have loss (an unusual loss function that I only use for not too deep networks), then the network has a great ability to evade probes and the probe still won’t be able to learn after retraining while is frozen.
It therefore appears as if probe evasion training works best for polynomial probes or inherently interpretable probes, so prove evasion training might still be used to hamper interpretability.
Probe evasion training might also work best for polynomial probes because polynomial probes are linearly dependent, so the network only needs to be able to evade finitely many probes. The linear dependence of polynomials is probably why it is easier to evade polynomial probes.