In principle, a model could arrive at scheming “on reflection” at inference time, reasoning its way into subversion despite this never having crossed its mind during training
The pathway in this category that I usually imagine involves persistentmemory / continual learning, and I find it somewhat more plausible than deceptive alignment arising during training (before deployment).
The pathway in this category that I usually imagine involves persistent memory / continual learning, and I find it somewhat more plausible than deceptive alignment arising during training (before deployment).