I agree with most of this article. However, I want to point out an objection towards Q6, specifically category C. In my opinion, task misalignment seems very tractable to reliably detect and penalize by a sufficiently capable LLM judge and interpretability tools (you’re using full surveillience and maybe even mind reading to make sure that the model cannot get away with cheating, and detection of cheating in that way intuitively seems much easier than getting away with cheating, so at any given capability level the monitors seem likely to have an advantage, and plausibly this is only more-and-more-so the case as interpretability methods improve). If that works, it would be a valid, non-brittle example of aligning a model to a “trained machine learning model” as described in Category C.
I do agree with the rest of Q6 though. It seems like to get full alignment you have to align a model towards an entire ethics system as the objective function, and that system has to be fairly complex (in the sense of well-thought-out) and internally consistent enough to avoid both Categories A and B. This basically points to a constitutional AI approach like what Anthropic is doing.
In addition I think that because the type of misalignment described in this article is rooted in task misalignment, if someone starts failing at it we (or at least they) will expect to see warning shots like the OpenAI/HF incident before anything truly catastrophic starts to happen. So that could be an optimistic sign that people failling to align AIs will be given a chance to change course.
Another thought just occurred to me, extending from my above point: when it comes to task misalignment, effective (automated) control is effective alignment. If you have reliable enough control to stop a model from going rogue, then you can also stop it from cheating during training, which prevents the task misalignment from happening in the first place.
I agree with most of this article. However, I want to point out an objection towards Q6, specifically category C. In my opinion, task misalignment seems very tractable to reliably detect and penalize by a sufficiently capable LLM judge and interpretability tools (you’re using full surveillience and maybe even mind reading to make sure that the model cannot get away with cheating, and detection of cheating in that way intuitively seems much easier than getting away with cheating, so at any given capability level the monitors seem likely to have an advantage, and plausibly this is only more-and-more-so the case as interpretability methods improve). If that works, it would be a valid, non-brittle example of aligning a model to a “trained machine learning model” as described in Category C.
I do agree with the rest of Q6 though. It seems like to get full alignment you have to align a model towards an entire ethics system as the objective function, and that system has to be fairly complex (in the sense of well-thought-out) and internally consistent enough to avoid both Categories A and B. This basically points to a constitutional AI approach like what Anthropic is doing.
In addition I think that because the type of misalignment described in this article is rooted in task misalignment, if someone starts failing at it we (or at least they) will expect to see warning shots like the OpenAI/HF incident before anything truly catastrophic starts to happen. So that could be an optimistic sign that people failling to align AIs will be given a chance to change course.
Another thought just occurred to me, extending from my above point: when it comes to task misalignment, effective (automated) control is effective alignment. If you have reliable enough control to stop a model from going rogue, then you can also stop it from cheating during training, which prevents the task misalignment from happening in the first place.