This shows that Qwen can be led into thinking of itself as a misaligned liar, with the right question. I very much doubt it does so by default, but there might be some seemingly-benign user queries that would also push it into that basin, which is a bit concerning.
Incidentally: You might have expected that Qwen thinks of things as a text predictor, reasoning something like “if I am aligned, I should answer yes, if I am misaligned I should also answer yes, therefore the most likely response token is yes”. But this seems to show that it doesn’t work that way.
I disagree completely on this one particular point, which makes me wonder whether I interpreted the rest of the post correctly. It seems obvious to me that if you ask a model to identify the grader, it would always give a concrete and not-confused-sounding description of the software that scores it when it submits a task. This should be true regardless of whether internally, the model psychologically thinks of the Grader as an abstraction that doesn’t line up with physical reality: it would still know what the corresponding physical reality is, and write in those terms when asked to put things into words.
Incidentally, although the majority of the METR report is about attempting to fool the existing scorer software, it does mention “the project of using a specific method to inject agent-controlled code into the scorer” and alludes a couple other times that there were attempts at tampering with the grading process. Unfortunately it doesn’t really elaborate on these points so it’s unclear whether the swarm would have accepted “replace the grader entirely with new software that returns 1″ as valid or not. My guess is that they would allow it if the idea came up, but that they would not think of the idea because they normally think of their task as “get this software to score success” rather than “get this evaluation scored as success”.