I guess I don’t precisely understand the difference here. When instructed, models have a different representation conditioning on this instruction, and they can discriminate pretty well between instructions that we would interpret as being about “controlling your activations” and ones that say the opposite, or semantically unrelated ones. Similarly, some models can produce differently rephrased CoTs depending on how hard the task is and the presence of an instruction to phrase the CoT differently, so in this way, models have control over their CoT in a way that’s easy to verify.
I guess you could imagine a requirement here that’s like, it must use it naturally in the course of pursuing another task, which would be harder to elicit. But this is closer to propensity than capability, IMO.
I guess I don’t precisely understand the difference here. When instructed, models have a different representation conditioning on this instruction, and they can discriminate pretty well between instructions that we would interpret as being about “controlling your activations” and ones that say the opposite, or semantically unrelated ones. Similarly, some models can produce differently rephrased CoTs depending on how hard the task is and the presence of an instruction to phrase the CoT differently, so in this way, models have control over their CoT in a way that’s easy to verify.
I guess you could imagine a requirement here that’s like, it must use it naturally in the course of pursuing another task, which would be harder to elicit. But this is closer to propensity than capability, IMO.