that seems within the spirit of what we’re trying to measure (a model could have tools like THINK(bread) and THINK_HARD(bread),
This makes sense to me. In this case, it seems like a stretch to claim that this is a “benchmark to measure how well models can control their activations”, because these results seem consistent with a weaker mechanism where different prompts simply induce different task-conditioned representations.
I guess I don’t precisely understand the difference here. When instructed, models have a different representation conditioning on this instruction, and they can discriminate pretty well between instructions that we would interpret as being about “controlling your activations” and ones that say the opposite, or semantically unrelated ones. Similarly, some models can produce differently rephrased CoTs depending on how hard the task is and the presence of an instruction to phrase the CoT differently, so in this way, models have control over their CoT in a way that’s easy to verify.
I guess you could imagine a requirement here that’s like, it must use it naturally in the course of pursuing another task, which would be harder to elicit. But this is closer to propensity than capability, IMO.
This makes sense to me. In this case, it seems like a stretch to claim that this is a “benchmark to measure how well models can control their activations”, because these results seem consistent with a weaker mechanism where different prompts simply induce different task-conditioned representations.
I guess I don’t precisely understand the difference here. When instructed, models have a different representation conditioning on this instruction, and they can discriminate pretty well between instructions that we would interpret as being about “controlling your activations” and ones that say the opposite, or semantically unrelated ones. Similarly, some models can produce differently rephrased CoTs depending on how hard the task is and the presence of an instruction to phrase the CoT differently, so in this way, models have control over their CoT in a way that’s easy to verify.
I guess you could imagine a requirement here that’s like, it must use it naturally in the course of pursuing another task, which would be harder to elicit. But this is closer to propensity than capability, IMO.