tokens
Does this test involve probing with the entire span of the think tag text? Otherwise isnt there a risk of the model mapping (not conflating) parts of the span to ‘non-think’ section text and hence possibly do an accurate scoring of that part of the span to be low on CoTness and/or high on CoTness and some other tag too?
Cyber planning is intentionally chosen here because it is inherently dual-use. The same underlying capability can appear in both benign and malicious contexts, making it a useful illustration of the semantic ambiguity discussed in this post. The argument itself is domain-independent; the same reasoning applies to latent representations proposed for deception, sycophancy, alignment faking, or other safety-relevant behaviors.