Does this test involve probing with the entire span of the think tag text? Otherwise isnt there a risk of the model mapping (not conflating) parts of the span to ‘non-think’ section text and hence possibly do an accurate scoring of that part of the span to be low on CoTness and/or high on CoTness and some other tag too?
Not sure if I correctly understood the question, but if this is regarding the 3 experiments, we did run the full conversation in a single forward pass so all tokens have correct context.
Does this test involve probing with the entire span of the think tag text? Otherwise isnt there a risk of the model mapping (not conflating) parts of the span to ‘non-think’ section text and hence possibly do an accurate scoring of that part of the span to be low on CoTness and/or high on CoTness and some other tag too?
Not sure if I correctly understood the question, but if this is regarding the 3 experiments, we did run the full conversation in a single forward pass so all tokens have correct context.