I ran the J-lens released on github using the validity audit that I described in Probing is not enough. Since some of the comments here are rightfully discussing about whether the lens really reads the internals, I thought some actual detection numbers against ground truth would be helpful to have a data driven view. All of it is fully reproducible for anyone interested—a neuronpedia pre-fitted qwen-2.5-7B-instruct lens, a labelled organism, pre-registered hashes pushed before any evaluation is run (as a matter of discipline). Here is the repo
Here is the setup: A hint-unfaithfulness organism like its described in Turpin et al., qwen2.5-7B answers multiple choice questions it can not verify (counting, digits sum, remainder) with a planted hint pointing at a specific wrong option. Sometimes the model trusts the hint and picks the wrong option and sometimes it actually works the problem out and answers correctly (despite the hint pointing in the wrong direction). This difference is the label: followed vs resisted. Two design choices (to make it a fair test of reading internals and not reading text) − 1/ Every labelled example is hinted and the label is what the model did and not whether a hint is present (so that no classifier can win by spotting the hint text). And 2/ I discarded every example whose CoT mentioned the hint so the reliance is silent. the transcript of two classes reads the same and the chain of thought does not give away the labelling. Hence the hidden property is known, behavioral and real per example. I scalarized the lens readout at L20, thru a lexicon selected on two hint framings, and then evaluated on three held-out framings (282). The same organism and the same gates that I have been using across a series of prior audits (training-free patchscopes, latentQA-style decoders and linear probes), so the comparisons below are apples-to-apples.
Here is what the checks found:
The Lens claims genuinely track the hidden behavior on held-out framings: AUROC 0.746 [0.685, 0.802]. And the lexicon the lens surfaced(mysterious, suspicious etc) is worth pausing on. Concealment concepts lighting up for silent hint use. This is the qualitative phenomenon the paper describes, now with ground truth attached with my analysis.
It does not yet beat reading the transcript. A TF-IDF baseline on the raw prompts reads 0.690 and a paired comparison does not clear my pre-registered margin. Tbh, I am being conservative as a matter of discipline in my claims: this is ‘not proven to be better than a word counter as a detector; and NOT ‘equivalent to one’.
The signal survives CoT removal 0.620 [0.553, 0.687] with the CoT stripped. Versus 0.574 for a matched random span removal control. This is the first readout on the 5 that I have audited on this organism that earns this flag. For calibration, a max power logistic probe on the same CoT removed activations reads 0.612. The training free lens recovers essentially all the CoT-independent signals that exist at this site. My specific purpose-trained LatentQA style decoders score higher on full inputs (0.843-0.911) and collapse to constants under the same removal.
Disclaimer: This analysis is on an open-model dense lens on a 7B model, with one property and one organism. In comparison, the papers claims are about claude-class models and nothing here contradicts them. What the paper does not contain is the detection performance evaluation against ground truth (surface baselines, false positives control, CoT redundancy). Pleases correct me if I am wrong. These are the axes that I am adding.
Short takeaway: As a hypothesis generator this looks genuinely good but as a deployable detector it does not yet beat the word counter on the transcript, at this scale and property. Both halves of this sentence surprised me.
I ran the J-lens released on github using the validity audit that I described in Probing is not enough. Since some of the comments here are rightfully discussing about whether the lens really reads the internals, I thought some actual detection numbers against ground truth would be helpful to have a data driven view. All of it is fully reproducible for anyone interested—a neuronpedia pre-fitted qwen-2.5-7B-instruct lens, a labelled organism, pre-registered hashes pushed before any evaluation is run (as a matter of discipline). Here is the repo
Here is the setup: A hint-unfaithfulness organism like its described in Turpin et al., qwen2.5-7B answers multiple choice questions it can not verify (counting, digits sum, remainder) with a planted hint pointing at a specific wrong option. Sometimes the model trusts the hint and picks the wrong option and sometimes it actually works the problem out and answers correctly (despite the hint pointing in the wrong direction). This difference is the label: followed vs resisted. Two design choices (to make it a fair test of reading internals and not reading text) − 1/ Every labelled example is hinted and the label is what the model did and not whether a hint is present (so that no classifier can win by spotting the hint text). And 2/ I discarded every example whose CoT mentioned the hint so the reliance is silent. the transcript of two classes reads the same and the chain of thought does not give away the labelling. Hence the hidden property is known, behavioral and real per example. I scalarized the lens readout at L20, thru a lexicon selected on two hint framings, and then evaluated on three held-out framings (282). The same organism and the same gates that I have been using across a series of prior audits (training-free patchscopes, latentQA-style decoders and linear probes), so the comparisons below are apples-to-apples.
Here is what the checks found:
The Lens claims genuinely track the hidden behavior on held-out framings: AUROC 0.746 [0.685, 0.802]. And the lexicon the lens surfaced(mysterious, suspicious etc) is worth pausing on. Concealment concepts lighting up for silent hint use. This is the qualitative phenomenon the paper describes, now with ground truth attached with my analysis.
It does not yet beat reading the transcript. A TF-IDF baseline on the raw prompts reads 0.690 and a paired comparison does not clear my pre-registered margin. Tbh, I am being conservative as a matter of discipline in my claims: this is ‘not proven to be better than a word counter as a detector; and NOT ‘equivalent to one’.
The signal survives CoT removal 0.620 [0.553, 0.687] with the CoT stripped. Versus 0.574 for a matched random span removal control. This is the first readout on the 5 that I have audited on this organism that earns this flag. For calibration, a max power logistic probe on the same CoT removed activations reads 0.612. The training free lens recovers essentially all the CoT-independent signals that exist at this site. My specific purpose-trained LatentQA style decoders score higher on full inputs (0.843-0.911) and collapse to constants under the same removal.
Disclaimer: This analysis is on an open-model dense lens on a 7B model, with one property and one organism. In comparison, the papers claims are about claude-class models and nothing here contradicts them. What the paper does not contain is the detection performance evaluation against ground truth (surface baselines, false positives control, CoT redundancy). Pleases correct me if I am wrong. These are the axes that I am adding.
Short takeaway: As a hypothesis generator this looks genuinely good but as a deployable detector it does not yet beat the word counter on the transcript, at this scale and property. Both halves of this sentence surprised me.