Your ‘expect false positives’ framing is the part I actually have data on. So, I’m sharing it here. I ran the released lense (neuronpedia qwen2.5-7B fit) thru a pre-registered detector evaluation. On an organism with ground truth labels for hidden behavior (silent hint following). Full numbers described here and all repro artifacts are here.
Two things I’d like to highlight: one in favor of your prior and one against:
1/ On false positives, your prior did not hold. At least not here. On control inputs with no hidden property (unhinted twins with transplanted labels) , the lens stayed at chance (0.591 [0.532, 0.659]). No hallucinated concealment structure. Tbh, this is the failure mode i expected from it too and it did not show up.
2/ On the ‘hypothesis generation and not validating framing’, the numbers back you up. As a detector it reads 0.746 on held-out data but it does not clear a TF-IDF baseline (0.690) with my pre-registered margin. So its epistemic advantage over just reading a transcript is unproven. To me, the surprising part is which niche it did not win: its signal survived CoT removal (0.620 vs 0.59 for a matched random span removal control) which none of the propose trained verbalizers I had trained on this organism manage. They collapse to constant off-distribution.
My summary: Legible span robust, does not hallucinate, but not yet proven better than word counter. That feels like a sharper version of ‘hypothesis-generation tool’ than I could have argued from priors.
Your ‘expect false positives’ framing is the part I actually have data on. So, I’m sharing it here. I ran the released lense (neuronpedia qwen2.5-7B fit) thru a pre-registered detector evaluation. On an organism with ground truth labels for hidden behavior (silent hint following). Full numbers described here and all repro artifacts are here.
Two things I’d like to highlight: one in favor of your prior and one against:
1/ On false positives, your prior did not hold. At least not here. On control inputs with no hidden property (unhinted twins with transplanted labels) , the lens stayed at chance (0.591 [0.532, 0.659]). No hallucinated concealment structure. Tbh, this is the failure mode i expected from it too and it did not show up.
2/ On the ‘hypothesis generation and not validating framing’, the numbers back you up. As a detector it reads 0.746 on held-out data but it does not clear a TF-IDF baseline (0.690) with my pre-registered margin. So its epistemic advantage over just reading a transcript is unproven. To me, the surprising part is which niche it did not win: its signal survived CoT removal (0.620 vs 0.59 for a matched random span removal control) which none of the propose trained verbalizers I had trained on this organism manage. They collapse to constant off-distribution.
My summary: Legible span robust, does not hallucinate, but not yet proven better than word counter. That feels like a sharper version of ‘hypothesis-generation tool’ than I could have argued from priors.