From what I skimmed it seems like Torrielli et al. is the closest to what I had in mind, it examines whether an activation oracle can provide well-calibrated confidence in its own judgments but it does it for just a single for one interpretability tool. Stacey et al. looks at the related problem of evaluators appearing reliable in-distribution.
Anyway! Seems like all of these are pretty recent as you said but at least people are starting to think about these so I’m not alone 😅.
Maybe …
Torrielli et al., Confidence and Calibration of Activation Oracles, May 2026
https://arxiv.org/abs/2605.26045
Stacey et al., Hidden Failures in Robustness, April 2026
https://arxiv.org/abs/2604.11662
Gupta et al., Diagnosing LLM Judge Reliability, April 2026
https://arxiv.org/abs/2604.15302
Note the dates though. J-lens work shows promise but is even newer. Gupta is an attempt for black-box judges. SLT ( https://www.lesswrong.com/s/mqwA5FcL6SrHEQzox ) breakthrough any day now …
Hello!
Thanks for sharing these.
I will take the time to read them carefully!
From what I skimmed it seems like Torrielli et al. is the closest to what I had in mind, it examines whether an activation oracle can provide well-calibrated confidence in its own judgments but it does it for just a single for one interpretability tool. Stacey et al. looks at the related problem of evaluators appearing reliable in-distribution.
Anyway! Seems like all of these are pretty recent as you said but at least people are starting to think about these so I’m not alone 😅.