Hi, thanks for this interesting work. You may also be interested in our new work where we investigate whether internal linear probes (before an answer is produced) capture whether a model is going to answer correctly: https://www.lesswrong.com/posts/KwYpFHAJrh6C84ShD/no-answer-needed-predicting-llm-answer-accuracy-from
We also compare that with verbalised self-confidence and we find internals have more predictive power, so potentially you can apply internals to your setup
Really cool work! I guess a step towards more realism is getting traces from real incidents and manipulating them to plan some things, and then seeing if investigator agents can spot those. Do you think this is a promising approach (of course, dependent on being able to actually get access to traces)?