Pretty cool read. I wonder how much of this is testing investigation vs reconstruction. In the METR case, a lot of the hard part seemed to be noticing the initial dataset was incomplete, asking for more data, checking provenance, dealing with spoofed or missing logs, and updating the story as new evidence came in.
Here, the model mostly gets a fixed dataset and is scored on recovering findings from the final human report. Would the results look very different if it had to decide what evidence was missing, what to request next, and how much to trust the logs?
This is pretty Cool! Maybe continuation messages should be treated as part of the eval itself rather than as a neutral implementation detail. Also, I wonder if the effect holds up for models like Astra and Fable.