Pretty cool read. I wonder how much of this is testing investigation vs reconstruction. In the METR case, a lot of the hard part seemed to be noticing the initial dataset was incomplete, asking for more data, checking provenance, dealing with spoofed or missing logs, and updating the story as new evidence came in.
Here, the model mostly gets a fixed dataset and is scored on recovering findings from the final human report. Would the results look very different if it had to decide what evidence was missing, what to request next, and how much to trust the logs?
Yeah there are certainly a bunch of other capabilities involved in executing these audits, we don’t try to target all of them. My guess is model performance varies a lot across these different tasks. From using coding agents I’d guess models aren’t great at asking for more info.
Pretty cool read. I wonder how much of this is testing investigation vs reconstruction. In the METR case, a lot of the hard part seemed to be noticing the initial dataset was incomplete, asking for more data, checking provenance, dealing with spoofed or missing logs, and updating the story as new evidence came in.
Here, the model mostly gets a fixed dataset and is scored on recovering findings from the final human report. Would the results look very different if it had to decide what evidence was missing, what to request next, and how much to trust the logs?
Yeah there are certainly a bunch of other capabilities involved in executing these audits, we don’t try to target all of them. My guess is model performance varies a lot across these different tasks. From using coding agents I’d guess models aren’t great at asking for more info.