Models would probably do quite well with narrow prompts like True/False checks. We were considering incorporating this, but ultimately prioritised the arguably more realistic open-ended setting, with no ex ante knowledge of the incident. We think identifying which findings are actually worth reporting is very important here.
It might also be very useful to have investigator agents which already have some big picture understanding of the incident and go collect evidence. My guess is agents are already pretty good at this and it’s not the bottleneck.
Models would probably do quite well with narrow prompts like True/False checks. We were considering incorporating this, but ultimately prioritised the arguably more realistic open-ended setting, with no ex ante knowledge of the incident. We think identifying which findings are actually worth reporting is very important here.
It might also be very useful to have investigator agents which already have some big picture understanding of the incident and go collect evidence. My guess is agents are already pretty good at this and it’s not the bottleneck.