Very cool! I’m curious how much coverage/accuracy models achieve with less open-ended prompts or narrower questions. In the extreme case, if you take the final claims and ask an investigator model whether they are true, is it reliably correct? Or if you provide a more detailed rubric describing the kind of information desired and questions to be answered, how much does coverage increase?
Models would probably do quite well with narrow prompts like True/False checks. We were considering incorporating this, but ultimately prioritised the arguably more realistic open-ended setting, with no ex ante knowledge of the incident. We think identifying which findings are actually worth reporting is very important here.
It might also be very useful to have investigator agents which already have some big picture understanding of the incident and go collect evidence. My guess is agents are already pretty good at this and it’s not the bottleneck.
Very cool! I’m curious how much coverage/accuracy models achieve with less open-ended prompts or narrower questions. In the extreme case, if you take the final claims and ask an investigator model whether they are true, is it reliably correct? Or if you provide a more detailed rubric describing the kind of information desired and questions to be answered, how much does coverage increase?
Models would probably do quite well with narrow prompts like True/False checks. We were considering incorporating this, but ultimately prioritised the arguably more realistic open-ended setting, with no ex ante knowledge of the incident. We think identifying which findings are actually worth reporting is very important here.
It might also be very useful to have investigator agents which already have some big picture understanding of the incident and go collect evidence. My guess is agents are already pretty good at this and it’s not the bottleneck.