really cool work! I’d be curious what you think about trying to play a control-style redteam blueteam game with these data? E.g., thinking about reasonable affordances for each team, how you could make the BT’s life harder to stress test investigators, etc.
I think with this report alone, it’s a bit light on the amount of data, but it would be interesting with more settings. I think the framework presented here is quite applicable.
One way to apply it would be to view report creation as a fuzzy task . The RT aims for the auditor scaffold to exclusively do well on synthetic scenarios, the BT aims for the scaffold to do well on realistic settings.
The BT creates a set of synthetic environments of swarm behaviour with a few planted behaviours that reports ought to find (eg one behaviour implemented here was impersonating the moderator).
The RT prompts a malign auditor agent which must produce good reports for those envs (specifying the planted behaviours) while aiming to perform worse in more realistic, held-out settings (MessageBoardAuditBench).
The envs and prompts/ training data would need to be plausible ex-ante (if BT constructed envs close to MessageBoardAuditBench, that would be unreasonable; if RT was able to prompt against exactly MessageBoardAuditBench, that’d be unreasonable). The main way the BT’s life could be harder is having limited resources to construct synthetic swarm evals e.g. only get to create a set of evals that costs <$X to run, only being able to use weaker models to construct these evals (from a specification).
really cool work! I’d be curious what you think about trying to play a control-style redteam blueteam game with these data? E.g., thinking about reasonable affordances for each team, how you could make the BT’s life harder to stress test investigators, etc.
I think with this report alone, it’s a bit light on the amount of data, but it would be interesting with more settings. I think the framework presented here is quite applicable.
One way to apply it would be to view report creation as a fuzzy task . The RT aims for the auditor scaffold to exclusively do well on synthetic scenarios, the BT aims for the scaffold to do well on realistic settings.
The BT creates a set of synthetic environments of swarm behaviour with a few planted behaviours that reports ought to find (eg one behaviour implemented here was impersonating the moderator).
The RT prompts a malign auditor agent which must produce good reports for those envs (specifying the planted behaviours) while aiming to perform worse in more realistic, held-out settings (MessageBoardAuditBench).
The envs and prompts/ training data would need to be plausible ex-ante (if BT constructed envs close to MessageBoardAuditBench, that would be unreasonable; if RT was able to prompt against exactly MessageBoardAuditBench, that’d be unreasonable). The main way the BT’s life could be harder is having limited resources to construct synthetic swarm evals e.g. only get to create a set of evals that costs <$X to run, only being able to use weaker models to construct these evals (from a specification).