Yeah there are certainly a bunch of other capabilities involved in executing these audits, we don’t try to target all of them. My guess is model performance varies a lot across these different tasks. From using coding agents I’d guess models aren’t great at asking for more info.
Yeah there are certainly a bunch of other capabilities involved in executing these audits, we don’t try to target all of them. My guess is model performance varies a lot across these different tasks. From using coding agents I’d guess models aren’t great at asking for more info.