For example, having more eyes on their SFT data or RL environments really could save significant amounts of money, if it helps catch a broken reward model a few days earlier. I’m a bit skeptical of this argument. The properties that developers have a clear incentive to care about (e.g. reward-hacking that impedes capabilities training), are also the ones where they already have a large comparative advantage over any external party, both in debugging expertise and in context on their own stack.
Thanks for this, I agree with almost all of it, and strongly with the central claim that final-checkpoint evals are not sufficient and that we should be building the third-party TRA ecosystem now. I am currently writing something on why RL environment audits in particular should involve third parties, so I would love your feedback on it.
And one small pushback, on this:
I do not think this is necessarily true. Developers do search hard for hack paths and find many of them, but I expect they still miss a lot, since a single internal team shares its own blind spots and cannot cover its own tail just by searching harder. In security, organizations with strong internal teams still run external penetration tests and bug bounties for exactly this reason. Security and AI development are of course very different, but I suspect developers could benefit from third-party environment audits even in the cases they are already incentivized to look at, not only the low-incentive ones.
Thanks for this, I agree with almost all of it, and strongly with the central claim that final-checkpoint evals are not sufficient and that we should be building the third-party TRA ecosystem now. I am currently writing something on why RL environment audits in particular should involve third parties, so I would love your feedback on it.
And one small pushback, on this:
I do not think this is necessarily true. Developers do search hard for hack paths and find many of them, but I expect they still miss a lot, since a single internal team shares its own blind spots and cannot cover its own tail just by searching harder. In security, organizations with strong internal teams still run external penetration tests and bug bounties for exactly this reason. Security and AI development are of course very different, but I suspect developers could benefit from third-party environment audits even in the cases they are already incentivized to look at, not only the low-incentive ones.