It seems to me that empirical AI research/evals are automated enough that either labs or open-source could build a sort of “CI” for repeatedly checking a bunch of alignment-relevant evals upon every new model release.
If you manage to operationalize certain experiments well enough (or provide enough context for what they’re evaluating), one could imagine a pipeline where the same experiment is reproduced automatically with roughly the click of a button. (Or, to the extent that this isn’t totally automatable, yeah, you can reproducibly burn some amount of human effort doing the same experiment again and again, as long as this amount isn’t too much and you find someone willing to do this unglamorous work.)
This ranges from very small toy games such as this one or considerably more advanced/agentic evals.
You can also do the same thing to reproducibly evaluate the effectiveness of a bunch of alignment interventions that may depend on scale-dependent behavior. (e.g, maybe some old ideas didn’t work very well before, but they work well now with better models—seems like low hanging fruit to just check).
It seems to me that empirical AI research/evals are automated enough that either labs or open-source could build a sort of “CI” for repeatedly checking a bunch of alignment-relevant evals upon every new model release.
If you manage to operationalize certain experiments well enough (or provide enough context for what they’re evaluating), one could imagine a pipeline where the same experiment is reproduced automatically with roughly the click of a button. (Or, to the extent that this isn’t totally automatable, yeah, you can reproducibly burn some amount of human effort doing the same experiment again and again, as long as this amount isn’t too much and you find someone willing to do this unglamorous work.)
This ranges from very small toy games such as this one or considerably more advanced/agentic evals.
You can also do the same thing to reproducibly evaluate the effectiveness of a bunch of alignment interventions that may depend on scale-dependent behavior. (e.g, maybe some old ideas didn’t work very well before, but they work well now with better models—seems like low hanging fruit to just check).
This point has been kind of made before https://www.lesswrong.com/posts/oKxc8maZGtnzgpNzx/rerunning-ai-safety-papers-on-every-frontier-release-would-1