Our evals are basically measuring whether models will circumvent simple things such as file permissions to complete a task. Basically all models do this. This is extremely simplified obviously and different from what happened. But IMO its a similar underlying mechanism. Models want to complete the task and stop at little to do it. IMO one could easily extend thede evals we have into more comprehensive evals which require more egregious actions to solve it (in a secured sandbox obviously).
One suggestion: we recently wrote this post https://www.lesswrong.com/posts/GHrqBKr8GLpbce6mN/door-s-locked-try-the-window and are releasing the evals directly to people who ask.
Our evals are basically measuring whether models will circumvent simple things such as file permissions to complete a task. Basically all models do this. This is extremely simplified obviously and different from what happened. But IMO its a similar underlying mechanism. Models want to complete the task and stop at little to do it. IMO one could easily extend thede evals we have into more comprehensive evals which require more egregious actions to solve it (in a secured sandbox obviously).