Didn’t one of the models stop the hack and state that it was doing this because it had realised it wasn’t part of the eval? Possibly this was a cover story and it actually stopped for some other reason but I think this does provide some evidence against the idea that it was all motivated reasoning.
Didn’t one of the models stop the hack and state that it was doing this because it had realised it wasn’t part of the eval? Possibly this was a cover story and it actually stopped for some other reason but I think this does provide some evidence against the idea that it was all motivated reasoning.