Another data point! I wish that we could compare it with Mythos Preview or with GPT-5.6 which METR scolded off for wholesale cheating, and hope that someone at METR (e.g. @GradientDissenter?) comments on this...
Another data point! I wish that we could compare it with Mythos Preview or with GPT-5.6 which METR scolded off for wholesale cheating, and hope that someone at METR (e.g. @GradientDissenter?) comments on this...