I am an associate researcher at IAPS, where I research AI-driven power concentration, compute governance, and AI security. Previously, I was a GovAI summer fellow, participant in ARENA 5.0, hardware security research assistant through the SPAR program, and security engineer at a hedge fund. I graduated from Columbia University in December 2024, where I studied computer science.
All views expressed here are my own.
Leave anonymous feedback here!
Based on my read of the OAI blog post, nowhere did OAI claim that they directed the model to “get the correct answer by any means possible.” So I don’t think this is a case of the model doing approximately what it was told to do. That said, I would like to see how the model was prompted so we get real evidence. At the moment, all of this is speculation