I’m excited that Transluce got this piece out! We see behavior evaluations as an important lever on AI governance, and are actively investing in this area with a focus on cognitive security, evaluation gaming, and agentic honesty.
In the course of doing this work, we’ve learned a lot about the challenges, both at a technical level and an ecosystem level, and we’re hoping that sharing some of them will help the rest of the field.
We’ll also have some concrete technical reports out soon; stay tuned!
Hi mods, got a couple weird formatting things above, where all the code blocks display as code twice; not totally sure how to fix it! [Edit: fixed by converting the code to unformatted blocks without the pretty-printing; preserving the comment for posterity.]