Great post! Regarding “Roles as a cognitive window” in Auditing Language Models for Hidden Objectives, the authors sample from an LLM starting from the user[1] tag as a way to extract information from the model that it otherwise kept secret. This technique allowed them to audit a model organism which was trained to have a secret goal (being sycophantic to reward models).
Great post! Regarding “Roles as a cognitive window” in Auditing Language Models for Hidden Objectives, the authors sample from an LLM starting from the user[1] tag as a way to extract information from the model that it otherwise kept secret. This technique allowed them to audit a model organism which was trained to have a secret goal (being sycophantic to reward models).
I think in their case “human” since it’s claude