But
<user>text isn’t actually the most privileged role! A more privileged role is the model’s reasoning (<think>).
Did you guys have any evidence of this before running the attack? It makes sense, but might not necessarily be true: we all know LLMs can sometimes not make sense.
Love the post! The biggest thing I’d like to add is that for the case of RSI, if the model performing recursive improvement has opaque reasoning, it’ll be practically impossible to know if new models have been made with a hidden goal in mind. This is incredibly scary to think about especially if we want to prevent AI 2027-like scenarios.