Thank you, this is very meaningful to hear.
I tend to think in simple intuition pumps (like the single stream of text). I think most other researchers do too, but it can be intimidating to lead with broader intuition, since it opens more surface area to criticism. It’s safer and more defensible to focus on results and methodology without narrative.
Writing is also just hard! This post took around 30h to write.
I suspect heavy adversarial training leads models to no longer implicitly trust their CoT, and to downweight role privileges in general. Anecdotally, there seems to be strong correlation between an LLM’s prefill tampering awareness and its prompt injection defense. I think this is ultimately a bad-for-safety solution—it leads to models not faithfully behaving in line with their verbalized reasoning, and similarly erodes control/interpretability for other roles.