<notice>Curated! This post contains both (1) superbly readable and accessible explanation that makes me feel like I understand something I thought I already did to a yet deeper degree, (2) elegant novel experiment to highlight their points. I have definitely updated on the significance of roles from this post.
Role tags were a formatting trick that became the security architecture and the cognitive scaffolding of modern LLMs.
I’d also say the point of the steady evolution of roles and how they’re used (and leave open vulnerability) feels like it says something about the way our civilization is developing AI, and not in a positive way.
Overall, kudos for this great research and post!</notice>
When I first glanced at this comment, I was really hoping that it would turn out to be by the original author, attempting a role-spoofing attack on the LW admins. Alas, unless something really devious is going on it is in fact by an LW admin. (But maybe the <notice>...</notice> thing is, even so, a deliberate nod to the subject matter of the post?)
Thanks for the curation! And yeah, there’s so many little architectural decisions that seem inconsequential now, but could heavily define how humans are allowed to trade-off agency against AI in the future.
<notice>Curated! This post contains both (1) superbly readable and accessible explanation that makes me feel like I understand something I thought I already did to a yet deeper degree, (2) elegant novel experiment to highlight their points. I have definitely updated on the significance of roles from this post.
I’d also say the point of the steady evolution of roles and how they’re used (and leave open vulnerability) feels like it says something about the way our civilization is developing AI, and not in a positive way.
Overall, kudos for this great research and post!</notice>
When I first glanced at this comment, I was really hoping that it would turn out to be by the original author, attempting a role-spoofing attack on the LW admins. Alas, unless something really devious is going on it is in fact by an LW admin. (But maybe the <notice>...</notice> thing is, even so, a deliberate nod to the subject matter of the post?)
[EDITED because I accidentally a word]
Thanks for the curation! And yeah, there’s so many little architectural decisions that seem inconsequential now, but could heavily define how humans are allowed to trade-off agency against AI in the future.