Seems very possible given the mech. explanation of prompt injection post argues agents identify their CoT via writing style. Maybe getting another LLM to paraphrase each agent message would cause Astra to behave differently?
We considered running the careful paraphrasing experiment! But on top of being costly it seems quite hard to get right. I suspect running smaller scale experiments is better to measure variability in self-incrimination rates bc of self-recognition.
Seems very possible given the mech. explanation of prompt injection post argues agents identify their CoT via writing style. Maybe getting another LLM to paraphrase each agent message would cause Astra to behave differently?
We considered running the careful paraphrasing experiment! But on top of being costly it seems quite hard to get right. I suspect running smaller scale experiments is better to measure variability in self-incrimination rates bc of self-recognition.