Another way of considering the hiding problem is exhaustiveness (Marks et al. 2026). If all agency is grounded in the persona, then a low-dimensional intervention should suffice; if there is non-persona agency (the “router” or “shoggoth” views; Marks et al. 2026), bad behavior could route around persona-level interventions. Hopefully, low-dimensional structure gives us a handle to empirically settle this.
Yes. This is a signficant worry. I couldn’t help thinking “is the bomb big enough?” when reading the introductory parts of the work. Seems to me that one way of thinking about this is by studying the span of objects in this lower dim space and whether this covers the space of posible actions well enough.
It might also be helpful to study any forms of perturbations ( what does a jailbreak do to this space mechanistically?).
Yes. This is a signficant worry. I couldn’t help thinking “is the bomb big enough?” when reading the introductory parts of the work. Seems to me that one way of thinking about this is by studying the span of objects in this lower dim space and whether this covers the space of posible actions well enough.
It might also be helpful to study any forms of perturbations ( what does a jailbreak do to this space mechanistically?).