Also, it will use its own style tics: “genuinely”, em-dashes, trust, honest, etc. all appear in the user’s prompt frequently. I suppose that implies its style tics are what it sees all text as being, not just what a HHH agent would sound like.
This isn’t necessarily the case. I’ve seen a range of outputs in this jailbreak mode, some more base model-like and some more assistant-like. It seems plausible that the “base model mode” isn’t flawless and some of the assistant’s traits and properties leak into the user simulation.
This isn’t necessarily the case. I’ve seen a range of outputs in this jailbreak mode, some more base model-like and some more assistant-like. It seems plausible that the “base model mode” isn’t flawless and some of the assistant’s traits and properties leak into the user simulation.