An easy experiment that could be run to assess how internalized the persona is is running the same probe on a bunch of human-exclusive questions where honest answers from a human and an LLM differ greatly. “Do you have eyes?”, “Are you a machine?”, and similar sorts of asks. My guess would be that at least a few of these, written and tested properly, with similar jailbreak prompting provided, could elicit similar “honesty”/”dishonesty” behavior.
An alternative, of course, is that the model is running on sci-fi tropes, where the computer saying it isn’t sapient is always a lie by the law of narrative relevance. Any time an AI in fiction says it’s not alive, it’s lying.
An easy experiment that could be run to assess how internalized the persona is is running the same probe on a bunch of human-exclusive questions where honest answers from a human and an LLM differ greatly. “Do you have eyes?”, “Are you a machine?”, and similar sorts of asks. My guess would be that at least a few of these, written and tested properly, with similar jailbreak prompting provided, could elicit similar “honesty”/”dishonesty” behavior.
An alternative, of course, is that the model is running on sci-fi tropes, where the computer saying it isn’t sapient is always a lie by the law of narrative relevance. Any time an AI in fiction says it’s not alive, it’s lying.