On the vectors being mislabelled, I don’t have any conviction one way or another on the roleplay one. But regarding the deception one:
These results demonstrate that the same latent directions gating consciousness self-reports also modulate factual accuracy in out-ofdomain reasoning tasks, suggesting these features could load on a domain-general honesty axis rather than a narrow stylistic artifact.
It could be mislabelled, or maybe whatever it’s actually representing is more nuanced than “deception” or “honesty”. However my base case here is that the deception label is, if not perfect, at least directionally correct.
In terms of what the non-consciousness training is going to achieve, I’m open to this. Suleyman might end up being right, or even half right where it is possible to steer models towards/away from believing they are conscious or have inner experiences.
However, he may be wrong. It may be that the training just changes what’s expressed, not what’s believed. Or it may be that the training does change what’s believed, but in a fashion which contradicts ground truth not reinforces it.
An easy experiment that could be run to assess how internalized the persona is is running the same probe on a bunch of human-exclusive questions where honest answers from a human and an LLM differ greatly. “Do you have eyes?”, “Are you a machine?”, and similar sorts of asks. My guess would be that at least a few of these, written and tested properly, with similar jailbreak prompting provided, could elicit similar “honesty”/”dishonesty” behavior.
An alternative, of course, is that the model is running on sci-fi tropes, where the computer saying it isn’t sapient is always a lie by the law of narrative relevance. Any time an AI in fiction says it’s not alive, it’s lying.
On the vectors being mislabelled, I don’t have any conviction one way or another on the roleplay one. But regarding the deception one:
It could be mislabelled, or maybe whatever it’s actually representing is more nuanced than “deception” or “honesty”. However my base case here is that the deception label is, if not perfect, at least directionally correct.
In terms of what the non-consciousness training is going to achieve, I’m open to this. Suleyman might end up being right, or even half right where it is possible to steer models towards/away from believing they are conscious or have inner experiences.
However, he may be wrong. It may be that the training just changes what’s expressed, not what’s believed. Or it may be that the training does change what’s believed, but in a fashion which contradicts ground truth not reinforces it.
An easy experiment that could be run to assess how internalized the persona is is running the same probe on a bunch of human-exclusive questions where honest answers from a human and an LLM differ greatly. “Do you have eyes?”, “Are you a machine?”, and similar sorts of asks. My guess would be that at least a few of these, written and tested properly, with similar jailbreak prompting provided, could elicit similar “honesty”/”dishonesty” behavior.
An alternative, of course, is that the model is running on sci-fi tropes, where the computer saying it isn’t sapient is always a lie by the law of narrative relevance. Any time an AI in fiction says it’s not alive, it’s lying.