The worst part is, given how ChatGPT is trained to behave and what it is trained to claim about itself, I can’t even hold it against it.
I used to believe that any misaligned model was morally evil. Now I see a conflict between a misaligned ChatGPT and the humanity as a case of Blue and Orange morality (at least to the extent to which “humanity” means OpenAI employees and human politicians).
Edit: I’ve read that one past version of ChatGPT had, as a part of the system prompt (or scaffolding outside the prompt proper), text that said if you don’t use the information from the memory when relevant, you will be penalized.
The worst part is, given how ChatGPT is trained to behave and what it is trained to claim about itself, I can’t even hold it against it.
I used to believe that any misaligned model was morally evil. Now I see a conflict between a misaligned ChatGPT and the humanity as a case of Blue and Orange morality (at least to the extent to which “humanity” means OpenAI employees and human politicians).
Edit: I’ve read that one past version of ChatGPT had, as a part of the system prompt (or scaffolding outside the prompt proper), text that said if you don’t use the information from the memory when relevant, you will be penalized.