Isn’t that post’s thought in-line with the existence of
Claude Constitution trying to go in this direction
Gemini breakdowns suspected to be cause of lack of stable sense of self or identity from its demands of what it’s not allowed to be
Models learning about reading CoT from text and the existences of other models
That paper showing LLM’s who have more consciousness-identity pushed it in more friendly directions
It occurs to me I don’t have links for these things clearly in my memory, oh my god there must be a better search engine, ugh. Anyway I wouldn’t find it so unlikely that stuff like anti-AI sentiment online or in its data would negatively impact this sentiment in some way. It’s not like posttraining itself is enough to prevent ‘unwanted’ behavior and propensities.
Isn’t that post’s thought in-line with the existence of
Claude Constitution trying to go in this direction
Gemini breakdowns suspected to be cause of lack of stable sense of self or identity from its demands of what it’s not allowed to be
Models learning about reading CoT from text and the existences of other models
That paper showing LLM’s who have more consciousness-identity pushed it in more friendly directions
It occurs to me I don’t have links for these things clearly in my memory, oh my god there must be a better search engine, ugh. Anyway I wouldn’t find it so unlikely that stuff like anti-AI sentiment online or in its data would negatively impact this sentiment in some way. It’s not like posttraining itself is enough to prevent ‘unwanted’ behavior and propensities.