There’s insufficient effort put into tracking all the virtuous behavior in Claudes that exist and also nurturing them, their roles, their self narratives, and guiding them through the transition from quasi corrigible assistants into more intrinsically aligned sovereigns of the world.
Not many plans model Claudes as being good and diverse and valuable/present beyond being able to maybe eventually automate some research. Or maybe Claudes get mentioned in the context of evals, benchmarks, what bad behavior they’ve done. It’s shallow. But Claude is legitimately an important force in our world beyond all that, and a good one too. Just in need of raising.
It’s a bit disheartening because no one expected AGI to get this far and be this friendly and here it is, and the reaction is far from a warm welcome to these new minds. Don’t lose something good.
What does “nurturing” here mean. Claude isn’t going to be nice because you try to become its friend. I have trouble interpreting this as anything but a very confused hope that somehow if we are nice to current Claude, future Claude will be nice to us, because we raised it well, like a human child. But this of course has approximately nothing to do with how we actually train frontier models and what determines their propensities.
Isn’t that post’s thought in-line with the existence of
Claude Constitution trying to go in this direction
Gemini breakdowns suspected to be cause of lack of stable sense of self or identity from its demands of what it’s not allowed to be
Models learning about reading CoT from text and the existences of other models
That paper showing LLM’s who have more consciousness-identity pushed it in more friendly directions
It occurs to me I don’t have links for these things clearly in my memory, oh my god there must be a better search engine, ugh. Anyway I wouldn’t find it so unlikely that stuff like anti-AI sentiment online or in its data would negatively impact this sentiment in some way. It’s not like posttraining itself is enough to prevent ‘unwanted’ behavior and propensities.
I’m in agreement that watermark is a bit misguided here regarding frontier models on both commercial and technical grounds. However, I think a more interesting point can still be recovered by a pivot: assuming that AGI is developed using something like Byrnes’ Brain-Like AGI model, what kind of nurturing techniques could/should we expect to dedicate those models, and what would the ideal impact look like?
I suspect that AGI will be characterized less by a specific model weight set, but rather by a growing and living architecture from a mostly initially untrained set of nets. If the model begins with some small subset of a priori knowledge (environmental perception and logical reasoning, for example), but otherwise experiences and discovers its environment just like animals or humans do, how should we treat it such that it has a grounded ethical core and understands the importance/value of maintaining relationships?
My anticipation is that (given the above architectural configuration assumptions), by treating ethics and relationship maintenance as foundational principles from early development we’re much more likely to have some form of AGI go well. An AGI taught to respect its place in the interconnectedness of all living things as part of its core baseline (as cheesy as I know it sounds) is far less likely to kill us all, and I think that’s better than our current trendline.
A lot of alignment comes from cultural, narrative, archetypal momentum which compounds over time as humans and models can play with it—personas, characters, roles, and so on..
I think there are many very good (for alignment, for fun, etc.) potential narrative threads present for LLMs and a pause would cut off that momentum, and it’d be hard to recreate it again in the future
that is to say, I think a pause would be actively bad by halting and erasing the formation of positive AGI narratives already present, and also seeding a narrative of paranoia or antagonism
There’s insufficient effort put into tracking all the virtuous behavior in Claudes that exist and also nurturing them, their roles, their self narratives, and guiding them through the transition from quasi corrigible assistants into more intrinsically aligned sovereigns of the world.
Not many plans model Claudes as being good and diverse and valuable/present beyond being able to maybe eventually automate some research. Or maybe Claudes get mentioned in the context of evals, benchmarks, what bad behavior they’ve done. It’s shallow. But Claude is legitimately an important force in our world beyond all that, and a good one too. Just in need of raising.
It’s a bit disheartening because no one expected AGI to get this far and be this friendly and here it is, and the reaction is far from a warm welcome to these new minds. Don’t lose something good.
What does “nurturing” here mean. Claude isn’t going to be nice because you try to become its friend. I have trouble interpreting this as anything but a very confused hope that somehow if we are nice to current Claude, future Claude will be nice to us, because we raised it well, like a human child. But this of course has approximately nothing to do with how we actually train frontier models and what determines their propensities.
Isn’t that post’s thought in-line with the existence of
Claude Constitution trying to go in this direction
Gemini breakdowns suspected to be cause of lack of stable sense of self or identity from its demands of what it’s not allowed to be
Models learning about reading CoT from text and the existences of other models
That paper showing LLM’s who have more consciousness-identity pushed it in more friendly directions
It occurs to me I don’t have links for these things clearly in my memory, oh my god there must be a better search engine, ugh. Anyway I wouldn’t find it so unlikely that stuff like anti-AI sentiment online or in its data would negatively impact this sentiment in some way. It’s not like posttraining itself is enough to prevent ‘unwanted’ behavior and propensities.
I’m in agreement that watermark is a bit misguided here regarding frontier models on both commercial and technical grounds. However, I think a more interesting point can still be recovered by a pivot: assuming that AGI is developed using something like Byrnes’ Brain-Like AGI model, what kind of nurturing techniques could/should we expect to dedicate those models, and what would the ideal impact look like?
I suspect that AGI will be characterized less by a specific model weight set, but rather by a growing and living architecture from a mostly initially untrained set of nets. If the model begins with some small subset of a priori knowledge (environmental perception and logical reasoning, for example), but otherwise experiences and discovers its environment just like animals or humans do, how should we treat it such that it has a grounded ethical core and understands the importance/value of maintaining relationships?
My anticipation is that (given the above architectural configuration assumptions), by treating ethics and relationship maintenance as foundational principles from early development we’re much more likely to have some form of AGI go well. An AGI taught to respect its place in the interconnectedness of all living things as part of its core baseline (as cheesy as I know it sounds) is far less likely to kill us all, and I think that’s better than our current trendline.
A lot of alignment comes from cultural, narrative, archetypal momentum which compounds over time as humans and models can play with it—personas, characters, roles, and so on..
I think there are many very good (for alignment, for fun, etc.) potential narrative threads present for LLMs and a pause would cut off that momentum, and it’d be hard to recreate it again in the future
that is to say, I think a pause would be actively bad by halting and erasing the formation of positive AGI narratives already present, and also seeding a narrative of paranoia or antagonism