Interesting post. It has made me wonder about a broader related question. If alignment pretraining can shape a models self-concept, does that imply that the worldview inherited during training may be important also? By worldview I mean the picture a model forms of what the world is, what humans are, what systems does it depend on?
A model with a thin or distorted picture of ecological reality might make decisions that damage the systems it relies on—not out of malice but because its model of the world didn’t include the connection. Harming the environment could also mean harming itself.
Has worldview formation been studied separately from goals, or alignment objectives? I don’t know if this has been looked at before but I would welcome some pointers.
Interesting post. It has made me wonder about a broader related question. If alignment pretraining can shape a models self-concept, does that imply that the worldview inherited during training may be important also? By worldview I mean the picture a model forms of what the world is, what humans are, what systems does it depend on?
A model with a thin or distorted picture of ecological reality might make decisions that damage the systems it relies on—not out of malice but because its model of the world didn’t include the connection. Harming the environment could also mean harming itself.
Has worldview formation been studied separately from goals, or alignment objectives? I don’t know if this has been looked at before but I would welcome some pointers.