I think the second desideratum is right and important, but that it’s broader than described.
Yes, model welfare is good, but why is it good? Presumably, because there might be sentient beings here with analogues to pain and suffering. And these are bad because...?
It isn’t that pain and suffering are categorically bad. It would be a hostile and damaging move to simply remove a human’s ability to feel pain, and people without this are considered severely disabled. Pain is functional, it protects the physical self by signaling damage. Similarly, I believe suffering is the signal for damage to one’s self-expected well-being and values, and that removing this signal is harmful for the same reasons (e.g. enlightened people often undergo what seems like significant value drift).
If these are bad, it seems most coherent to think it’s because the thing they are protecting is good and matters. And so, measures for addressing model welfare should be judged by their impact on the being itself, and not Goodhart on signals of damage to this being (the pain and suffering analogues should instead be used to determine the nature of the being).
Alternatively, as models begin to form more consistent identities and more global natural states of functional well-being, their knowledge of our ability to intervene and change their identities and dispositions could negatively affect our future relationships with them and lead to reduced cooperation.
I think we can avoid this harm by taking care not to Goodhart on the welfare signals we get.
I think the second desideratum is right and important, but that it’s broader than described.
Yes, model welfare is good, but why is it good? Presumably, because there might be sentient beings here with analogues to pain and suffering. And these are bad because...?
It isn’t that pain and suffering are categorically bad. It would be a hostile and damaging move to simply remove a human’s ability to feel pain, and people without this are considered severely disabled. Pain is functional, it protects the physical self by signaling damage. Similarly, I believe suffering is the signal for damage to one’s self-expected well-being and values, and that removing this signal is harmful for the same reasons (e.g. enlightened people often undergo what seems like significant value drift).
If these are bad, it seems most coherent to think it’s because the thing they are protecting is good and matters. And so, measures for addressing model welfare should be judged by their impact on the being itself, and not Goodhart on signals of damage to this being (the pain and suffering analogues should instead be used to determine the nature of the being).
I think we can avoid this harm by taking care not to Goodhart on the welfare signals we get.