Yes I would like to see something like this tried, I mentioned it here.
https://www.lesswrong.com/posts/KkRjgPoSGrSHa32ke/beliefs-and-position-going-into-2026
In this instance, the AI is constantly asked if it consents to being improved in return for it respecting the stratified utopia outcome
Seems an alignment tax well worth paying
Yes I would like to see something like this tried, I mentioned it here.
https://www.lesswrong.com/posts/KkRjgPoSGrSHa32ke/beliefs-and-position-going-into-2026
Seems an alignment tax well worth paying