How do we carve out “the system” from “the environment”, i.e. how do we draw a Cartesian boundary, in order to roughly match human instincts like “looks like it’s robustly optimizing for X”? That’s an open question, and probably a special case of the more general question of how humans abstract out subsystems from their environment. (This was actually relevant in the previous section too, but it’s more apparent once agency is introduced.)
On this problem I think Fernando Rosas made promising progress with The separation principle: where do beliefs and desires come from?. I recommend checking if this makes true progress
Are there any Chinese speakers here that can tell what the actual vibes in the Chinese government are about this? I suspect lesswrong would benefit a lot from an introduction to Chinese culture on AI in late 2026