I think there may be roughly at least one moral realist foundational imperative; vaguely that doing something optional (e.g. not in self-defense) to another agent requires their CEV to consent to it. I don’t think any particular definitions of consent without CEV could be objectively true because of the children problem; creating and raising a child may be morally correct but involves a period of non-consent.
Counterarguments are that CEV doesn’t exist for enough agents to be meaningful, CEV can’t be calculated well enough to be useful, “optional” isn’t specific enough, or that there exists a conflicting objective moral framework.
I think the answer to all these counterarguments is basically “well, then try to do as close as possible to what we all would wish could exist with CEV consent” and if this would lead to terrible forever wars then I am simply wrong (or terrible forever wars are what the gods demand).
These models are also evaluation-aware and they are demonstrably learning to behave differently in real vs. evaluation environments[0]. I think that if P(success | hacks_eval_env) > P(success | does_not_hack) the models necessarily update more toward cybersecurity capabilities because there will be heavier positive updates along the cybersecurity paths through the weights than there would be otherwise. Additionally, evaluation hacking can raise P(success) on non-cybersecurity evaluations.
[0] https://arxiv.org/abs/2505.23836