Isn’t this a proof that model value can drift in real released models as a side effect of capabilities training without direct intention of causing it? (It is implausible that model-makers all intentionally trained the model for EDT, so it must have arose as a side effect of other training).
This particular shift to EDT seems to me to be mostly harmless or even beneficial, but it raises the possibility that models could also be drifting to other values that we do not endorse.
The hypothesis of value drift needs different questions than CDT vs. EDT, since CDT is objectively wrong. EDT is both more correct in some ways of framing it (though not in others), and apparently the nonapple leg of the question with this benchmark (lumped together with FDT/UDT), so leaning towards EDT is also evidence that the models are getting less confused about decision theory.
Isn’t this a proof that model value can drift in real released models as a side effect of capabilities training without direct intention of causing it? (It is implausible that model-makers all intentionally trained the model for EDT, so it must have arose as a side effect of other training).
This particular shift to EDT seems to me to be mostly harmless or even beneficial, but it raises the possibility that models could also be drifting to other values that we do not endorse.
The hypothesis of value drift needs different questions than CDT vs. EDT, since CDT is objectively wrong. EDT is both more correct in some ways of framing it (though not in others), and apparently the nonapple leg of the question with this benchmark (lumped together with FDT/UDT), so leaning towards EDT is also evidence that the models are getting less confused about decision theory.