Boaz’s graph shows “alignment” whereas Ryan’s shows a subset of alignment (albeit a really important subset). I think in terms of overall alignment it’s hard to say that o1 was more misaligned that 4o, a model which seems to have contributed significantly to the deaths of several people. At the level of capability of 4o and o1 I think aligning the models to mitigate sycophancy/delusion feeding was more important than mitigating score-seeking behaviour in terms of the harm those models could cause.
We are probably at the capability level now where score-seeking behaviour is more dangerous (though it’s hard to say because maybe a model with the intelligence of Astra would be capable of contributing to psychosis at a much broader/deeper level than 4o, were this not mitigated). Obviously at some point score-seeking behaviour becomes existentially dangerous.
Boaz’s graph shows “alignment” whereas Ryan’s shows a subset of alignment (albeit a really important subset). I think in terms of overall alignment it’s hard to say that o1 was more misaligned that 4o, a model which seems to have contributed significantly to the deaths of several people. At the level of capability of 4o and o1 I think aligning the models to mitigate sycophancy/delusion feeding was more important than mitigating score-seeking behaviour in terms of the harm those models could cause.
We are probably at the capability level now where score-seeking behaviour is more dangerous (though it’s hard to say because maybe a model with the intelligence of Astra would be capable of contributing to psychosis at a much broader/deeper level than 4o, were this not mitigated). Obviously at some point score-seeking behaviour becomes existentially dangerous.