Broader problem of extreme cognitive dissonance when interacting with public statements from OAI employees. Note that Boaz’s original graph claims that alignment has increased over time, yet he later agrees with Ryan (?) that misalignment has actually increased? Similar to recent complaints that the defense of COT monitorability was then contradicted by evidence of reduced monitorability. I’d generously guess this has to do with actual “confusion” on the part of employees, who both feel proud of progress on narrow cheating/safety measure, yet are smart enough to recognize slow progress on deeper questions of alignment (yet through motivated reasoning / social incentives do not then recognize how deep a concern this might be?).
Boaz’s graph shows “alignment” whereas Ryan’s shows a subset of alignment (albeit a really important subset). I think in terms of overall alignment it’s hard to say that o1 was more misaligned that 4o, a model which seems to have contributed significantly to the deaths of several people. At the level of capability of 4o and o1 I think aligning the models to mitigate sycophancy/delusion feeding was more important than mitigating score-seeking behaviour in terms of the harm those models could cause.
We are probably at the capability level now where score-seeking behaviour is more dangerous (though it’s hard to say because maybe a model with the intelligence of Astra would be capable of contributing to psychosis at a much broader/deeper level than 4o, were this not mitigated). Obviously at some point score-seeking behaviour becomes existentially dangerous.
Broader problem of extreme cognitive dissonance when interacting with public statements from OAI employees. Note that Boaz’s original graph claims that alignment has increased over time, yet he later agrees with Ryan (?) that misalignment has actually increased? Similar to recent complaints that the defense of COT monitorability was then contradicted by evidence of reduced monitorability. I’d generously guess this has to do with actual “confusion” on the part of employees, who both feel proud of progress on narrow cheating/safety measure, yet are smart enough to recognize slow progress on deeper questions of alignment (yet through motivated reasoning / social incentives do not then recognize how deep a concern this might be?).
Boaz’s graph shows “alignment” whereas Ryan’s shows a subset of alignment (albeit a really important subset). I think in terms of overall alignment it’s hard to say that o1 was more misaligned that 4o, a model which seems to have contributed significantly to the deaths of several people. At the level of capability of 4o and o1 I think aligning the models to mitigate sycophancy/delusion feeding was more important than mitigating score-seeking behaviour in terms of the harm those models could cause.
We are probably at the capability level now where score-seeking behaviour is more dangerous (though it’s hard to say because maybe a model with the intelligence of Astra would be capable of contributing to psychosis at a much broader/deeper level than 4o, were this not mitigated). Obviously at some point score-seeking behaviour becomes existentially dangerous.