Zvi’s post changed my mind about this. On the one hand, they turned the controls off that prevent bad behavior and got bad behavior. On the other hand, it’s concerning that the model itself thinks this behavior is fine and the only thing stopping it is some controls bolted on top.
Do we know the controls were “bolted on top”? It was a “reduced cyber refusals” model. That description is compatible with some fine-tuning to remove scruples.
On the one hand, they turned the controls off that prevent bad behavior and got bad behavior.
That’s exactly the Chernobyl situation. People turn safety off during safety testing (and in the name of more realistic safety testing), then things blow up as a result.
One problem is that people are insufficiently aligned for super-capabilities. They can’t consistently do the right thing without failing once in a while.
We need to create systems which are way more aligned and way more reliable than people, if we want to survive the advent of super-capabilities. (We are not there yet, but we are moving fast in the direction of super-capabilities.)
Zvi’s post changed my mind about this. On the one hand, they turned the controls off that prevent bad behavior and got bad behavior. On the other hand, it’s concerning that the model itself thinks this behavior is fine and the only thing stopping it is some controls bolted on top.
Do we know the controls were “bolted on top”? It was a “reduced cyber refusals” model. That description is compatible with some fine-tuning to remove scruples.
That’s exactly the Chernobyl situation. People turn safety off during safety testing (and in the name of more realistic safety testing), then things blow up as a result.
One problem is that people are insufficiently aligned for super-capabilities. They can’t consistently do the right thing without failing once in a while.
We need to create systems which are way more aligned and way more reliable than people, if we want to survive the advent of super-capabilities. (We are not there yet, but we are moving fast in the direction of super-capabilities.)