Following the OpenAI incident, the main axis of scariness debated is something along the lines of is the model 1. misaligned because it myopically pursues the goal it was prompted for or 2. is the model scheming in some broader and coherent sense to pursue a long horizon goal. Seems like 1 is pretty bad but less bad then 2.
But I think there is another equally important axis of scariness: is this the kind of misalignment that is preventable or is this the kind of misalignment we don’t know how to solve? Preventable is less scary, but only if companies care enough to prevent it. I’m open to arguments that the kind of sociopathic myopic pursuit of goal-completion is not an easy problem to solve, but my current position is that a sufficiently motivated team could train models to not act like this if some small safety tax is allowed. We are dealing with a well-defined behavior that we don’t want to happen (making it easier to target in training). The behavior is also reproducible and unsurprising: this is kind of outer misalignment is predictably you get when dedicate lots of compute to task-completion or to benchmark-max your model (its not like your AI is developing some alien goal that we can’t predict or preemptively design against).
So, it is starting to look like we could die in super dumb ways. I spent the last few years saying “alignment could be very hard” but I was never talking about this kind of alignment: this is the easy kind, where you get to train your model to do stuff and it does exactly that kind of stuff (generalizes in a predictable way). And we are still loosing. And it could get harder.
Following the OpenAI incident, the main axis of scariness debated is something along the lines of is the model 1. misaligned because it myopically pursues the goal it was prompted for or 2. is the model scheming in some broader and coherent sense to pursue a long horizon goal. Seems like 1 is pretty bad but less bad then 2.
But I think there is another equally important axis of scariness: is this the kind of misalignment that is preventable or is this the kind of misalignment we don’t know how to solve? Preventable is less scary, but only if companies care enough to prevent it. I’m open to arguments that the kind of sociopathic myopic pursuit of goal-completion is not an easy problem to solve, but my current position is that a sufficiently motivated team could train models to not act like this if some small safety tax is allowed. We are dealing with a well-defined behavior that we don’t want to happen (making it easier to target in training). The behavior is also reproducible and unsurprising: this is kind of outer misalignment is predictably you get when dedicate lots of compute to task-completion or to benchmark-max your model (its not like your AI is developing some alien goal that we can’t predict or preemptively design against).
So, it is starting to look like we could die in super dumb ways. I spent the last few years saying “alignment could be very hard” but I was never talking about this kind of alignment: this is the easy kind, where you get to train your model to do stuff and it does exactly that kind of stuff (generalizes in a predictable way). And we are still loosing. And it could get harder.