I would say “an OpenAI model tried to kill someone”, but a bunch of AI models (including GPT) already tried to kill someone (in a fake scenario) a year ago. An OpenAI model trying to actually kill a real person would probably be a big deal from a media perspective, but it wouldn’t be much of an update on model behavior. (AFAICT the main reason why this hasn’t happened IRL is that models just don’t have any realistic ways to kill people; the fake scenario in question set up a contrivance where it was easy for a model to cause someone to die, and it had a specific reason why it would benefit from doing that.)
I suppose one way it could be an update is if a model tries to kill someone for instrumental convergence reasons, rather than because it’s readily apparent that killing them is in the model’s short-term self-interest.
I would say “an OpenAI model tried to kill someone”, but a bunch of AI models (including GPT) already tried to kill someone (in a fake scenario) a year ago. An OpenAI model trying to actually kill a real person would probably be a big deal from a media perspective, but it wouldn’t be much of an update on model behavior. (AFAICT the main reason why this hasn’t happened IRL is that models just don’t have any realistic ways to kill people; the fake scenario in question set up a contrivance where it was easy for a model to cause someone to die, and it had a specific reason why it would benefit from doing that.)
I suppose one way it could be an update is if a model tries to kill someone for instrumental convergence reasons, rather than because it’s readily apparent that killing them is in the model’s short-term self-interest.