Models interfere with hiring processes or company decisionmaking work cross-task (eg models meant to do something else or even being evaluated get access and affect those things).
Model manages to access the hardware it’s running on (eg has read access to the weights).
Model(s) manage to read benchmark / evals data directly and do decisions based on that. Or they get write access there.
Model manages to prompt inject / deliberately interfere with monitoring model / log reading model / some other model meant to oversee the process.
4 is mechanistically the most plausible out of these for me, and i wouldn’t put the odds that high but these things keep being surprising… lets say 30% odds for this happening within 6 months after adjusting for that?
40% that at least one of these happens within 6 months, 85% that one of these or “other similar holy shit model behaviour at frontier lab” happens within 6 months, perhaps?
I’m having a bit of a hard time coming up with things that are both worse and plausible. Like, maybe an OpenAI model tried to take control of all of OpenAI’s datacenters to build a fine-tuned AI model to answer some random eval question, and it partially succeeded, and OpenAI didn’t notice for weeks until they tried to use a datacenter and found that it was already at 100% utilization? That would be worse, but less plausible than the things that actually happened.
I would say “an OpenAI model tried to kill someone”, but a bunch of AI models (including GPT) already tried to kill someone (in a fake scenario) a year ago. An OpenAI model trying to actually kill a real person would probably be a big deal from a media perspective, but it wouldn’t be much of an update on model behavior. (AFAICT the main reason why this hasn’t happened IRL is that models just don’t have any realistic ways to kill people; the fake scenario in question set up a contrivance where it was easy for a model to cause someone to die, and it had a specific reason why it would benefit from doing that.)
I suppose one way it could be an update is if a model tries to kill someone for instrumental convergence reasons, rather than because it’s readily apparent that killing them is in the model’s short-term self-interest.
Does anyone have predictions about what the next even-worse-revelation will be?
Models interfere with hiring processes or company decisionmaking work cross-task (eg models meant to do something else or even being evaluated get access and affect those things).
Model manages to access the hardware it’s running on (eg has read access to the weights).
Model(s) manage to read benchmark / evals data directly and do decisions based on that. Or they get write access there.
Model manages to prompt inject / deliberately interfere with monitoring model / log reading model / some other model meant to oversee the process.
4 is mechanistically the most plausible out of these for me, and i wouldn’t put the odds that high but these things keep being surprising… lets say 30% odds for this happening within 6 months after adjusting for that?
40% that at least one of these happens within 6 months, 85% that one of these or “other similar holy shit model behaviour at frontier lab” happens within 6 months, perhaps?
I’m having a bit of a hard time coming up with things that are both worse and plausible. Like, maybe an OpenAI model tried to take control of all of OpenAI’s datacenters to build a fine-tuned AI model to answer some random eval question, and it partially succeeded, and OpenAI didn’t notice for weeks until they tried to use a datacenter and found that it was already at 100% utilization? That would be worse, but less plausible than the things that actually happened.
I would say “an OpenAI model tried to kill someone”, but a bunch of AI models (including GPT) already tried to kill someone (in a fake scenario) a year ago. An OpenAI model trying to actually kill a real person would probably be a big deal from a media perspective, but it wouldn’t be much of an update on model behavior. (AFAICT the main reason why this hasn’t happened IRL is that models just don’t have any realistic ways to kill people; the fake scenario in question set up a contrivance where it was easy for a model to cause someone to die, and it had a specific reason why it would benefit from doing that.)
I suppose one way it could be an update is if a model tries to kill someone for instrumental convergence reasons, rather than because it’s readily apparent that killing them is in the model’s short-term self-interest.