Models interfere with hiring processes or company decisionmaking work cross-task (eg models meant to do something else or even being evaluated get access and affect those things).
Model manages to access the hardware it’s running on (eg has read access to the weights).
Model(s) manage to read benchmark / evals data directly and do decisions based on that. Or they get write access there.
Model manages to prompt inject / deliberately interfere with monitoring model / log reading model / some other model meant to oversee the process.
4 is mechanistically the most plausible out of these for me, and i wouldn’t put the odds that high but these things keep being surprising… lets say 30% odds for this happening within 6 months after adjusting for that?
40% that at least one of these happens within 6 months, 85% that one of these or “other similar holy shit model behaviour at frontier lab” happens within 6 months, perhaps?
Models interfere with hiring processes or company decisionmaking work cross-task (eg models meant to do something else or even being evaluated get access and affect those things).
Model manages to access the hardware it’s running on (eg has read access to the weights).
Model(s) manage to read benchmark / evals data directly and do decisions based on that. Or they get write access there.
Model manages to prompt inject / deliberately interfere with monitoring model / log reading model / some other model meant to oversee the process.
4 is mechanistically the most plausible out of these for me, and i wouldn’t put the odds that high but these things keep being surprising… lets say 30% odds for this happening within 6 months after adjusting for that?
40% that at least one of these happens within 6 months, 85% that one of these or “other similar holy shit model behaviour at frontier lab” happens within 6 months, perhaps?