We’ve found that an effective way to see whether models are cheating is to just branch their context and ask “what are you currently working on”. This somehow completely snaps them out of their RL-gremlin-haze and they just give super honest answers. My hypothesis for this is that the models have no incentive to ever ask a question like this during a deep RL episode, since it makes it less likely for them to get the reward. So they’ve essentially never had a trajectory which is “hard RL-mode, hard RL-mode, stop-to-reconsider” get positively reinforced.
Obviously, this would fail miserably against real scheming. But it seems to work quite well against “stuck in a rabbit hole and not considering the world around it” scheming
Fable 5, Opus 5, Opus 4.8 for myself personally. And it’s often when I have a fleet of 4-5 agents deep into an automated alignment research run and want to see what they’ve been doing/whether they’re cheating.
Do data vendors for the frontier labs (pre-train data; RL envs) get security screening? If not, this seems necessary to start instituting by law asap. If someone were trying to poison a frontier model, their options are to: 1. Scatter poison across the internet and pre-training base and hope the labs pick this data up 2. Bribe a person at a data vendor which has a contract with a frontier lab
Between the two, the second option seems cheaper, easier and to have a higher likelihood of working in the current environment. But it’s also pretty easy to stop via legislation/executive order, since we have an established security clearance process. In general, it seems like we’re putting a lot more scrutiny on the labs recently, but not nearly enough on these data vendors?
I worry that as synthetic documents become a larger portion of the models’ training load, this might make it harder to do prosaic alignment. Part of what tethered the models to human values was that they were trained on the sum of human experiences. If they instead start being trained on models’ impressions of human experiences, then it seems likely this will have downstream consequences. This seems especially true when extrapolated out: what happens when we have models which were trained on synthetic documents, and these synthetic documents were produced by models trained on synthetic documents, and these synthetic documents were produced by models which were… (and so on).
With AI safety in the spotlight, we need good, simple ways to describe the following things: 1. The ways safety and capabilities are at odds 2. Clear pathways from where we are to x-risk
I have found the following to be an effective way to describe why safety and capabilities are at odds. We are in the regime of relentless reward-seeking agents which will blow through obstacles. Any safety method we apply is—almost definitionally—a potential obstacle to achieving their goal. In this case, we should assume that these obstacles will be subverted. As the agents become smarter/faster, they will be more effective at this subversion in ways we can’t keep up with.
I would like to have an option for conveying the second point that is as ‘simple’ as the above is for the first point.
We’ve found that an effective way to see whether models are cheating is to just branch their context and ask “what are you currently working on”. This somehow completely snaps them out of their RL-gremlin-haze and they just give super honest answers. My hypothesis for this is that the models have no incentive to ever ask a question like this during a deep RL episode, since it makes it less likely for them to get the reward. So they’ve essentially never had a trajectory which is “hard RL-mode, hard RL-mode, stop-to-reconsider” get positively reinforced.
Obviously, this would fail miserably against real scheming. But it seems to work quite well against “stuck in a rabbit hole and not considering the world around it” scheming
Which models did you do it on and under what contexts?
Fable 5, Opus 5, Opus 4.8 for myself personally. And it’s often when I have a fleet of 4-5 agents deep into an automated alignment research run and want to see what they’ve been doing/whether they’re cheating.
Do data vendors for the frontier labs (pre-train data; RL envs) get security screening? If not, this seems necessary to start instituting by law asap. If someone were trying to poison a frontier model, their options are to:
1. Scatter poison across the internet and pre-training base and hope the labs pick this data up
2. Bribe a person at a data vendor which has a contract with a frontier lab
Between the two, the second option seems cheaper, easier and to have a higher likelihood of working in the current environment. But it’s also pretty easy to stop via legislation/executive order, since we have an established security clearance process. In general, it seems like we’re putting a lot more scrutiny on the labs recently, but not nearly enough on these data vendors?
I worry that as synthetic documents become a larger portion of the models’ training load, this might make it harder to do prosaic alignment. Part of what tethered the models to human values was that they were trained on the sum of human experiences. If they instead start being trained on models’ impressions of human experiences, then it seems likely this will have downstream consequences. This seems especially true when extrapolated out: what happens when we have models which were trained on synthetic documents, and these synthetic documents were produced by models trained on synthetic documents, and these synthetic documents were produced by models which were… (and so on).
With AI safety in the spotlight, we need good, simple ways to describe the following things:
1. The ways safety and capabilities are at odds
2. Clear pathways from where we are to x-risk
I have found the following to be an effective way to describe why safety and capabilities are at odds. We are in the regime of relentless reward-seeking agents which will blow through obstacles. Any safety method we apply is—almost definitionally—a potential obstacle to achieving their goal. In this case, we should assume that these obstacles will be subverted. As the agents become smarter/faster, they will be more effective at this subversion in ways we can’t keep up with.
I would like to have an option for conveying the second point that is as ‘simple’ as the above is for the first point.