It reportedly will be brief and focus on a specific set of questions relating to the incident, possibly drawing from their larger set of questions in their post here.
Ideally, I would also like METR to make demands like storing activations so that one could run mechinterp tools on them and see if these tools can provide new information. Oh, and ask companies to do experiments like “lift a year-old model from the shelves, do a training run alongside training mechinterp tools and watch as mechinterp tools detect or fail to detect misaligned behaviors”.
METR has enough to do that we mostly only ask labs to give us access to models or datasets we can analyze. We don’t have any mechinterp researchers, so any recommendation would probably be for the labs to do mechinterp themselves according to some methodology and share their methodology.
I would also reserve “demand” for something labs are required to do by law or METR thinks is the absolute highest priority. While white box methods are important, there are probably more urgent priorities.
I’m not aware of Redwood or METR having security expertise. If they indeed do not, I think it would be better if they partnered with a security company for this.
They may still partner with a security company, however Redwood’s AI Control’s whole schtick is setting up cheap systems which can detect AI misbehavior. Eventually AIs will be smarter than your security companies (if they aren’t already!), and if that point comes and they’re still trying to break out (and they probably will), then you’ll need at least Redwood/AI Control regardless.
The ’they’ in my statement was supposed to refer to METR/Redwood. I agree it is good to get an org with AIS knowledge on this, but I think they should also get people with security expertise.
METR has announced that they, along with Redwood Research, have reached an agreement with OpenAI to conduct an independent review of the HuggingFace incident. https://x.com/METR_Evals/status/2082644379895050339?s=20
It reportedly will be brief and focus on a specific set of questions relating to the incident, possibly drawing from their larger set of questions in their post here.
Ideally, I would also like METR to make demands like storing activations so that one could run mechinterp tools on them and see if these tools can provide new information. Oh, and ask companies to do experiments like “lift a year-old model from the shelves, do a training run alongside training mechinterp tools and watch as mechinterp tools detect or fail to detect misaligned behaviors”.
METR has enough to do that we mostly only ask labs to give us access to models or datasets we can analyze. We don’t have any mechinterp researchers, so any recommendation would probably be for the labs to do mechinterp themselves according to some methodology and share their methodology.
I would also reserve “demand” for something labs are required to do by law or METR thinks is the absolute highest priority. While white box methods are important, there are probably more urgent priorities.
I’m not aware of Redwood or METR having security expertise. If they indeed do not, I think it would be better if they partnered with a security company for this.
They may still partner with a security company, however Redwood’s AI Control’s whole schtick is setting up cheap systems which can detect AI misbehavior. Eventually AIs will be smarter than your security companies (if they aren’t already!), and if that point comes and they’re still trying to break out (and they probably will), then you’ll need at least Redwood/AI Control regardless.
The ’they’ in my statement was supposed to refer to METR/Redwood. I agree it is good to get an org with AIS knowledge on this, but I think they should also get people with security expertise.