I think the business PoV is a serious issue here that would need addressing, but the technical PoV problem proves too much. Yes, you are filtering for models that fuck up in less visible ways. But that is true for any form of monitoring. I think the correct lesson from this is “Monitoring is insufficient to save us”, but that does not necessarily lead to “We should not do monitoring.” I don’t see how this is any worse, for this problem, than CoT monitoring by the labs themselves (and is actually better, since it’s harder to get a lot of bits of optimisation pressure from this setup that is hopefully flagging things very rarely!)
If you think it’s worth doing CoT monitoring at all, the technical issue here shouldn’t dissuade you. If you do think we shouldn’t be monitoring at all because the good of monitoring is outweighed by the bad of filtering for less visibility, I disagree. While filtering for models that fuck up less visibly is bad, it’s bad because monitoring can’t catch it any more. Whereas if you don’t do monitoring at all, your visible-to-monitors fuckups may as well be invisible already.
In theory, yes. In practice, I am unsure this would be a concern for this particular idea.
It is unlikely that this org would be reading transcripts in real time such that it blocked an agent from running without approval. This org would not have been able to block the HuggingFace Incident, but could potentially have caught it as it went underway and told OpenAI. Thus, if the org caught the HuggingFace Incident, I suspect the result would have been:
The HF incident / wiki incidents / Artifactory incident / research cluster incident / all the other incidents we don’t know about but they probably exist on priors would have been caught earlier and been less severe.
We would have had way more knowledge into what was actually happening and how this problem arose than OpenAI willingly provided to us.
It is not at all clear to me that this makes us worse off.