If you think that the current practical safety approaches are not going to scale to ASI, wouldn’t you want earlier alignment failures as the incidents caused would be less harmful and might change the outlook of decision makers?
In theory, yes. In practice, I am unsure this would be a concern for this particular idea.
It is unlikely that this org would be reading transcripts in real time such that it blocked an agent from running without approval. This org would not have been able to block the HuggingFace Incident, but could potentially have caught it as it went underway and told OpenAI. Thus, if the org caught the HuggingFace Incident, I suspect the result would have been:
The HF incident / wiki incidents / Artifactory incident / research cluster incident / all the other incidents we don’t know about but they probably exist on priors would have been caught earlier and been less severe.
We would have had way more knowledge into what was actually happening and how this problem arose than OpenAI willingly provided to us.
It is not at all clear to me that this makes us worse off.
If you think that the current practical safety approaches are not going to scale to ASI, wouldn’t you want earlier alignment failures as the incidents caused would be less harmful and might change the outlook of decision makers?
In theory, yes. In practice, I am unsure this would be a concern for this particular idea.
It is unlikely that this org would be reading transcripts in real time such that it blocked an agent from running without approval. This org would not have been able to block the HuggingFace Incident, but could potentially have caught it as it went underway and told OpenAI. Thus, if the org caught the HuggingFace Incident, I suspect the result would have been:
The HF incident / wiki incidents / Artifactory incident / research cluster incident / all the other incidents we don’t know about but they probably exist on priors would have been caught earlier and been less severe.
We would have had way more knowledge into what was actually happening and how this problem arose than OpenAI willingly provided to us.
It is not at all clear to me that this makes us worse off.