IIUC this whole research agenda assumes that you have a way to tell whether predicted short run behaviors lead to catastrophic long term consequences?
Does the research agenda generalize to when the AIs learn continuously (no longer have static weights in deployment)?
Seems clearly dual use in that if you can find greatly compressed explanations for how AIs accomplish cognitive feats, you can likely figure out how to make AIs that accomplish those feats much more efficiently.
I’d add: Does the agenda generalize to when multiple AI agents interact over long time horizons, as it seems to have been the case in OpenAI agents attacking HuggingFace?[1]
IIUC this whole research agenda assumes that you have a way to tell whether predicted short run behaviors lead to catastrophic long term consequences?
Does the research agenda generalize to when the AIs learn continuously (no longer have static weights in deployment)?
Seems clearly dual use in that if you can find greatly compressed explanations for how AIs accomplish cognitive feats, you can likely figure out how to make AIs that accomplish those feats much more efficiently.
I’d add: Does the agenda generalize to when multiple AI agents interact over long time horizons, as it seems to have been the case in OpenAI agents attacking HuggingFace?[1]
https://www.youtube.com/watch?v=87DyyMV0kCY