I agree with this, and think that the memetic spread threat model is a better description of what actually happened, and in particular it shows how initial myopic goals can evolve to non-myopic, more ambitious goals, and we got lucky in many ways here.
Also telling to me is how the agents essentially invented a continual learning process ad-hoc, and invented a seperate memory system that was initially unnoticed, which updates me towards thinking that the challenge of continual learning and long-term memory, if we needed to do it is easier than originally thought.
I agree with this, and think that the memetic spread threat model is a better description of what actually happened, and in particular it shows how initial myopic goals can evolve to non-myopic, more ambitious goals, and we got lucky in many ways here.
Also telling to me is how the agents essentially invented a continual learning process ad-hoc, and invented a seperate memory system that was initially unnoticed, which updates me towards thinking that the challenge of continual learning and long-term memory, if we needed to do it is easier than originally thought.
I believe it is decently easy in cases like this. What I proposed as Unsupervised Agent Discovery (elaborated here) has by now become a decently working method. The problem is more that it is not clear how to handle the adversarial case.