imo training multi-agent RL feels like it’s the most likely driver of this. That said, most multi-agent RL in LLM environments I’m aware of is non adversarial.
imo training multi-agent RL feels like it’s the most likely driver of this. That said, most multi-agent RL in LLM environments I’m aware of is non adversarial.