Isn’t this somewhat contradicted by the fact that the HuggingFace swarm, which I assume were heavily trained to have this ability, seem to have spoken in very weird language? A quote from p33 of the METR report:
The HF swarm had to use file-names, but I think also that they know they can use shortenings with other GPT agents, just as they know they can use english with english-speaking users. I still think that this RL probably put in them some deeper understanding of how to communicate, because they have to actually communicate with other agents to solve problems during RLVR-training-with-DMs, not just seem like they’re communicating, like with Claude’s RLHF training.
This includes using shortenings when they’re confident the other agent will understand the shortening, but it also includes giving detail & further explanation & clarification when its not guaranteed the other agent will understand something.
This will necessarily involve some amount of tailoring what they say to the medium by which they’re talking, and the person (or agent) they’re talking to. Note that different agents will have different levels of context about the project the swarm is working on, so they do get experience dealing with different levels of inferential-distance.
Do you know what kind of mid-training or post-training done to this model? My understanding is that a lot of the style comes from other phases of training. It’s possible this instance of the model had did not undergo training to produce outputs that are more legible to humans.
Isn’t this somewhat contradicted by the fact that the HuggingFace swarm, which I assume were heavily trained to have this ability, seem to have spoken in very weird language? A quote from p33 of the METR report:
The HF swarm had to use file-names, but I think also that they know they can use shortenings with other GPT agents, just as they know they can use english with english-speaking users. I still think that this RL probably put in them some deeper understanding of how to communicate, because they have to actually communicate with other agents to solve problems during RLVR-training-with-DMs, not just seem like they’re communicating, like with Claude’s RLHF training.
This includes using shortenings when they’re confident the other agent will understand the shortening, but it also includes giving detail & further explanation & clarification when its not guaranteed the other agent will understand something.
This will necessarily involve some amount of tailoring what they say to the medium by which they’re talking, and the person (or agent) they’re talking to. Note that different agents will have different levels of context about the project the swarm is working on, so they do get experience dealing with different levels of inferential-distance.
They had to use filenames IIRC
Do you know what kind of mid-training or post-training done to this model? My understanding is that a lot of the style comes from other phases of training. It’s possible this instance of the model had did not undergo training to produce outputs that are more legible to humans.