@_nwyin
nwyin.com
nwyin
We urgently need to explore the swarm scaling laws and swarm behavior in the open (I’d focus on coding and cybersecurity benchmarks first). I think we can learn a lot from constrained environments, using a fairly minimal harness and system prompt. I’d be focused on understanding the messaging patterns and how swarms solve problems before turning to finding hacks.
We’ve seen that, given enough time and tokens, models will eventually find and abuse exploits.
Currently, it seems like you need budgets in the range of $10-100ks to run swarms on tasks to gather sufficient data. This is out of budget for most safety researchers. Seems bad!
Only a few organizations or individuals will have the budget to further develop our understanding of swarm behavior. I don’t expect OpenAI or Anthropic to be fully transparent about their interrogations about swarm behavior based on what they’ve been publishing so far about these incidents.
[Linkpost] Looking into the Swarm’s Eye
I’m updating my beliefs on the hardware gap between the USA (Nvidia) and China. I’ve been arguing the belief that Huawei would be able to close the gap in H100 equivalents with the US in ~5 years. As in, by 2032-ish, China would be able to bring online fleets of data centers with enough accelerators to rival the US buildout.
EpochAI has a post exploring the hardware growth trajectories in more detail. It’s convinced me that the hardware gap will remain for longer than I originally thought; I’m updating my belief that there will be a gap in H100 equivalents for another 7-10 years.
Multi-Agent Coordination Lets Us Pour More Compute Into Post-Training
nwyin’s Shortform
Do you know what kind of mid-training or post-training done to this model? My understanding is that a lot of the style comes from other phases of training. It’s possible this instance of the model had did not undergo training to produce outputs that are more legible to humans.
We also observe that GPTs are much better at communicating with humans than Claudes are.
Could you expand more on why you think this is the case?
The other view is that this isn’t severe enough that it won’t generate enough noise or concern for teams to take larger action (e.g. intl slowdown). Faced with competitive external pressures, leaders will decide this is benign and manageable enough that it just requires a relatively small pause/adjustment to security posture.

Thank you! sorry I didn’t catch this. Editing the original post.