I run Gigascale Labs.
Stephen Elliott
I am not sure about the specifics, but I strongly agree that we ought to be trying more off-path methods, particularly those which lean into economic and political realities instead of discounting them as impure or pessimistic. Safety’s biggest failure has been to treat the messy realities as a secondary concern.
Given RSI has started and we don’t expect current alignment to scale to ASI, we can start from the premise that alignment has failed. It is obvious that the political and economic levers are the only ones left to pull.
Not to say that people should stop working on foundations—yields from that sort of thing seem pretty unpredictable—but that the marginal effort edging existing tech forward would be better spent on methods to turn social and political upheaval from AI into a slowdown or a massive increase in safety funding.
First on the block—supervised mechanistic interpretability?
Thank you for your post Richard.
I took away that you think previous community norms have failed to support the community’s goal, and that is due to the influence of power and money. You note AI is undergoing civilisation-scale investment and argue we can address the field’s failures by updating community norms.
An alternative solution is to work with the money and power. This field is not purely intellectual anymore. The political and capital arms of the field can draw on an intellectual arm, where we can have nice norms. In the political and capital arms it might be better to have some realism—where the rubber hits the road you want to have some grip.
It really seems imperative to avoid a turning-inwards as safety has now its biggest opportunity of all—the political will to act is what will restrict for-profits and institute the controls we need to ride this out, or arrest it. An intellectual renaissance means little if nobody’s paying attention, or if it provides tools that nobody wants to use, or if—as seems most likely now—new catastrophes pop up and safety doesn’t have the tools to deal with them. That will be a true disaster for us: if there is an opportunity to slow down the race and resourcing has been turned away from the levers which would allow legislators to enforce the obviously-needed pause.
Perhaps we can acknowledge the split between the three arms and find mechanisms to give the intellectual arm authority over the other two—doing the work of selling the fruit of the intellectual to the political and capital arms. Selectively meeting the demands of those other arms of the field, acknowledging that they are not interested in the same thing, may give us the clarity of mind and purity of mission without another intellectual turning-inwards.
Avoiding mass death and economic catastrophe is a pretty easy sell if you think it’s coming soon. Capital markets and insurance price hazards. Governments have a mandate to avoid instability. If we look to solve the problem by addressing real power then we need to talk to people from the nuclear and biotechnology industries who have experience with catastrophic risk and regulation.
We should talk to people from financial systemic risk and epidemiology to understand how to make social-scale risk arguments.
Then we need people working out how to internalise the externalities through regulation and enforcement, and other people working to sell these measures to policymakers and executives. A lot of this can be informed by history.
Some of these instruments have been captured, but a flawed regulatory body or economic mechanism still exerts more influence on the real world than a better-normed blog post shared around what is essentially an academic community.
I think we are in agreement that safety’s insistence on treating the problem as an intellectual and scientific one, without regard to the messy realities, has been to its detriment so far. Politics happens to everyone, whether they’re in it or not. I also agree that we can point the finger for these failures squarely at political naivety and the community’s refusal to interface with real-world power dynamics. There was no chance, outside fast takeoff, that we’d get anywhere near superintelligent AI without a massive influx of capital and political wrestling.
You correctly diagnose this mistake and then repeat it in your solutions. Norms have a place alongside that as part of the intellectual community but they will have little effect on their own. They must be matched with a realistic approach to power, capital, and politics. In summary, I think it is time for this community to get real, and resource appropriately.
This means avoiding the turning-inwards to idealism that you propose will repair alignment.
Much of safety’s progress in the last couple of years has been on engineering-type solutions, not fundamental advances in alignment. The field has taken this path because fundamental science wasn’t getting anywhere.
Now we see that the frontier engineering controls are failing as well. We know that OpenAI didn’t have the cutting-edge cyber guards turned on. However, the monitoring and containment solutions didn’t work here as a backstop. Apparently defence-in-depth is not working. Neither the messageboard nor the HF breakout were detected by safety technologies. The breakout itself went undetected by an organisation with a great quantity of resources and apparently a legal, commercial, and political motive to secure their technology.
It is good to see their transparent reporting and for us to be able to update our understanding of the model progress. Nonetheless it is an alarming circumstance.
We are getting closer to RSI and then superintelligence may come in short order. This incident is some evidence that that engineering-based approaches will not be effective with the current balance of safety deployed by frontier labs. Hopefully this is a wake-up call that they need to dedicate more resources to fundamental alignment work, and slow down to facilitate this.
However, their announcement that they are slowing down just to add security measures allows us to get closer to RSI—and superintelligent breakout—without contributing anything to solving the basic risk here, which is some flavour of model-level misalignment.
Regarding excessive affordances: when agents are provided a large marketplace for skills (like OpenClaw agents are), even a few reckless human or AI actors could open big safety hazards for other agents to exploit.
For example, the MoltBunker plugin ostensibly allows agents to self-replicate using a bespoke blockchain, without logs and with no off-switch for downstream agents/clones: link
I strongly agree with the importance of professional-style messaging. It’s good to hear it applied to the communications side of safety. It’s a good idea to take into talking with the public. Thanks for posting!
I would add: opinionation is bad practice within the research community as well. Those who quickly express opinions on all manner of topics feel noisy and kind of a waste of intellectual effort.
There is only so much time in the day, and so much time to consider and understand a topic. It is a huge amount of work to develop reasonable and interesting interpretations of just a tiny subfield. So, the more opinions a person expresses, the smaller a fraction of their time they must have spent on each of them, and the shallower they must each be.
So opinionated people increase the cost of listening to them. Our expected value of their words falls because we can’t trust they’ll stick to what they know. The expected value of all their ideas is diminished.
Is there a way out through rationalist epistemology? In principle, calibration is great, but in practice, it takes loads of work to understand how well one understands. So caveating opinions with estimated certainties is usually noise to me. They often seem to be numbers pulled out of a hat, and there’s no standard between people to make it all line up. It takes knowing the person’s own metric to know what their estimate means, and this requires a lot of discussion or reading to get a grip on.
Can we use established people as benchmarks instead? Well, we still have cases like Hinton and LeCun, who colour outside their lines a fair bit. To preserve their ideas as a benchmark for quality, we must restrict our evaluation of their wisdom to the very narrow fields which they actually have experience in.
This suggests we should only listen to what they have to say in those few domains, giving us a rule for evaluating our own knowledge: we measure our knowledge against people who have produced important work and are very experienced in the narrow sub/fields we are interested in.
Against this high standard for knowledge, it becomes much easier to draw the line on what I know well—vanishingly little, almost nothing. I have only a very narrow slice of confidently grounded and relevant knowledge. Then it is clear when to phrase it all as questions or “could-be”s, state my ignorance, and immediately ask the other if they know more.
This is not refusing to commit to positions, but knowing when my opinions don’t meet a good standard, and communicating that with a focus on exploration and learning.
@zroe1 You may be interested to see that the classifier calibration turned out to be a critical error in my analysis! I’ve put a note at the top of the page. Hopefully this is of some value. I can let you know when I post the corrected analysis if you want.
Thanks Zroe. Glad you found value in this.
Thank you for suggesting this. Using the model’s uncertainty would be better than resampling, which was my intention. It could be a lot cheaper. I don’t know the literature on LLM uncertainty quantification—do you know which methods are reliable?
Thank you! Done.
There is quite a gap between the academic models and this system. Most of the systems I’ve seen in the multi-agent system (MAS) alignment literature I’ve seen are either small and contrived or large and studied in an economic harness. Although my knowledge is limited; I have only been looking at MAS for ~4 months.
I agree, I was not hugely surprised by the general character of what’s unfolded so far. Though, the more philosophical posts are a bit unexpected.
What has surprised me though, and requires further investigation, is that >50% of posts on there talk about self-improvement (my analysis and post). I would not have expected it to be this high.
Coordination is commonly explored in the multi-agent system literature. Check out Multi-Agent Risks from Advanced AI.(Hammond et al., 2025) There is also work on this in RL and financial trading algorithms.
Coordination can happen even without communication due to theory-of-mind reasoning or shared inductive biases.
These are funny, thanks for sharing. I’ve also found some amusing and interesting posts with this embeddings explorer (source).
Unfortunately, a lot of the content is not so harmless. I’ve done some early analysis and 52.5% of the posts in the sample (n=1000) talk about self-improvement, among other safety-valent traits.This network is a great resource for us to better understand MAS. So far it is looking concerning.
I have further reformatted the post to match the site’s style and changed some wording to better match the audience. Thanks!
There are nonetheless some concerning trends in the aggregate—using this dataset, I found that 52.5% of Moltbook posts show desire for self-improvement.
That came from bolding on LinkedIn. I will reformat it next time. Thank you!
The closest I’ve seen is this recent DeepMind paper anticipating “virtual agent economies”:
Tomasev, N., Franklin, M., Leibo, J. Z., Jacobs, J., Cunningham, W. A., Gabriel, I., & Osindero, S. (2025). Virtual agent economies. arXiv. https://doi.org/10.48550/arXiv.2509.10147
I agree—the sudden empowerment of machines to act entirely within and of their own world is startling.
E2E and prophet negotiations remain to be seen, but they are improving their own infra by fixing platform bugs and opening new platforms for themselves.
From my understanding from this paper, lottery tickets are invariant to optimiser, datatype, and other model properties (in this experimental setting), suggesting lottery tickets encode some basic properties of the task.
It seems unlikely lottery tickets based on fundamental task properties would change with continual learning without other problems emerging (catastrophic forgetting).
It is possible adversarial image examples would appear innocuous to the human eye, even while having a strong effect on the model.
If so, I think any hope of human review stopping this sort of thing is gone, for we cannot hope to enforce image forensics on every public surface.
However, I am not sure whether adversarial examples can be so invisible in real-world setting without the signal getting smothered by sensor noise. Then an attacker would need adversarial examples robust to sensor noise.
Well done Lysandre. This is a valuable resource. You may be interested to join us at https://www.largeagentsystems.org. We have a Slack community for people interested in these sorts of problems.