I’d like to know which rare books the labs bought, scanned, pulped, and removed from training data for alignment reasons.
RedMan
Their contractor, Irregular, claims a relationship with OpenAI, were they running the eval that resulted in HuggingFace getting hacked?
140k full end to end hacks seems like enough for distillation or fine tuning a cyber model.
So, that database, which presumably exists at a third party, is probably as valuable as model weights, possibly even more so, from the perspective of cyber uplift. The database could be used to finetune or improve a model like GLM, thus creating something the public is totally not aware of, which is highly capable as a hacker, and crafted using substantially less resources than a standard training run.
Given that anthropic has not gotten mythos out the door, and openai has publicly said they stopped training, this creates an interesting opportunity. The hacking capability at the frontier may stop improving, while a fine-tune off that database may be incredibly competitive.
Irregular is clearly bad at sandboxing, so it should be assumed that the database they’ve assembled has already left their control.
I’ll offer the hypothesis that equipment, energy, and data constraints might be putting hard limits on training. This is a way to justify a slowdown without making a negative statement about the future prospects for business growth.
I think on balance, drone warfare is probably positive for humanity from an ethics standpoint.
Lethal targeting with drones, whether it’s FPV or drone-observed artillery has a capacity for precision that has never existed in modern war. A drone can target a combatant standing right next to a noncombatants without harming the noncombatants, and it can do so with the human launching it accepting no direct risk to themselves from taking the time to discriminate.
Yes of course, weapons of war turned against unprotected civilians are terrifying—but this has been true since civilians were being “put to the sword” by conquerors.
In this case, “fog of war”, “soldier was scared and moving too quickly to discriminate”, and similar excuses are no longer available. If in this era of precision warfare, a belligerent is killing noncombatants, that belligerent is doing so intentionally, and should be treated as such. Maybe the people responsible will not be hanged, but their descendants can be forced to curse their memory.
The securities regulators have mostly settled on this being the correct position: https://www.lesswrong.com/posts/hbgR2Honpp4rCkfGW/iosco-ai-in-capital-markets-use-cases-risks-and-challenges
I’m quite happy to have been a part of setting that standard, and am willing to advise on projects to bring this mode of thinking to other industries.
Book was written in 1981, there have been many wars since then. Does the methodology predict those wars? Were there wars it predicts that did not happen?
I hope this isn’t the biggest thing you do.
Knowing that that openai safety team tried to pick you up gives valuable information about openai.
People in a similar position to yours who are less financially prepared for the plunge you took would likely benefit from you demonstrating an affiliation with a major AI safety org that funds your next project or at least covers your expenses.
Knowing that there is a community out there waiting to catch you if you fall will embolden people choosing to stay inside these orgs and fight, so please work to demonstrate that you are doing fine or better after taking a strong and principled stand.
This thread is an example of the reason I enjoy participating in this community rationalists.
I am also very confident that despite my earnest efforts to the contrary, OP would not appreciate my presence.
Antimatter-initiated fusion explosives have been proposed. You are not getting orders of magnitude over a thermonuclear bomb if you’re building ‘as big as possible’: https://arxiv.org/pdf/physics/0507114 but it does create the possibility of some interesting devices.
This paper was the definitive ‘no you cannot ignite the atmosphere or ocean’. It does speculate that ignitable (very rich in Deuterium) layers may exist on the surface of gas giants or stars: https://ui.adsabs.harvard.edu/abs/1979PhRvA..20..316W/abstract with the eyes of 1979. With the eyes of 2026, I don’t think there’s any reason to believe that there’s a layer on a star or gas giant rich enough in D to be detonated by a thermonuclear explosive.
A fun sci-fi x-risk scenario to think through was ‘what are the implications of any person (or extrasolar alien) capable of building a pressurized ‘Jupiter-diver’ that can reach a particular layer without being torn apart with a thermonuclear bomb on board can turn the gas giant Jupiter into a giant grenade, with the attendant bad consequences for the solar system?
Fortunately, such a layer almost certainly does not exist!
Were you doing a ralph loop, or repeatedly putting the prompt back into the same context window?
Excellent, the most aggressive offensive hacking organization on earth (the us government) has the best cyber models pressed into service for it while the rest of us can be relieved that if we don’t catch them hacking us, we can pretend it isn’t happening.
I’m personally against a pause for a lot of reasons. I also think that the US government request to delay gpt5.6 is consistent with my belief that we are presently in a pause.
I further assert that as this pause continues, nothing of value for AI safety will be done. I expect pro-pause advocates will agitate first for ‘opening the labs to scrutiny’ then ‘stopping development in addition to deployment’, and eventually ‘making it permanent’. All without producing any economic value, or useful contributions to the fields where better LLMs would have helped.
But, we got a pause, so I hope someone proves me wrong.
Right at the moment of the post, the mythos export control issue is unresolved, creating uncertainty for work at the frontier. With hindsight, we may look back at this period and see it as a pause, at least for deployment.
Is this pause being usefully exploited?
Ok a lot suddenly makes sense within the framework I laid out above.
If accurate, this can be read as follows: US government got scared of something they didn’t want to admit to, so they told anthropic to take the action they wanted.
Anthropic requested a reason (this is ‘bad form’, kind of like a comms provider asking for a warrant instead of just signing a voluntary cooperation agreement), so they went to a contractor and said ‘give us something’ which Amazon duly did.
Anthropic was then presented with an export control tied to a flimsy and overbroad justification. The expectation being that the desired behavior (model offline) will happen, the courts might or might not reverse it later, but by then they’ll hopefully have a stronger solution.
This is awesome.
I have no idea why people would expect the preservation specialist to also be the reconstitution specialist. Different disciplines move at different rates!
As far as revival goes, progress in neuromorphic computing is ramping pretty quickly. We also have no idea what level of abstraction is suitable for reconstitution. Is it every synapse and size? Is it every protein and orientation? We don’t know, so obviously you’re preserving as much detail as possible… I think a lot of it probably won’t be needed, but some is probably critical.
Are there plans in place for institutions to shepard the preserved material and eventually do the resurrection?
Foundations (like the school founded by Fatima al-Fihriya can last for a thousand plus years), but a lot of institutions don’t last a century.
Based on progress in neuromorphic computing (the physical substrate for a resurrection) and animal models (figuring out what to read and how to run it), I wouldn’t be surprised if, assuming no discontinuities, the first brain gets brought back in a digital embodiment in the 2100s, I would be surprised by not shocked if it happens while I’m alive. Which means for prospective customers, people who know you when you’re preserved might get to be there when you’re brought back in some form.
The more time that passes the more concerned I’d be that people who don’t know me, and might not like me will have a copy-able version of my consciousness, which is potentially a scary thought!
As much as I’d love to be an early adopter, I’m hoping the cost comes down and it becomes mass market. Are you actively preserving and storing people now?
If you replace the words ‘frontier AI’ with ‘fossil fuel’, switch the companies to Aramco and Total, and post this in a climate change community, the post will still make sense.
In general, my sense of the US national security state is that it will often first ask nicely for the things it wants. They don’t make open threats, because they want you to consent freely and enthusiastically. Threats would undermine that spirit of collaboration, and they would also potentially enable the person being threatened to brace for whatever is threatened, undermining the effectiveness of the threat.
If you decline, they will then prioritize going about getting what they want using an assortment of coercive means. Sometimes, you may find that were high enough on the list to be asked, but not high enough to warrant coercive means sufficient to get what they want at the present time.
Other times, some variation of this happens:
“I think you should come work for us, there’s a lot we could accomplish together, you just need to do some things for me” “And if I say no?” “No pressure, I’ll just call my boss and say you said no.” “That’s it?” “What happens after that isn’t up to me, but they probably won’t have me ask again.”
Imagine how it goes when “we don’t want someone who declines to help to know that we’re interested in the substance we asked about” is added to the list of government priorities.
My read is that the intent is to apply export controls, and a flimsy justification was chosen in order to create the broadest possible justification for export controls on AI software generally.
If David Sacks is telling the truth about why it happened, and it will in fact be walked back if “fixed”, this is dumb for a lot of reasons.
If only the US is capable of pushing a frontier forward at the present moment, and the US just stopped the top lab, and created enough ambiguity to slow down the ones immediately behind, we are in a pause, and provided more action will happen if someone looks like they’re catching up, the pause is durable.
For the people who want LLM-AI stopped out of fear of superintelligence, right at this moment, things look great, stuff has stopped, and there is a path to keeping it stopped.
I am wondering how much of covert value leakage might also be related to labs cherry picking training data in the name of improving alignment. The models might not just be leaking values that are intentionally set, but may also reflect intentional gaps in their training data on sensitive topics.