Open-Weights Mythos Capabilities Are Coming. We’re Not Ready.

Long story short: in my assessment, there is an 85% chance we will end up, in the next 24 months, with an open-weights model, or system thereof, capable of “The Juice” that models such as Mythos have, with respect to cybersecurity at the very least. This post goes into why that will likely happen, what the implications are, and how we, as a society and as individuals, can respond to it if/​when it does.

First off: why do I say it’s so likely? Like, couldn’t China just...ban open-weights models, and then that solves the problem? Not so fast. Yes, China currently dominates the open-weights frontier. However, there are also open-weights AI labs in plenty of other countries (US, France, the UAE, South Korea, and Canada come to mind, and I’m sure there are others). Yes, some of these are substantially behind the frontier, but each of the aforementioned countries has an open-weights model no more than 24 months behind the current frontier (that’s why I said 24 months earlier); remember, in Aug 2024, 24 months ago as of when I am writing this, the strongest models were Sonnet 3.5 and GPT 4o.

It seems highly unlikely that all of these different countries, with their disparate political systems, geopolitical ties, regulatory environments, and more, will all independently adopt the exact same policy of fully banning all open-weights model releases, at least not until it’s too late and an open-weights model capable of Mythos-level capabilities has already been released. And, given the current geopolitical situation, it seems highly unlikely that there will be any international agreement to this effect (at least not until a disaster has already occurred and it’s already too late).

But let’s say we do manage to successfully conduct an international ban on open-weights Mythos-level models. At that point, you would also need to assume that each and every single closed-weights lab with Mythos-level capabilities also has sufficiently competent cybersecurity to be able to reliably not have their model weights exfiltrated. This seems highly unlikely across so many different labs, and an affirmative duty of having sufficient cybersecurity seems like an even taller order than a ban on open-weights training, which, as discussed before, is, in and of itself, highly unlikely to be possible on an international basis.

Mythos-level capabilities are likely possible with today’s open-weights frontier

(epistemic status: I’m not as confident in this part as I am in the rest of this post; I genuinely want to hear others’ thoughts here)

This next part is a bit more controversial, but I’d argue that the current frontier of open-weights models (probably Kimi K3), as it is now, with sufficient fine-tuning, inference-time compute, and harness improvements, could be used to get to a Mythos level in cyber capabilities (or capabilities in any other verifiable domain). Most obviously, a model could be fine-tuned to enhance dangerous capabilities in an area like cyber; this is well-established in the research and has already been done to some extent; see models like Hackphyr for example.

Additionally, inference-time compute has the potential to substantially enhance the capabilities of Kimi K3. For example, there are approaches like best-of-N, multi-agent debate, tree-of-thought, etc. It seems that a lot of why weaker models lack “The Juice” is because they, at a certain point, get off-track in a way that prevents them from recovering, and approaches like best-of-N could sharply reduce those problematic points.

You could even achieve a similar result without anything fancy: groups like AISLE Security have shown that open-weights models can replicate much of Mythos’s analysis if pointed at the correct file. You could use Dynamic Workflows or something similar and just point subagents at all of the applicable files (keep in mind you realistically wouldn’t have to do a subagent for every single file in a large codebase, as it’s only a subset of those files that are realistically likely to be exploitable).

Likewise, I expect substantial improvements to be possible just from better harnesses; for example, better memory systems, better multi-agent systems, etc.

Because of those factors, I think it’s quite likely that, even with current open-weights models as they are today, we will end up in a situation where attackers have access to Mythos-level cyber capabilities. Keep in mind that cybercrime is extraordinarily lucrative (breaches like Bybit have yielded attackers as much as $1.5B), so attackers would gladly be willing to pay even truly exorbitant token expenditures in order to conduct an attack; cost is not a meaningful bottleneck here, and with Cerebras/​Taalas/​etc., token rate almost certainly won’t be a bottleneck either. The one saving grace, as discussed earlier, is that this is only applicable to verifiable tasks like cyber; it is less applicable for less-verifiable tasks like bio, where harnesses/​inference-time compute/​etc. play less of a role and it’s more of a “either you know it or you don’t” sort of situation. But, for the reasons discussed below, open-weights Mythos capabilities, even just for cyber, are likely to be extremely harmful.

This will be very bad

The fact is that many of our most important organizations (both in terms of how severe the impact would be if they were compromised and in terms of their attractiveness to hackers) are fundamentally not ready for a world where anyone can conduct a cyberattack. I actually expect FAANGs/​other big tech companies to be able to prepare reasonably well from a defensive standpoint using some combination of closed-weight frontier models (which, at that point, will be far stronger than Mythos 5) along with just normal cybersecurity best-practices dialed up to 11 and implemented well. Same is likely true for network and telecom companies; yes, there have been incidents, but these places by-and-large have competent cyber teams. Physical infrastructure (e.g. power plants) I am not quite as confident on, but it seems like the sort of risk that, if nothing else, government natsec-type people would be able to recognize and work with the relevant entities to handle properly (e.g. by airgapping critical systems from the network, etc.).

But there are two areas that I am much more bearish on: financial and healthcare. These places frequently lack any meaningful cyberdefense capabilities, nor do they make serious efforts to attract top-level technical talent (just check out levels.fyi and compare SWE roles at a bank/​hospital/​etc. versus an FAANG). In fact, they are already breached frequently; the only reason why it isn’t even more frequent is that cyberoffense remains substantially bottlenecked by human talent, a bottleneck that will go away with sufficient AI capabilities and is already starting to go away, as shown by the recent FortiGate incident. And these institutions are very slow-moving (many banks still have much of their code in COBOL!), meaning they are likely going to be, in many cases, structurally unable to integrate defensive use of AI to a sufficient extent prior to when attackers can access sufficient AI capabilities to be able to compromise them. Yes, these industries are also highly risk-averse, but not in a way that is helpful here unfortunately.

This could have very severe implications. Consider what happens if a bank suffers a cyberattack sufficiently severe to lead to its insolvency (either directly or by harming depositor confidence enough to cause a massive outflow of funds). At that point, the FDIC saves the day, right? Not necessarily. A sufficient portion of all US banks failing in a correlated way all at once is not what the FDIC is built for. The FDIC has $157.4 billion in the DIF against roughly $11 trillion in insured deposits. So this could not only take out banks but also compromise the backstop that is normally used to prevent people from losing their life savings in such situations, leaving countless people broke.

Even worse would be if a hospital or other healthcare facility was compromised. This could bring down key systems required to provide effective care (e.g. medical records), risking severe patient harm or death. This has already happened in some cases, when ransomware has paralyzed hospitals and patients have died as a result. And, when cyberoffense can be automated, this could become far more common.

How can we prepare for this?

So, long story short, open-weights Mythos is very likely coming and will be very bad when it does. What can be done about this by society in advance? What can you, as an individual, do to prepare for it?

Starting with the society-level, I think a substantial amount can be done with just airgapping alone, especially in areas like healthcare and energy infrastructure. This doesn’t seem like it would be especially hard to get mandated through regulation (except in the sense that governments struggle to get things done in a general sense). It is not a hot-button partisan issue, nor can I think of any powerful lobbying group that would be sufficiently annoyed by it to push back strongly against it. I actually think airgapping mandates for safety-critical systems could be quite politically viable, especially since it could be framed as a US-China national security issue.

For areas like the financial sector, it’s more complicated, as their critical systems genuinely do have to be network-connected in many cases. I think there’s still some room to prepare, but that much of the work here will unfortunately have to be done after a crisis has already begun (which, sadly, will likely mean a lot of people losing serious amounts of money in the meantime). I suppose legacy corporations do listen to strategy consultants (e.g. MBB) and can make big changes fairly quickly in response, so it’s theoretically possible that, if the right consultants convince the right decision-makers at the right time, at least some financial institutions may be able to get their acts together, but that seems unlikely. It’s also possible that, if SWEs get displaced from more dynamic companies that are capable of automating work more quickly, you could have extreme competition for the remaining seats at less-dynamic companies that have yet to implement the same AI automation, paradoxically leading to an influx of talent. But I wouldn’t count on that either.

So the next question becomes: what can you do as an individual, assuming you don’t want to wake up broke one day? Big banks and smaller regional banks are probably roughly equally risky here; the bigger banks are bigger targets but also probably have slightly-less-awful cybersecurity. Self-custody crypto is a terrible, terrible choice here; these get compromised regularly as it is, and there is absolutely zero recourse, even in theory, if it does (there are crypto depeg/​hack insurance providers out there, but these will almost certainly go under if there is a sudden, correlated influx of claims). I actually think that physical cash (stored somewhere secure) could be semi-reasonable here; if the >=M1 money supply was reduced by compromises of financial infrastructure, that would likely be deflationary. That said, I think this is problematic for other reasons; after all, AI has, and likely will continue to, increase the S&P, meaning you miss out on that if you keep it all in cash. Precious metals have a similar issue to cash, in that they are largely uncorrelated with the sorts of things you’d expect to go up with AI, so you’d be missing out on substantial opportunity there. I suppose it could still make sense to keep some percent of one’s portfolio in something physical, but I wouldn’t overindex on that (not financial advice, talk to a financial advisor first). Paper stock certificates do technically still exist but are usually very rare and exorbitantly expensive even when they are available. I suppose the standard stuff (2FA, strong passwords, etc.) could help slightly on the margin, but, in the end, if the bank’s systems themselves are compromised, all of those account-level protections are moot. Keeping paper copies of financial documentation showing one’s assets could help to some extent if it becomes necessary to prove that one had the assets they claim to have had, although it would presumably have to be something officially certified or notarized (not just a standard printed-out statement) if it is to have any probative value. But I honestly don’t think there is a clean way to fully protect against this as an individual. The bottom line is that, if cyberoffense can access Mythos-level capabilities, this will, by default, be catastrophic for many areas of our society. And I’m not sure how, or if, one can fully prevent that (but of course happy to hear your thoughts in the comments).

Note: this post, including the ideas and all of the writing, is my own. After writing, I lightly revised it using AI (Opus 5) to improve clarity.