I can't yet clearly imagine a correlated attack on a large percentage of the banking system. Will the criminals wire a few trillion dollars... somewhere? Can't the US government say "no, this huge transfer of dollars is not ok"?
Strong agree. I’ve been working on a similar post - I’ve been calling this “the Hackening” - so I’ll put the thoughts I’ve had below.
As I understand it, pointing lesser agents at a section of code and saying “find the exploit”, in parallel, with high false positive rates, is not the same thing as a better agent grokking the whole environment and finding and chaining exploits together, at low false positive rates, into working steps of a cyberattack. The latter is the Juice.
I agree that harnesses can often make up the gap between open and frontier, and I haven’t seen anything that rules out that Kimi K3 swarms guided by T3MP3ST, or suchlike, could be worming their way across the Internet very soon.
Speaking of worms. Right now I classify threat actors into three fuzzy groups:
As for how many hacks can happen. Sol is telling me - so take this with heaps of salt - using Mythos’ weights, indexing off of The Last Ones, you get roughly .1-10 pwns of weak, flat-topography enterprises per cluster of 32 H200s per hour. There are dozens of tail compute holders that have hundreds of H200 equivalents. Mid tiers have thousands, or tens or hundreds of thousands. Hyperscalers have millions.
In dollar terms, that’s maybe $25-250 per pwn.
If an entity can simply steal compute - LLMjacking is a thing - then it gets it for free.
But datacenters that aren’t part of a nation’s arsenal won’t put up with this. The severity of the scenario depends a lot on how well they get their act together and shut out compute thieves, and either voluntarily - or under pressure from governments - start looking for and shutting out malicious paid inference.
Because the alternative is: there’s - again, Sol’s wild guess - perhaps 2 million easily pwnable orgs in the US, out of ~4.5 million. So, one month of the full compute of a single compromised or paid fleet of 100,000 H200 equivalents pwns ~half the US.
…Or maybe it’s a lot more. Or a lot less. I don’t know! I’d really like to see someone make a serious model of this.
Other considerations:
Overlap. If the open weights become common knowledge at the same time, possibly hundreds or thousands of actors will be fighting each other to get at the same pool of targets. The least defended targets will go down first, and the competition will work its way up from there.
Distillation. The smaller a model the Juice can run on, the cheaper attacks get. But if someone can get below certain thresholds - I believe 70B or 30B parameters - that opens up a lot more marginal compute to host malicious AIs, and greatly reduces the degree to which datacenter discipline can control outbreaks. Universities, workstations, hobbyists. Quite possible over the next few years at current scaling and algorithmic improvement trends. If it can fit on 7B - i.e. a phone- it is even worse.
Core services. AI suggests to me that if you can route your IT through a trusted security layer - formally verified, memory safe, frontier AI patrolled - that greatly reduces a lot of classes of vulnerability. Of course, the remaining attack efforts will concentrate on the remaining weak points. But this could help shift the balance.
Business realities. If an attack takes down an economic chokepoint, or a cluster of orgs that collectively act as a chokepoint, that could cascade through the economy. Also, ~half of businesses don’t have enough cash flow to deal with a month of down time and would simply collapse. Insurance will start forcing businesses to remediate their technical debt as soon as they start seeing the writing on the wall. If the (perception of the) business environment, or stability of digital banking, gets bad enough, this could trigger a credit contraction. And from there, a recession.
Attribution. If the chaos gets bad enough, or exploit windows start closing as things get patched, governments might decide now is the time to make their moves and accomplish cyber objectives they would’ve never considered otherwise. Many actors will be motivated to false flag themselves, and might succeed. This could lead to heightened international tensions.
New normal. The persistence of AI worms, and holdout copies of the weights among criminals and terrorists, means this won’t go away. Even within trust boundaries, there could be flareups as hibernating harnesses holding a cache of resources periodically reactivate themselves.
The labs. What happens? Do their IPOs break down because the cost of business suddenly goes up? Or public sentiment turns against them? Or the downturn resulting from this slows revenue growth and deployment? Do the widespread hacks disrupt the supply chain they rely on? Do they get nationalized? Do they get even more funding because now everyone wants their AI-driven security and rebuilding services?
And lastly, AI governance. If it gets bad enough, it will, in my opinion, the first real fire alarm that everyone can coordinate on. The AI safety community should be prepared for this. Policymakers will be in extraordinary mode, and will pick up whatever solutions are lying around. Hopefully, they will have good ones.
I agree there's a big challenge coming, but we also have an opportunity to improve cybersecurity, as I've written about elsewhere.
Briefly: your narrative presents computer systems like biological systems, implicitly created by an evolutionary process we have no control over, so of course obscure vulnerabilities are inevitable. However, we have the chance to integrate new vulnerability-checking into the software-development lifecycle. Even better, we can lean into formal verification to prove mathematically that systems are secure, so future models can't find bugs of certain kinds. They both sound like hard goals from a year-2020 perspective, but generative AI is accelerating both dramatically. In fact, I recently argued here that the time has nearly come to routinely throw away all old code and regenerate from scratch following new wisdom in security and other requirements dimensions -- because the cost of reliable software engineering should drop so dramatically.
indeed, many nations will be motivated to acquire this level of technology just to have a seat at the table.
Long story short: in my assessment, there is an 85% chance we will end up, in the next 24 months, with an open-weights model, or system thereof, capable of "The Juice" that models such as Mythos have, with respect to cybersecurity at the very least. This post goes into why that will likely happen, what the implications are, and how we, as a society and as individuals, can respond to it if/when it does.
First off: why do I say it's so likely? Like, couldn't China just...ban open-weights models, and then that solves the problem? Not so fast. Yes, China currently dominates the open-weights frontier. However, there are also open-weights AI labs in plenty of other countries (US, France, the UAE, South Korea, and Canada come to mind, and I'm sure there are others). Yes, some of these are substantially behind the frontier, but each of the aforementioned countries has an open-weights model no more than 24 months behind the current frontier (that's why I said 24 months earlier); remember, in Aug 2024, 24 months ago as of when I am writing this, the strongest models were Sonnet 3.5 and GPT 4o.
It seems highly unlikely that all of these different countries, with their disparate political systems, geopolitical ties, regulatory environments, and more, will all independently adopt the exact same policy of fully banning all open-weights model releases, at least not until it's too late and an open-weights model capable of Mythos-level capabilities has already been released. And, given the current geopolitical situation, it seems highly unlikely that there will be any international agreement to this effect (at least not until a disaster has already occurred and it's already too late).
But let's say we do manage to successfully conduct an international ban on open-weights Mythos-level models. At that point, you would also need to assume that each and every single closed-weights lab with Mythos-level capabilities also has sufficiently competent cybersecurity to be able to reliably not have their model weights exfiltrated. This seems highly unlikely across so many different labs, and an affirmative duty of having sufficient cybersecurity seems like an even taller order than a ban on open-weights training, which, as discussed before, is, in and of itself, highly unlikely to be possible on an international basis.
Mythos-level capabilities are likely possible with today's open-weights frontier
(epistemic status: I'm not as confident in this part as I am in the rest of this post; I genuinely want to hear others' thoughts here)
This next part is a bit more controversial, but I'd argue that the current frontier of open-weights models (probably Kimi K3), as it is now, with sufficient fine-tuning, inference-time compute, and harness improvements, could be used to get to a Mythos level in cyber capabilities (or capabilities in any other verifiable domain). Most obviously, a model could be fine-tuned to enhance dangerous capabilities in an area like cyber; this is well-established in the research and has already been done to some extent; see models like Hackphyr for example.
Additionally, inference-time compute has the potential to substantially enhance the capabilities of Kimi K3. For example, there are approaches like best-of-N, multi-agent debate, tree-of-thought, etc. It seems that a lot of why weaker models lack "The Juice" is because they, at a certain point, get off-track in a way that prevents them from recovering, and approaches like best-of-N could sharply reduce those problematic points.
You could even achieve a similar result without anything fancy: groups like AISLE Security have shown that open-weights models can replicate much of Mythos's analysis if pointed at the correct file. You could use Dynamic Workflows or something similar and just point subagents at all of the applicable files (keep in mind you realistically wouldn't have to do a subagent for every single file in a large codebase, as it's only a subset of those files that are realistically likely to be exploitable).
Likewise, I expect substantial improvements to be possible just from better harnesses; for example, better memory systems, better multi-agent systems, etc.
Because of those factors, I think it's quite likely that, even with current open-weights models as they are today, we will end up in a situation where attackers have access to Mythos-level cyber capabilities. Keep in mind that cybercrime is extraordinarily lucrative (breaches like Bybit have yielded attackers as much as $1.5B), so attackers would gladly be willing to pay even truly exorbitant token expenditures in order to conduct an attack; cost is not a meaningful bottleneck here, and with Cerebras/Taalas/etc., token rate almost certainly won't be a bottleneck either. The one saving grace, as discussed earlier, is that this is only applicable to verifiable tasks like cyber; it is less applicable for less-verifiable tasks like bio, where harnesses/inference-time compute/etc. play less of a role and it's more of a "either you know it or you don't" sort of situation. But, for the reasons discussed below, open-weights Mythos capabilities, even just for cyber, are likely to be extremely harmful.
This will be very bad
The fact is that many of our most important organizations (both in terms of how severe the impact would be if they were compromised and in terms of their attractiveness to hackers) are fundamentally not ready for a world where anyone can conduct a cyberattack. I actually expect FAANGs/other big tech companies to be able to prepare reasonably well from a defensive standpoint using some combination of closed-weight frontier models (which, at that point, will be far stronger than Mythos 5) along with just normal cybersecurity best-practices dialed up to 11 and implemented well. Same is likely true for network and telecom companies; yes, there have been incidents, but these places by-and-large have competent cyber teams. Physical infrastructure (e.g. power plants) I am not quite as confident on, but it seems like the sort of risk that, if nothing else, government natsec-type people would be able to recognize and work with the relevant entities to handle properly (e.g. by airgapping critical systems from the network, etc.).
But there are two areas that I am much more bearish on: financial and healthcare. These places frequently lack any meaningful cyberdefense capabilities, nor do they make serious efforts to attract top-level technical talent (just check out levels.fyi and compare SWE roles at a bank/hospital/etc. versus an FAANG). In fact, they are already breached frequently; the only reason why it isn't even more frequent is that cyberoffense remains substantially bottlenecked by human talent, a bottleneck that will go away with sufficient AI capabilities and is already starting to go away, as shown by the recent FortiGate incident. And these institutions are very slow-moving (many banks still have much of their code in COBOL!), meaning they are likely going to be, in many cases, structurally unable to integrate defensive use of AI to a sufficient extent prior to when attackers can access sufficient AI capabilities to be able to compromise them. Yes, these industries are also highly risk-averse, but not in a way that is helpful here unfortunately.
This could have very severe implications. Consider what happens if a bank suffers a cyberattack sufficiently severe to lead to its insolvency (either directly or by harming depositor confidence enough to cause a massive outflow of funds). At that point, the FDIC saves the day, right? Not necessarily. A sufficient portion of all US banks failing in a correlated way all at once is not what the FDIC is built for. The FDIC has $157.4 billion in the DIF against roughly $11 trillion in insured deposits. So this could not only take out banks but also compromise the backstop that is normally used to prevent people from losing their life savings in such situations, leaving countless people broke.
Even worse would be if a hospital or other healthcare facility was compromised. This could bring down key systems required to provide effective care (e.g. medical records), risking severe patient harm or death. This has already happened in some cases, when ransomware has paralyzed hospitals and patients have died as a result. And, when cyberoffense can be automated, this could become far more common.
How can we prepare for this?
So, long story short, open-weights Mythos is very likely coming and will be very bad when it does. What can be done about this by society in advance? What can you, as an individual, do to prepare for it?
Starting with the society-level, I think a substantial amount can be done with just airgapping alone, especially in areas like healthcare and energy infrastructure. This doesn't seem like it would be especially hard to get mandated through regulation (except in the sense that governments struggle to get things done in a general sense). It is not a hot-button partisan issue, nor can I think of any powerful lobbying group that would be sufficiently annoyed by it to push back strongly against it. I actually think airgapping mandates for safety-critical systems could be quite politically viable, especially since it could be framed as a US-China national security issue.
For areas like the financial sector, it's more complicated, as their critical systems genuinely do have to be network-connected in many cases. I think there's still some room to prepare, but that much of the work here will unfortunately have to be done after a crisis has already begun (which, sadly, will likely mean a lot of people losing serious amounts of money in the meantime). I suppose legacy corporations do listen to strategy consultants (e.g. MBB) and can make big changes fairly quickly in response, so it's theoretically possible that, if the right consultants convince the right decision-makers at the right time, at least some financial institutions may be able to get their acts together, but that seems unlikely. It's also possible that, if SWEs get displaced from more dynamic companies that are capable of automating work more quickly, you could have extreme competition for the remaining seats at less-dynamic companies that have yet to implement the same AI automation, paradoxically leading to an influx of talent. But I wouldn't count on that either.
So the next question becomes: what can you do as an individual, assuming you don't want to wake up broke one day? Big banks and smaller regional banks are probably roughly equally risky here; the bigger banks are bigger targets but also probably have slightly-less-awful cybersecurity. Self-custody crypto is a terrible, terrible choice here; these get compromised regularly as it is, and there is absolutely zero recourse, even in theory, if it does (there are crypto depeg/hack insurance providers out there, but these will almost certainly go under if there is a sudden, correlated influx of claims). I actually think that physical cash (stored somewhere secure) could be semi-reasonable here; if the >=M1 money supply was reduced by compromises of financial infrastructure, that would likely be deflationary. That said, I think this is problematic for other reasons; after all, AI has, and likely will continue to, increase the S&P, meaning you miss out on that if you keep it all in cash. Precious metals have a similar issue to cash, in that they are largely uncorrelated with the sorts of things you'd expect to go up with AI, so you'd be missing out on substantial opportunity there. I suppose it could still make sense to keep some percent of one's portfolio in something physical, but I wouldn't overindex on that (not financial advice, talk to a financial advisor first). Paper stock certificates do technically still exist but are usually very rare and exorbitantly expensive even when they are available. I suppose the standard stuff (2FA, strong passwords, etc.) could help slightly on the margin, but, in the end, if the bank's systems themselves are compromised, all of those account-level protections are moot. Keeping paper copies of financial documentation showing one's assets could help to some extent if it becomes necessary to prove that one had the assets they claim to have had, although it would presumably have to be something officially certified or notarized (not just a standard printed-out statement) if it is to have any probative value. But I honestly don't think there is a clean way to fully protect against this as an individual. The bottom line is that, if cyberoffense can access Mythos-level capabilities, this will, by default, be catastrophic for many areas of our society. And I'm not sure how, or if, one can fully prevent that (but of course happy to hear your thoughts in the comments).
Note: this post, including the ideas and all of the writing, is my own. After writing, I lightly revised it using AI (Opus 5) to improve clarity.