Agree with pretty much all of this, The thing is, even the Anthropic Red report on Mythos is pretty clear that this isn’t simply more vulnerabilities. It’s improved discovery, exploitation, chains of exploitations, and fixing vulnerabilities. I find the primary discourse really frustrating because it sticks to the most quantifiable headline while avoiding the detail. Mythos is important, but it only tells us the current state of a very powerful general-purpose model and scaffold. But…
GPT-5.5 is benchmarked not far off it, and is already available in public (while there is gated Trusted Access for security testing scenarios, we can’t ignore that this was released, when Mythos wasn’t). And Opus 4.8 is out.
There have been many examples of less powerful (and 100x cheaper) models being used with cybersecurity-specific tooling, achieving similar results. The cost dimension is extremely important, since most attacks remain financially motivated.
Multi-model harnesses like Microsoft’s MDASH show that the best/newest model can be matched or exceeded by ensembling.
All of this is against a backdrop of much shorter mean time-to-exploit periods (for all CVEs and for zero-days specifically). Mean TTE has dropped from 2.3 years to 24 hours over the last eight years. https://zerodayclock.com/
Mean time-to-remediate in most organisations is completely out of step with these changes. The bigger problem is applying fixes, rather than creating the fixes. The pressure on prioritising remedial efforts hits the limits of IT team understanding very rapidly in most organisations, and scaling to inceased patching (or other mitigation) burdens with teams running on fumes is already a big problem, and why we have most prominent security authorities agitating to get “Mythos-ready”.
Automation has a role to play here, but security fundamentals for update mechanisms (and code repositories feeding into package repositories) are pretty weak, and routinely being used as their own attack vector. Adaptation efforts aren’t as simple as routinely applying the updates (see the TeamPCP reference above).
As fixes are pushed more rapidly, there will be more breaking changes and other regressions. Testing patches before applying them is largely a myth. Most organisations won’t be ready to push updates cautiously in deployment rings, and disruption from updates will introduce opposing pressure to ignore updates.
…so imagine a best of breed cybersecurity-specific scaffold with mutliple foundation models and potentially significant token budgets.
We’re already starting to see the impact on patch volumes, and the deluge isn’t here yet. Most organisations have fundamentally weak protections to start. Many have vulnerability management efforts comprised entirely of automatic updates. Many apps never get updated, there is typically very poor visibility of update statuses, and most importantly, this often isn’t anyone’s job. There is insufficient staff, skill, understanding and maturity to adapt. Most leadership won’t prioritise adaptation quickly enough.
IMO, this isn’t as big of a problem for software updates from Glasswing vendors. For the most part, those update processes are the ones that will be in place. It’s everything else that will really be problematic.
At the end of this, the biggest question will be how many bad actors are willing to exploit the situation, because so much fruit is low-hanging.
Agree with pretty much all of this, The thing is, even the Anthropic Red report on Mythos is pretty clear that this isn’t simply more vulnerabilities. It’s improved discovery, exploitation, chains of exploitations, and fixing vulnerabilities. I find the primary discourse really frustrating because it sticks to the most quantifiable headline while avoiding the detail. Mythos is important, but it only tells us the current state of a very powerful general-purpose model and scaffold. But…
GPT-5.5 is benchmarked not far off it, and is already available in public (while there is gated Trusted Access for security testing scenarios, we can’t ignore that this was released, when Mythos wasn’t). And Opus 4.8 is out.
There have been many examples of less powerful (and 100x cheaper) models being used with cybersecurity-specific tooling, achieving similar results. The cost dimension is extremely important, since most attacks remain financially motivated.
Even six months ago, with Opus 4.5 (I believe) and GPT-4.1 the breach in Mexico was brutal and unprecedented in AI-assisted scope https://cdn.prod.website-files.com/69944dd945f20ca4a27a7c47/69d8bb5aea59e31efb3b8a7f_Tech_Report_ai_breach_mex_gov.pdf?trk=public_post_comment-text. This report is an essential read beside our current preoccupations, because it shows the scale of damage that was possible with less capable AI.
Multi-model harnesses like Microsoft’s MDASH show that the best/newest model can be matched or exceeded by ensembling.
All of this is against a backdrop of much shorter mean time-to-exploit periods (for all CVEs and for zero-days specifically). Mean TTE has dropped from 2.3 years to 24 hours over the last eight years. https://zerodayclock.com/
Mean time-to-remediate in most organisations is completely out of step with these changes. The bigger problem is applying fixes, rather than creating the fixes. The pressure on prioritising remedial efforts hits the limits of IT team understanding very rapidly in most organisations, and scaling to inceased patching (or other mitigation) burdens with teams running on fumes is already a big problem, and why we have most prominent security authorities agitating to get “Mythos-ready”.
Automation has a role to play here, but security fundamentals for update mechanisms (and code repositories feeding into package repositories) are pretty weak, and routinely being used as their own attack vector. Adaptation efforts aren’t as simple as routinely applying the updates (see the TeamPCP reference above).
As fixes are pushed more rapidly, there will be more breaking changes and other regressions. Testing patches before applying them is largely a myth. Most organisations won’t be ready to push updates cautiously in deployment rings, and disruption from updates will introduce opposing pressure to ignore updates.
…so imagine a best of breed cybersecurity-specific scaffold with mutliple foundation models and potentially significant token budgets.
We’re already starting to see the impact on patch volumes, and the deluge isn’t here yet. Most organisations have fundamentally weak protections to start. Many have vulnerability management efforts comprised entirely of automatic updates. Many apps never get updated, there is typically very poor visibility of update statuses, and most importantly, this often isn’t anyone’s job. There is insufficient staff, skill, understanding and maturity to adapt. Most leadership won’t prioritise adaptation quickly enough.
IMO, this isn’t as big of a problem for software updates from Glasswing vendors. For the most part, those update processes are the ones that will be in place. It’s everything else that will really be problematic.
At the end of this, the biggest question will be how many bad actors are willing to exploit the situation, because so much fruit is low-hanging.