Applied AI security practitioner with a cybersecurity background, and formative rationalist grounding in philosophy and competitive debate
tr5tn
Interesting to note that GPT-5.6 was trained using GPT-Red as an adversary (current post is unclear on the precise method) and that GPT-Red itself was trained through self-play. I wonder if red team/blue team is a natural domain for wider debate testing. https://openai.com/index/unlocking-self-improvement-gpt-red/ Fuller details published later this week.
I just wanted to thank you for building in the initial framing for a wider audience. It’s not especially intuitive that debate research would focus on training rather than inference, so this is very welcome. And the work itself is super interesting. Really interested to see where your focus on the protocols might arrive.
@ethanelasky per the other thread, this was pretty insightful. Can I ask, how have you settled on the better judges and critics? Have you introduced anything specific to encourage those skills? And I’m a bit confused about the number of rebuttals, and the implementation of/results of the rebuttals in general. In other words, I can discern the (lack of) improvement from rebuttals, but not clear what was actually seen there. I couldn’t pick that out in the transcripts?
One thing that’s probably more immediate and direct than some of these possibilities is an AI safety corrolary of the attack surface of security technology. There’s a long and embarrasing history of highly-privileged security (and wider management/resilience) technologies introducing their own (sometimes inadequately protected) attack surface. AI safety processes and technologies come with some of those same risks. As the stakes get higher and the interventions more complex, we’ll be introducing a growing set of novel risks. This is probably most applicable to AI control, but certainly not exclusive to it.
Agent Identity Standardisation Efforts
Not disputing they might be able to fall back to it. But in the Glasswing context I would have expected them to have updated already, given the UK AISI results on The Last Ones. Just speculation of course, and I know some Glasswing partners weren’t (exclusively) using Mythos (see the Visa harness as one example) https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber-capability-advancing
Presumably they would have updated.
Thanks Ethan! I’ve just seen your research from this year, which I’m going to digest at a sensible pace. The more recent one looks especially interesting.
Totally agree about the conceptual update. I hadn’t seen that original research when I first wrote the post, and agree it would have given me much to chew on.
I also take the point on optimisation pressures, and have been thinking about that a lot since David’s comments. The more I think about the human debate references, the more I’m inclined to think the human protocols haven’t advanced enough. Law has plenty of its own problems. Large-scale formats like MUN descend into contests of vocal or rhetorical strength. When I look at all of that, I feel that competitive debate starts to address many of those challenges through strict rules, while admiting of imperfections that the constraints introduce (you can only reach limited depth, and there’s the artifice/flaws from optimisation for those rules). But in so many contexts those types of failure become a reason to add deeper refinements rather than reverting to a simpler form. Maybe we’ve reached the limit of what these imperfect games can teach us, even if the AI safety technique could go deeper in its own direction (maybe following complexity theory, as you say—a topic where I am completely out of my depth).
In any case, keen to get stuck in to yourrecent posts.
Agree with pretty much all of this, The thing is, even the Anthropic Red report on Mythos is pretty clear that this isn’t simply more vulnerabilities. It’s improved discovery, exploitation, chains of exploitations, and fixing vulnerabilities. I find the primary discourse really frustrating because it sticks to the most quantifiable headline while avoiding the detail. Mythos is important, but it only tells us the current state of a very powerful general-purpose model and scaffold. But…
GPT-5.5 is benchmarked not far off it, and is already available in public (while there is gated Trusted Access for security testing scenarios, we can’t ignore that this was released, when Mythos wasn’t). And Opus 4.8 is out.
There have been many examples of less powerful (and 100x cheaper) models being used with cybersecurity-specific tooling, achieving similar results. The cost dimension is extremely important, since most attacks remain financially motivated.
Even six months ago, with Opus 4.5 (I believe) and GPT-4.1 the breach in Mexico was brutal and unprecedented in AI-assisted scope https://cdn.prod.website-files.com/69944dd945f20ca4a27a7c47/69d8bb5aea59e31efb3b8a7f_Tech_Report_ai_breach_mex_gov.pdf?trk=public_post_comment-text. This report is an essential read beside our current preoccupations, because it shows the scale of damage that was possible with less capable AI.
Multi-model harnesses like Microsoft’s MDASH show that the best/newest model can be matched or exceeded by ensembling.
All of this is against a backdrop of much shorter mean time-to-exploit periods (for all CVEs and for zero-days specifically). Mean TTE has dropped from 2.3 years to 24 hours over the last eight years. https://zerodayclock.com/
Mean time-to-remediate in most organisations is completely out of step with these changes. The bigger problem is applying fixes, rather than creating the fixes. The pressure on prioritising remedial efforts hits the limits of IT team understanding very rapidly in most organisations, and scaling to inceased patching (or other mitigation) burdens with teams running on fumes is already a big problem, and why we have most prominent security authorities agitating to get “Mythos-ready”.
Automation has a role to play here, but security fundamentals for update mechanisms (and code repositories feeding into package repositories) are pretty weak, and routinely being used as their own attack vector. Adaptation efforts aren’t as simple as routinely applying the updates (see the TeamPCP reference above).
As fixes are pushed more rapidly, there will be more breaking changes and other regressions. Testing patches before applying them is largely a myth. Most organisations won’t be ready to push updates cautiously in deployment rings, and disruption from updates will introduce opposing pressure to ignore updates.
…so imagine a best of breed cybersecurity-specific scaffold with mutliple foundation models and potentially significant token budgets.
We’re already starting to see the impact on patch volumes, and the deluge isn’t here yet. Most organisations have fundamentally weak protections to start. Many have vulnerability management efforts comprised entirely of automatic updates. Many apps never get updated, there is typically very poor visibility of update statuses, and most importantly, this often isn’t anyone’s job. There is insufficient staff, skill, understanding and maturity to adapt. Most leadership won’t prioritise adaptation quickly enough.
IMO, this isn’t as big of a problem for software updates from Glasswing vendors. For the most part, those update processes are the ones that will be in place. It’s everything else that will really be problematic.
At the end of this, the biggest question will be how many bad actors are willing to exploit the situation, because so much fruit is low-hanging.
I take the point that in the self-play context this could drift off-course! I suppose (linking this back to the MATS research) I’m suggesting it would be good to measure that beside a more naïve protocol.
@David Africa thanks! Many of those points are certainly worth focusing on. For what it’s worth, I was also an awarded speaker in the Model UN, but I found that format to be far more arbitrary, susceptible to being gamed by speaking skill and rhetoric, and IMO less likely to arrive at something desirable (I led an uprising of militarised third-party countries to vote down all disarmament proposals).
Ultimately, the actual plans, counterplans, kritiks and topicality discussions in policy debate are ridiculous. Every debater I’ve met would acknowledge that. And ultimately, it is a game, so IMO that is to be expected. So I am certainly not agitating for AI Safety outcomes that resemble policy debate verdicts, but I think the game itself is good reference since one of the current problems in AI Safety debate protocols is that they are being gamed. Distinctions between Constructives, Rebuttals and Cross-Examination are really fundamental to policy debate, and we get similar constructs in legal proceedings.
I’m conscious this is all one American reference and I’m not offering empirical findings, but I do think the field should consider these games as reference protocols (as well as cross-cultural legal rules) since they are a result of refinement over decades. They are practices that already exist.
--
Edited to add: not wishing to be dismissive of your empirical findings. I’d love to read more about the difficulties with training or inference-time persona adoption, but also, I don’t know that current negative findings should preclude focus on those problems.
@ethanelasky @David Africa I found a pretty divergent approach to AI debate that may be of interest. It seems to focus on a lot of the historical shortcomings you mentioned originally, like struggling to assume the perspective of one side. They’ve introduced some interesting techniques like HEXIS, which modifies attention for a specific disposition (while remaining hidden from semantic memory). In any case, it’s an approach where the researchers have fully commited to policy debate structure and dialectical reasoning, and have had some positive results. Their open questions are intriguing too. https://www.debaterhub.com/blog/projects/dialectical-debate-ai