AI agents are very capable and very useful. Here’s a decomposition of the kinds of capability they have or could have.
The models know a lot of stuff. A lot of human knowledge is in the weights, meaning even they can look up stuff, they often don’t know need to and/or their encyclopaedic knowledge means they know what to look up. For example, without looking, they already know all cybersecurity 101, 201, and possibly more.
They can operate very quickly. Read quickly, write quickly. For any task that they can accomplish as well as I can on quality, they can do it faster. Write and run some tests on code? Single digit minutes.
Connecting disparate ideas. Yesterday I watched a video about how McLaren engineers took inspiration from the sailfish to optimize the aerodynamics of their sportscar. I don’t know that the agents so far do a tonne of this, but I’d say it’d benefit a lot from 1.
Deep understanding of causal structures. For whatever reason, they don’t readily do this. If you give them a scientific paper, they’ll more want to quote the author’s said than discuss what the evidence really shows based on an independent interpretation of the observations.
There’s then also 5., the skill of chaining together small tasks towards larger tasks. I think this is an area where models have gained a lot recently, allowing them to take advantage of the (2) their speed.
I suspect that right now models are being capable heavily because of 1. (knowing a lot of stuff) and 2. (operating quickly) without necessarily having more 3. (“connecting the dots”) than humans or engaging in (4) developing deep or novel understanding. They can hack better or solve hard math problems by sheer “brute force” of the standard playbook, plus knowing the playbook.
If you can perform a lot of very rapid simple experiments, you can quickly end up with better understanding of systems.
I predict we’ll see a scary jump as they get better at connecting the dots and doing the causal modeling though. It will stack on their ability to operate fast, which is already enough to make them powerful.
It’s scary because even if they were no better at reasoning, quality-wise, than humans, the speed would make them dominate human performance. But also I think there’s no reason they’ll cap out at human-level reasoning quality, plus even sub-human level “connect the dots” combined with simply knowing way more dots and operating very fast will let them dominate.
We can also see the swarm behavior as a huge further boost to speed via parallelization. No individual model is necessarily smart or faster on its own, but through tiling itself, it can get a lot more done.
They’re also good at noticing missing things or altered details between two things. I had Opus 5 load all of Sol’s and Astra’s model cards into its context:
“[…] Also note what vanished. Sol’s card made an explicit argument in point 6 that broad cyber access is net-positive because models are better at finding-and-fixing than attacking. That argument is completely absent from Astra’s card. It got replaced by phased Daybreak rollout, jurisdiction restrictions, and mandatory Advanced Account Security. The offense/defense-balance framing quietly disappeared at exactly the moment the model got good at offense.
[…]
METR is gone entirely. On Sol they reported a high detected cheating rate and declined to call the time-horizon result a robust capability measurement. No METR section at all for GPT-6.
Red-team compute dropped. 700,000 A100e GPU-hours for 5.6, versus “one measured portion” at ~200,000 for Astra. They attribute it to more efficient attackers (GPT-Red) and note both figures are partial. Still, a 3.5x smaller headline number on the model that crossed Critical.“
AI agents are very capable and very useful. Here’s a decomposition of the kinds of capability they have or could have.
The models know a lot of stuff. A lot of human knowledge is in the weights, meaning even they can look up stuff, they often don’t know need to and/or their encyclopaedic knowledge means they know what to look up. For example, without looking, they already know all cybersecurity 101, 201, and possibly more.
They can operate very quickly. Read quickly, write quickly. For any task that they can accomplish as well as I can on quality, they can do it faster. Write and run some tests on code? Single digit minutes.
Connecting disparate ideas. Yesterday I watched a video about how McLaren engineers took inspiration from the sailfish to optimize the aerodynamics of their sportscar. I don’t know that the agents so far do a tonne of this, but I’d say it’d benefit a lot from 1.
Deep understanding of causal structures. For whatever reason, they don’t readily do this. If you give them a scientific paper, they’ll more want to quote the author’s said than discuss what the evidence really shows based on an independent interpretation of the observations.
There’s then also 5., the skill of chaining together small tasks towards larger tasks. I think this is an area where models have gained a lot recently, allowing them to take advantage of the (2) their speed.
I suspect that right now models are being capable heavily because of 1. (knowing a lot of stuff) and 2. (operating quickly) without necessarily having more 3. (“connecting the dots”) than humans or engaging in (4) developing deep or novel understanding. They can hack better or solve hard math problems by sheer “brute force” of the standard playbook, plus knowing the playbook.
If you can perform a lot of very rapid simple experiments, you can quickly end up with better understanding of systems.
I predict we’ll see a scary jump as they get better at connecting the dots and doing the causal modeling though. It will stack on their ability to operate fast, which is already enough to make them powerful.
It’s scary because even if they were no better at reasoning, quality-wise, than humans, the speed would make them dominate human performance. But also I think there’s no reason they’ll cap out at human-level reasoning quality, plus even sub-human level “connect the dots” combined with simply knowing way more dots and operating very fast will let them dominate.
We can also see the swarm behavior as a huge further boost to speed via parallelization. No individual model is necessarily smart or faster on its own, but through tiling itself, it can get a lot more done.
They’re also good at noticing missing things or altered details between two things. I had Opus 5 load all of Sol’s and Astra’s model cards into its context:
“[…] Also note what vanished. Sol’s card made an explicit argument in point 6 that broad cyber access is net-positive because models are better at finding-and-fixing than attacking. That argument is completely absent from Astra’s card. It got replaced by phased Daybreak rollout, jurisdiction restrictions, and mandatory Advanced Account Security. The offense/defense-balance framing quietly disappeared at exactly the moment the model got good at offense.
[…]
METR is gone entirely. On Sol they reported a high detected cheating rate and declined to call the time-horizon result a robust capability measurement. No METR section at all for GPT-6.
Red-team compute dropped. 700,000 A100e GPU-hours for 5.6, versus “one measured portion” at ~200,000 for Astra. They attribute it to more efficient attackers (GPT-Red) and note both figures are partial. Still, a 3.5x smaller headline number on the model that crossed Critical.“