The chembio risks section of the Opus 5 system card contains serious rigor issues. While I agree with Anthropic’s bottom-line conclusions, their reasoning is deeply flawed.
Opus scored similarly to (or better than) Mythos 5 on ~every automated benchmark, and Anthropic didn’t run human-intensive evals (e.g., uplift studies) because they didn’t have time. Despite this, they assess that overall it’s probably not as dangerous as Mythos because it is a generally worse model, and an early checkpoint was qualitatively worse on a long-horizon biological task. Basing most of the risk assessment on qualitative assessments and providing a single example where you don’t even use a late-stage model checkpoint is an extremely uncomfortable precedent, and Anthropic’s risk conclusions are too confident.
Based on the evidence provided, it seems totally reasonable that scaling up bio RL (which they clearly did for Opus 5!) means that Opus performs better than Mythos for bioweapon uplift in particular, even if it’s worse at other agentic coding tasks.[1]
Generally, capabilities are weird enough across domains that if performance on all the automated benchmarks breach your risk levels, plausibly the non-automated ones might as well! You don’t just get to not run them and guess!
At the very least, the obvious thing to do once the automated evals aren’t providing signal is to run a fast human-intensive eval that you expect Mythos to do much better at as a smoke test (sorry). Or, simply broadcast more clearly that you’re somewhat uncertain in your assessment, and use evidence from other parts of the system card about long-run autonomy to support your qualitative capability claims.
I shared an edited version of this message in a private Slack, and was encouraged to share it publicly. Crosspost from Twitter.
[1] It is unclear whether Opus is worse than Mythos at agentic coding tasks. Opus is comparable to Mythos on AECI and across many agentic benchmarks and on UK AISI’s cyber ranges, but performs significantly worse on many cyber benchmarks. Opus seems clearly worse than Mythos on difficult math and QA tasks.
The chembio risks section of the Opus 5 system card contains serious rigor issues. While I agree with Anthropic’s bottom-line conclusions, their reasoning is deeply flawed.
Opus scored similarly to (or better than) Mythos 5 on ~every automated benchmark, and Anthropic didn’t run human-intensive evals (e.g., uplift studies) because they didn’t have time. Despite this, they assess that overall it’s probably not as dangerous as Mythos because it is a generally worse model, and an early checkpoint was qualitatively worse on a long-horizon biological task. Basing most of the risk assessment on qualitative assessments and providing a single example where you don’t even use a late-stage model checkpoint is an extremely uncomfortable precedent, and Anthropic’s risk conclusions are too confident.
Based on the evidence provided, it seems totally reasonable that scaling up bio RL (which they clearly did for Opus 5!) means that Opus performs better than Mythos for bioweapon uplift in particular, even if it’s worse at other agentic coding tasks.[1]
Generally, capabilities are weird enough across domains that if performance on all the automated benchmarks breach your risk levels, plausibly the non-automated ones might as well! You don’t just get to not run them and guess!
At the very least, the obvious thing to do once the automated evals aren’t providing signal is to run a fast human-intensive eval that you expect Mythos to do much better at as a smoke test (sorry). Or, simply broadcast more clearly that you’re somewhat uncertain in your assessment, and use evidence from other parts of the system card about long-run autonomy to support your qualitative capability claims.
I shared an edited version of this message in a private Slack, and was encouraged to share it publicly. Crosspost from Twitter.
[1] It is unclear whether Opus is worse than Mythos at agentic coding tasks. Opus is comparable to Mythos on AECI and across many agentic benchmarks and on UK AISI’s cyber ranges, but performs significantly worse on many cyber benchmarks. Opus seems clearly worse than Mythos on difficult math and QA tasks.
Note this is par for the course (at least as of a year ago, when I followed this stuff): AI companies’ eval reports mostly don’t support their claims.