in the graphs they show, mythos looks like a step change—its introduction dwarfs everything else on the y-axes.
the interpretation that anthropic wants the reader to make is like: “the y-axes are a proxy for d(capabilities)/dt. code is written faster → claudes get developed faster → coding is even faster → etc.”
however, the y-axes are also a measure of capabilities itself. (hence, under anthropic’s interpretation, dC/dt ~ C and C grows exponentially.) and i would bet that mythos is a much larger-scale model in terms of training compute, and that that is why it is so much more capable.
but the bottleneck on training compute is not code-writing or research; these things are necessary complements to increased compute capacity, needed to leverage the increased capacity effectively, but it’s still true that the data centers come online at their own pace and writing 10x more LoC won’t make them come any faster.
so a contrarian interpretation would be: “these plots show that to a first order approximation, training compute is all that matters for capabilities (or at least for the proxies on the y-axes), and automating coding/research with claude code does not magically give you more compute, so this extra coding/research velocity is not meaningfully ‘building claude faster.’ and that’s why the y-axes show a step change between two plateaus rather than a clean exponential trend—the quantities plotted on the y-axes did not feed back into the dynamics, we’re just seeing a larger-scale model arrive exogenously and cause a step change (which only looks somewhat smooth because it takes a finite time to learn how to use it fully and because the moving averages in the “session success rate” plot have an inherent smoothing effect)”
It’s not necessarily the case that code and training compute are complements in an economic sense, even if compute is the reason why Mythos > Opus. Compute-efficiency has increased at something like 10x/year so more code can absolutely substitute for compute—except to the extent that code is bottlenecked on experiment compute.
Sure, I think that’s an important consideration and I pin some hopes of slowdown on it.
I’ll just make my standard caveat that there are plenty of routes to improvement that don’t depend strictly on scaling. Particularly salient is improved memory and metacognition/executive function. Both areas of active research, documented in those posts, so both are being improved independently of gains from scaling.
This isn’t likely to produce a step change, since progress on both is in amount not kind, but it’s another factor that will drive improvements in competence and capabilities, that’s not bottlenecked by compute.
Yeah, this has been probably the single biggest update I’ve made over the past couple of years, and yeah the trendline of training compute increasing is basically my first-pass guess on why METR’s trend has been as stable as it is.
And yeah, the plot on here basically validates the story where training compute is by far the most important determinant of AI progress, with everything else like algorithms/data contributing a massively smaller share of improvements.
Thank you, I also wanted to write something similar because of a similarly linear “scaling law” of AECI over time. As far as I understand, the current architecture doesn’t allow anyone to do the RSI for reasons similar to why humans cannot do RSI on themselves: capabilities are increased linearly over epochs or logarithmically over lived experience, which in the AIs’ case is proportional to compute spent.
Under this interpretation, eventually someone will understand it and come up with alternate architectures with their scaling laws (neuralese trained from scratch? Multiple CoTs receiving tokens in a single forward pass? Gemini Diffusion-like models?). In this case, a lab which cares about alignment will begin research in order to understand scaling laws of the new architectures and the ways to prevent such architectures from scheming in unnoticeable ways (SAE? NLA? Reliance on low capabilities of CoTless skills?) and either invent schemes to reliably align the AIs with newfound capabilities or end up failing to notice that an AI began to scheme and took over.
random thought about anthropic’s RSI post[1]:
in the graphs they show, mythos looks like a step change—its introduction dwarfs everything else on the y-axes.
the interpretation that anthropic wants the reader to make is like: “the y-axes are a proxy for d(capabilities)/dt. code is written faster → claudes get developed faster → coding is even faster → etc.”
however, the y-axes are also a measure of capabilities itself. (hence, under anthropic’s interpretation, dC/dt ~ C and C grows exponentially.) and i would bet that mythos is a much larger-scale model in terms of training compute, and that that is why it is so much more capable.
but the bottleneck on training compute is not code-writing or research; these things are necessary complements to increased compute capacity, needed to leverage the increased capacity effectively, but it’s still true that the data centers come online at their own pace and writing 10x more LoC won’t make them come any faster.
so a contrarian interpretation would be: “these plots show that to a first order approximation, training compute is all that matters for capabilities (or at least for the proxies on the y-axes), and automating coding/research with claude code does not magically give you more compute, so this extra coding/research velocity is not meaningfully ‘building claude faster.’ and that’s why the y-axes show a step change between two plateaus rather than a clean exponential trend—the quantities plotted on the y-axes did not feed back into the dynamics, we’re just seeing a larger-scale model arrive exogenously and cause a step change (which only looks somewhat smooth because it takes a finite time to learn how to use it fully and because the moving averages in the “session success rate” plot have an inherent smoothing effect)”
copied from a message i sent in a messaging app, hence the lowercase and abbreviated style
It’s not necessarily the case that code and training compute are complements in an economic sense, even if compute is the reason why Mythos > Opus. Compute-efficiency has increased at something like 10x/year so more code can absolutely substitute for compute—except to the extent that code is bottlenecked on experiment compute.
Sure, I think that’s an important consideration and I pin some hopes of slowdown on it.
I’ll just make my standard caveat that there are plenty of routes to improvement that don’t depend strictly on scaling. Particularly salient is improved memory and metacognition/executive function. Both areas of active research, documented in those posts, so both are being improved independently of gains from scaling.
This isn’t likely to produce a step change, since progress on both is in amount not kind, but it’s another factor that will drive improvements in competence and capabilities, that’s not bottlenecked by compute.
Yeah, this has been probably the single biggest update I’ve made over the past couple of years, and yeah the trendline of training compute increasing is basically my first-pass guess on why METR’s trend has been as stable as it is.
And yeah, the plot on here basically validates the story where training compute is by far the most important determinant of AI progress, with everything else like algorithms/data contributing a massively smaller share of improvements.
Thank you, I also wanted to write something similar because of a similarly linear “scaling law” of AECI over time. As far as I understand, the current architecture doesn’t allow anyone to do the RSI for reasons similar to why humans cannot do RSI on themselves: capabilities are increased linearly over epochs or logarithmically over lived experience, which in the AIs’ case is proportional to compute spent.
Under this interpretation, eventually someone will understand it and come up with alternate architectures with their scaling laws (neuralese trained from scratch? Multiple CoTs receiving tokens in a single forward pass? Gemini Diffusion-like models?). In this case, a lab which cares about alignment will begin research in order to understand scaling laws of the new architectures and the ways to prevent such architectures from scheming in unnoticeable ways (SAE? NLA? Reliance on low capabilities of CoTless skills?) and either invent schemes to reliably align the AIs with newfound capabilities or end up failing to notice that an AI began to scheme and took over.