Thank you, I also wanted to write something similar because of a similarly linear “scaling law” of AECI over time. As far as I understand, the current architecture doesn’t allow anyone to do the RSI for reasons similar to why humans cannot do RSI on themselves: capabilities are increased linearly over epochs or logarithmically over lived experience, which in the AIs’ case is proportional to compute spent.
Under this interpretation, eventually someone will understand it and come up with alternate architectures with their scaling laws (neuralese trained from scratch? Multiple CoTs receiving tokens in a single forward pass? Gemini Diffusion-like models?). In this case, a lab which cares about alignment will begin research in order to understand scaling laws of the new architectures and the ways to prevent such architectures from scheming in unnoticeable ways (SAE? NLA? Reliance on low capabilities of CoTless skills?) and either invent schemes to reliably align the AIs with newfound capabilities or end up failing to notice that an AI began to scheme and took over.
Thank you, I also wanted to write something similar because of a similarly linear “scaling law” of AECI over time. As far as I understand, the current architecture doesn’t allow anyone to do the RSI for reasons similar to why humans cannot do RSI on themselves: capabilities are increased linearly over epochs or logarithmically over lived experience, which in the AIs’ case is proportional to compute spent.
Under this interpretation, eventually someone will understand it and come up with alternate architectures with their scaling laws (neuralese trained from scratch? Multiple CoTs receiving tokens in a single forward pass? Gemini Diffusion-like models?). In this case, a lab which cares about alignment will begin research in order to understand scaling laws of the new architectures and the ways to prevent such architectures from scheming in unnoticeable ways (SAE? NLA? Reliance on low capabilities of CoTless skills?) and either invent schemes to reliably align the AIs with newfound capabilities or end up failing to notice that an AI began to scheme and took over.