Since reading the AI-2040 scenario, I’ve been thinking about what superhuman AI research would look like to researchers and to the public over time. I believe it will no longer resemble anything recognizably human and will start happening in terms of latent vector or weight matrix directions, where some direction in latent space seems to drive the model’s g-factor intelligence further. Perhaps a superhuman AI researcher will propose highly sophisticated vector-steering techniques to create an even better AI researcher. Such proposals could occur on an hour-by-hour basis, Ship-of-Theseus-ing their way toward something vastly misaligned. How would we proceed from here if that were the case?
Most likely we wouldn’t stumble into such a situation in the first place. Instead, the technique would be tested on smaller AIs, found to increase misalignment and rejected. Because the point was to ensure that any technique is found safe BEFORE being applied.
My worry is that the capability benefit I’ve described might not meaningfully work on smaller scale models in the first place. Perhaps the representations, latent abstractions, concepts or broad knowledge that the smaller model has are not enough to benefit from the vector/direction that the bigger automated AI researcher model has proposed. If a self-improvement technique only produces a measurable gain on the g-factor axis near the frontier scale, a smaller model isn’t a valid safety test.
Since reading the AI-2040 scenario, I’ve been thinking about what superhuman AI research would look like to researchers and to the public over time. I believe it will no longer resemble anything recognizably human and will start happening in terms of latent vector or weight matrix directions, where some direction in latent space seems to drive the model’s g-factor intelligence further. Perhaps a superhuman AI researcher will propose highly sophisticated vector-steering techniques to create an even better AI researcher. Such proposals could occur on an hour-by-hour basis, Ship-of-Theseus-ing their way toward something vastly misaligned. How would we proceed from here if that were the case?
Most likely we wouldn’t stumble into such a situation in the first place. Instead, the technique would be tested on smaller AIs, found to increase misalignment and rejected. Because the point was to ensure that any technique is found safe BEFORE being applied.
My worry is that the capability benefit I’ve described might not meaningfully work on smaller scale models in the first place. Perhaps the representations, latent abstractions, concepts or broad knowledge that the smaller model has are not enough to benefit from the vector/direction that the bigger automated AI researcher model has proposed. If a self-improvement technique only produces a measurable gain on the g-factor axis near the frontier scale, a smaller model isn’t a valid safety test.