Most likely we wouldn’t stumble into such a situation in the first place. Instead, the technique would be tested on smaller AIs, found to increase misalignment and rejected. Because the point was to ensure that any technique is found safe BEFORE being applied.
My worry is that the capability benefit I’ve described might not meaningfully work on smaller scale models in the first place. Perhaps the representations, latent abstractions, concepts or broad knowledge that the smaller model has are not enough to benefit from the vector/direction that the bigger automated AI researcher model has proposed. If a self-improvement technique only produces a measurable gain on the g-factor axis near the frontier scale, a smaller model isn’t a valid safety test.
Most likely we wouldn’t stumble into such a situation in the first place. Instead, the technique would be tested on smaller AIs, found to increase misalignment and rejected. Because the point was to ensure that any technique is found safe BEFORE being applied.
My worry is that the capability benefit I’ve described might not meaningfully work on smaller scale models in the first place. Perhaps the representations, latent abstractions, concepts or broad knowledge that the smaller model has are not enough to benefit from the vector/direction that the bigger automated AI researcher model has proposed. If a self-improvement technique only produces a measurable gain on the g-factor axis near the frontier scale, a smaller model isn’t a valid safety test.