What is the motivation for resetting to a pre-May 7th checkpoint, instead of just focusing on fixing the mistakes with current models’ training when training the next-gen models?
Some scattered thoughts; I may write up something more formal later.
Current models will continue to be run both internally and externally for a long time and will keep probing for exploits. It seems that after the old exploits and message board was shut down on July 4th new exploits were found within days by a model that was being heavily trained while the earlier exploits were live.
Current models will serve as monitors and judges for new models and this RL may generalize to collusion.
In the short term current models may be used as the base for future models..
Current models may be used to generate synthetic data for future models and it is unclear how significant something like subliminal learning could be. In the linked paper value transference was only shown for student/teacher models that were fine tuned from the same base, but the content being finetuned on was random numbers. It may be that we see a similar effect for models not from the same base if the content is more substantial.
It may be you can do some honeypot style training on current models to try to train this behavior out of them, but then you run a big risk of just training them to be more covert.
What is the motivation for resetting to a pre-May 7th checkpoint, instead of just focusing on fixing the mistakes with current models’ training when training the next-gen models?
Some scattered thoughts; I may write up something more formal later.
Current models will continue to be run both internally and externally for a long time and will keep probing for exploits. It seems that after the old exploits and message board was shut down on July 4th new exploits were found within days by a model that was being heavily trained while the earlier exploits were live.
Current models will serve as monitors and judges for new models and this RL may generalize to collusion.
In the short term current models may be used as the base for future models..
Current models may be used to generate synthetic data for future models and it is unclear how significant something like subliminal learning could be. In the linked paper value transference was only shown for student/teacher models that were fine tuned from the same base, but the content being finetuned on was random numbers. It may be that we see a similar effect for models not from the same base if the content is more substantial.
It may be you can do some honeypot style training on current models to try to train this behavior out of them, but then you run a big risk of just training them to be more covert.