My problem with similar plans is the following. How likely is it that good-but-hacking AIs are more likely to become schemers than good AIs who are trained to resist the tendency to hack reward and to point out a rival’s attempts to hack reward? Or that good-but-reward-hacking AIs like the ones used in MATS 9 or by Greenblatt end up making it easier for Agent-4 to develop misaligned far-reaching goals or harder to catch Agent-4 because of things like accidental training on the CoT or Agent-3 being sloppy?
Wait can you explain the MATS 9 link? I couldn’t find any reference to reward hacking there.
For inoculation pretraining, I’m imagining we’d add data about good-but-reward-hacking AIs to pretraining, but still in post-training we’d try to train AIs to resist the tendency to hack reward and we’d try to train them to point out a rival’s attempts to hack reward, etc. The data about good-but-reward-hacking AIs is just there as a fallback in case we fail and accidentally train AIs to reward hack in post-training.
It’s possible that adding the data about good-but-reward-hacking AIs would nontrivially increase the probability that we fail, but I’m not sure. It seems probable to me that reward hacks are something that AIs will explore their way into, whether or not we add data about good-but-reward-hacking AIs to pretraining.
My problem with similar plans is the following. How likely is it that good-but-hacking AIs are more likely to become schemers than good AIs who are trained to resist the tendency to hack reward and to point out a rival’s attempts to hack reward? Or that good-but-reward-hacking AIs like the ones used in MATS 9 or by Greenblatt end up making it easier for Agent-4 to develop misaligned far-reaching goals or harder to catch Agent-4 because of things like accidental training on the CoT or Agent-3 being sloppy?
Wait can you explain the MATS 9 link? I couldn’t find any reference to reward hacking there.
For inoculation pretraining, I’m imagining we’d add data about good-but-reward-hacking AIs to pretraining, but still in post-training we’d try to train AIs to resist the tendency to hack reward and we’d try to train them to point out a rival’s attempts to hack reward, etc. The data about good-but-reward-hacking AIs is just there as a fallback in case we fail and accidentally train AIs to reward hack in post-training.
It’s possible that adding the data about good-but-reward-hacking AIs would nontrivially increase the probability that we fail, but I’m not sure. It seems probable to me that reward hacks are something that AIs will explore their way into, whether or not we add data about good-but-reward-hacking AIs to pretraining.