I think if no one knows the correct answer, it’s a different problem. It seems like a bad idea to use unsolved conjectures during online training (rather than first having the AI solve the conjecture, verifying the solution, then adding it to the training data). I would expect this to lead to a lot of reward hacking.
That said, I agree that giving the model a way to reward hack in a monitorable way is a good idea if you do this (please don’t do this).
I think if no one knows the correct answer, it’s a different problem. It seems like a bad idea to use unsolved conjectures during online training (rather than first having the AI solve the conjecture, verifying the solution, then adding it to the training data). I would expect this to lead to a lot of reward hacking.
That said, I agree that giving the model a way to reward hack in a monitorable way is a good idea if you do this (please don’t do this).