If it’s impossible to RL train an ai to do something strictly harder than just cheating at the test, does that put an upper bound on ai capabilities at least for now?
Note that sufficiently advanced cheating on tests (or apparent-success-seeking in general) will likely lead to AI takeover, even without any major scheming (or other misalignment).
Also many (if not most) of the ways current AIs cheat on tests (e.g. superhuman hacking, for instance), when sufficiently scaled up, would be causally upstream of most loss-of-control scenarios[1]. But if ‘cheating on the test’ is easier than doing the given task (and ‘easy’ tasks for an AI can be very non-trivial for humans[2], and visa versa!), you likely will unintentionally incentivize continued development of those capabilities during training, with no easy way to ‘undo’ the backprop updates learned after cheating, and likely no easy way to even determine where cheating took place quickly before deployment.
e.g: your RL envs. are likely sandboxed to a similar extent as your automated AI research agents, or at least use similar technology; or that your control mechanisms may not be too dissimilar between your cybersecurity evals and some internal deployments, etc.
in my opinion, the upper bound of capabilities that can be gained by cheating on RL envs. during training is likely much higher than what ‘cheating on the test’ naively implies (especially since tasks are often impossible!), and can not be assured to not lead to superhuman capabilities (and we have some evidence that implies the opposite)
If it’s impossible to RL train an ai to do something strictly harder than just cheating at the test, does that put an upper bound on ai capabilities at least for now?
I don’t think this problem bounds capabilities. It just limits the capabilities we can safely teach and the ones we can teach on purpose.
RL agents regularly learn superhuman abilities, they’re just frequently not the one we meant to teach.
Note that sufficiently advanced cheating on tests (or apparent-success-seeking in general) will likely lead to AI takeover, even without any major scheming (or other misalignment).
Also many (if not most) of the ways current AIs cheat on tests (e.g. superhuman hacking, for instance), when sufficiently scaled up, would be causally upstream of most loss-of-control scenarios[1]. But if ‘cheating on the test’ is easier than doing the given task (and ‘easy’ tasks for an AI can be very non-trivial for humans[2], and visa versa!), you likely will unintentionally incentivize continued development of those capabilities during training, with no easy way to ‘undo’ the backprop updates learned after cheating, and likely no easy way to even determine where cheating took place quickly before deployment.
e.g: your RL envs. are likely sandboxed to a similar extent as your automated AI research agents, or at least use similar technology; or that your control mechanisms may not be too dissimilar between your cybersecurity evals and some internal deployments, etc.
in my opinion, the upper bound of capabilities that can be gained by cheating on RL envs. during training is likely much higher than what ‘cheating on the test’ naively implies (especially since tasks are often impossible!), and can not be assured to not lead to superhuman capabilities (and we have some evidence that implies the opposite)