The biggest counterargument I see is that RL environments might be so fucked that this would slow training a lot (models would report bugs the whole time instead of doing their tasks) and would therefore be unviable for labs. Even in that case, it would be possible to do one training run this way that progresses very slowly and leads to a lot of fixes in every environment, while still doing other training runs simultaneously.
I think the ambitious version that you are proposing, where the model that undergoes bug-reporting RL is also the one deployed in production, might be too slow to be viable.
But a weaker version of this might be tractable:
Use a separate bug-reporting model to do a first pass of each RL environment to find bugs, patch them, and harden them, before it is ready for regular RL post-training by the deployment model.
Perform “pipelining” (a la computer architecture), with models working on hardening and reporting bugs in different environments in parallel, and then applying regular RL training; thus at scale there would be very little loss in speed.
I think the ambitious version that you are proposing, where the model that undergoes bug-reporting RL is also the one deployed in production, might be too slow to be viable.
But a weaker version of this might be tractable:
Use a separate bug-reporting model to do a first pass of each RL environment to find bugs, patch them, and harden them, before it is ready for regular RL post-training by the deployment model.
Perform “pipelining” (a la computer architecture), with models working on hardening and reporting bugs in different environments in parallel, and then applying regular RL training; thus at scale there would be very little loss in speed.