I suspect that RL with a more thorough reward could help. When a real-world human tries to present slop, slop takes a rather long time to be revealed to be slop and punished accordingly. What the LLM has is a context window which is graded according to the reward model. What if we ask the reward model to be a just-as-capable LLM and to use the code in the project, then to report all the failures to the instance which generated the code and have the instance finetuned on the user’s rant?
I suspect that RL with a more thorough reward could help. When a real-world human tries to present slop, slop takes a rather long time to be revealed to be slop and punished accordingly. What the LLM has is a context window which is graded according to the reward model. What if we ask the reward model to be a just-as-capable LLM and to use the code in the project, then to report all the failures to the instance which generated the code and have the instance finetuned on the user’s rant?