LLMs are often graded by an LLM judge that takes in the entire context of the rollout… naturally if you were trained for millions of RL episodes where you were rewarded on your entire context your prose generation mechanism would develop a prior that the counterparty has seen everything.
Yes! I wonder if models would be more articulate after RL graded by a judge who sees a random fragment instead of the whole context. It would force them to write in a way that “makes sense locally”.
Yes! I wonder if models would be more articulate after RL graded by a judge who sees a random fragment instead of the whole context. It would force them to write in a way that “makes sense locally”.