LLMs are often graded by an LLM judge that takes in the entire context of the rollout… naturally if you were trained for millions of RL episodes where you were rewarded on your entire context your prose generation mechanism would develop a prior that the counterparty has seen everything.
Yes! I wonder if models would be more articulate after RL graded by a judge who sees a random fragment instead of the whole context. It would force them to write in a way that “makes sense locally”.
The MAIS (Math for AI Safety) repo is my new effort to draw mathematicians into AI safety.
It started as a survey paper Math for AI Safety: An Invitation for Mathematicians. During the writing of that paper, Claude and GPT generated 100+ pages of open problems organized in eight research agendas. I’m releasing that material as a public hub for mathematicians to collaborate on AI safety problems.
The repo just went live today, and it’s very much a work in progress! I’d appreciate any feedback on how to make it better.