Nice! I suspect even post data-filtering, if the model is trained with RLVR it should might boost VEA & UVEA.
(during RLVR) My guess is that UVEA/VEA, grader-awareness arise mostly when the verifier fails to capture the full spectrum of possible solutions but focusses on a smaller subset, forcing the model to guess the right format/structure of the valid solutions which could get accepted, My current project is around this.
Nice! I suspect even post data-filtering, if the model is trained with RLVR it should might boost VEA & UVEA.
(during RLVR) My guess is that UVEA/VEA, grader-awareness arise mostly when the verifier fails to capture the full spectrum of possible solutions but focusses on a smaller subset, forcing the model to guess the right format/structure of the valid solutions which could get accepted, My current project is around this.