That’s GRPO, a RL algorithm, not the reward. If you used other RL algorithms like PPO, I assume you would get similar results. (GRPO can be justified as approximately unbiased relative to REINFORCE)
That’s GRPO, a RL algorithm, not the reward. If you used other RL algorithms like PPO, I assume you would get similar results. (GRPO can be justified as approximately unbiased relative to REINFORCE)