We do not see significant accuracy uplift from self play optimization overall.
Isn’t this trivial due to the anthropic principle as accuracy uplift from self-play optimisation would FOOM via RSI?
Isn’t this trivial due to the anthropic principle as accuracy uplift from self-play optimisation would FOOM via RSI?