As per twitter, Ayeja Cotra thinks OpenAI happened not to implement the causal part of the grader, but that is so strange to me given the impact on scoring is so large. Seems rather important.
OpenAI should explain why they didn’t implement this. They very likely used AI assistance to implement this, so is it possible that e.g. Sol thought it was too unfair or something and just decided to skip it? Could have even been good faith, in that the model explicitly identifies this piece as potentially crazy-making to RL agents, and hence they did something saner which the human engineer approved.

I believe the 70B threshold is more Antra’s number than the consensus at Anima Labs.