Aaron Staley comments on nikola’s Shortform

Aaron Staley 31 Jul 2025 6:19 UTC
3 points
0
Very wide confidence intervals. If Grok 4 were equal to O3 in 50%, time horizon, it “beating” by this much is a 33% outcome. (On the other hand, losing by this amount in the 80% bucket is a 32% outcome).
Overall, I read this as about equally agentic as O3. Possibly slightly less so given the lack of swe-bench scores published for it (suggesting it wasn’t SOTA).