The only benchmark I trust is deepswe. Better methodology, actually matches the feel in SWE tasks, tries to avoid benchmark poisoning, prevents models from cheating, allows them to write tests etc. GlM 5.1 is not that high up in there. They’re yet to test GLM 5.2, but I would guess it’s somewhere near 5.4 mini or optimistically between opus 4.6 high and 4.7 on max.
edit: I looked further into it GLM 5.2′s launch page . Has it listed at 46.2. Which puts it between 4.6 high and 4.7 max.
GLM 5.2 does a lot better than GLM 5.1 at Deep SWE according to GLM’s website z.ai/blog/glm-5.2. I admit their 46.2 score falls further behind GPT 5.5 (and still below Claude Opus). But it still beats Gemini, Claude Sonnet, Grok, and the other models.
Somehow Deep SWE’s website doesn’t include GLM 5.2 yet, so I’m not sure if the 46.2 score is official.
Edit: thank you for adding their graphs, it’s very helpful! One potentially misleading by them (not you) is that Claude’s 58.0 score in DeepSWE is by Opus 4.8 not Fable 5.
I just stumbled upon that right now from their github, included it in the original message, I think it might fare similarly in official bench, it’s not mythos class yet.
But then again, it has less than a trillion parameters, while Mythos has 10 trillion. It might become more capable if they merely scale it up. Though “merely scaling” is obviously easier said than done!
Deep SWE will not include GLM 5.2 before they include Composer 2.5. Even for SWE bench Pro, OpenAI self-reports their own numbers for Codex 5.5 on the public set instead of paying Scale for the private set. The only way to be sure of any of these bench numbers is to run the benches yourself and verify instead of trusting official numbers.
How did you come to trust deepswe? What do you mean by better methodology? You say that it matches the feel in SWE tasks, so I’ll take it that you’re going with vibes because the rankings are reflective of your experience with these models. I had the same questions as you when deepswe made the media rounds a few weeks ago. So I just asked Codex with a short prompt and it took a few minutes for it to tell me how sound their methodology was.
Try it yourself, ask your agent:
I’m considering using https://deepswe.datacurve.ai/ as a benchmark to compare LLMs for a coding-agent project. Before I trust the leaderboard, audit it skeptically: any methodological problems, missing disclosures, inconsistencies in what they publish, or things a careful reader should flag?
Use their public artifacts at /artifacts/*.json (summary, leaderboard, heatmap, tasks, trials) and the per-trial JSONs at /artifacts/trials/{trial_name}.json. Cite specific URLs and JSON fields for any issue you find.
The only benchmark I trust is deepswe. Better methodology, actually matches the feel in SWE tasks, tries to avoid benchmark poisoning, prevents models from cheating, allows them to write tests etc. GlM 5.1 is not that high up in there. They’re yet to test GLM 5.2, but I would guess it’s somewhere near 5.4 mini or optimistically between opus 4.6 high and 4.7 on max.
edit: I looked further into it GLM 5.2′s launch page . Has it listed at 46.2. Which puts it between 4.6 high and 4.7 max.
GLM 5.2 does a lot better than GLM 5.1 at Deep SWE according to GLM’s website z.ai/blog/glm-5.2. I admit their 46.2 score falls further behind GPT 5.5 (and still below Claude Opus). But it still beats Gemini, Claude Sonnet, Grok, and the other models.
Somehow Deep SWE’s website doesn’t include GLM 5.2 yet, so I’m not sure if the 46.2 score is official.
Edit: thank you for adding their graphs, it’s very helpful! One potentially misleading by them (not you) is that Claude’s 58.0 score in DeepSWE is by Opus 4.8 not Fable 5.
I just stumbled upon that right now from their github, included it in the original message, I think it might fare similarly in official bench, it’s not mythos class yet.
I agree it’s not Mythos class.
But then again, it has less than a trillion parameters, while Mythos has 10 trillion. It might become more capable if they merely scale it up. Though “merely scaling” is obviously easier said than done!
Deep SWE will not include GLM 5.2 before they include Composer 2.5. Even for SWE bench Pro, OpenAI self-reports their own numbers for Codex 5.5 on the public set instead of paying Scale for the private set. The only way to be sure of any of these bench numbers is to run the benches yourself and verify instead of trusting official numbers.
How did you come to trust deepswe? What do you mean by better methodology? You say that it matches the feel in SWE tasks, so I’ll take it that you’re going with vibes because the rankings are reflective of your experience with these models. I had the same questions as you when deepswe made the media rounds a few weeks ago. So I just asked Codex with a short prompt and it took a few minutes for it to tell me how sound their methodology was.
Try it yourself, ask your agent: