Examenos

SWE-bench Pro Results Across LLMs

Vendor-reported SWE-bench Pro results across showcased LLMs. Long-horizon software engineering on real GitHub issues beyond the classic SWE-bench Verified set.

Official / vendor Higher is better · unit: %

What this table contains

Long-horizon software engineering on real GitHub issues beyond the classic SWE-bench Verified set.

12 models with results · 18 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.

Interactive results & effort preference → · Results over time → · Methodology

SWE-bench Pro: vendor-reported results only
ModelHeadline scoreSource/record dateEvidence
Claude Opus 5.5Anthropic89.9%unknown effort2026-09-27
All 1 recorded result & sources

89.9% · raw 89.9 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI; card cites the Opus 5.5 system card.

https://naive.ai/en/research/
Claude Sonnet 5.5Anthropic81.3%max effort2026-09-28
All 1 recorded result & sources

81.3% · raw 81.3 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

SWE-bench Pro avg over 5 trials; standard config (adaptive thinking max)

https://www.anthropic.com/claude-sonnet-5-5-system-card
Naive-N0.5-FlashNaiveAI73.6%unknown effort2026-09-27
All 1 recorded result & sources

73.6% · raw 73.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-27

PDF chart text and figure pixels read 73.6; the release page alt text says 68.8 and looks stale. Chart value used; harness Claude Code 2.1.207.

https://naive.ai/en/research/
Qwen3.8 Max (0902)Alibaba (Qwen)67.7%unknown effort2026-09-02
All 3 recorded results & sources

67.7% · raw 67.7 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

SWE-Pro / SWE-bench Pro

https://www.datacamp.com/blog/qwen3-8-max

67.7% · raw 67.7 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp; vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value.

https://huggingface.co/nex-agi/Nex-N2.5-Pro

67.7% · raw 67.7 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted.

https://naive.ai/en/research/
Hy4 previewTencent65.7%unknown effort2026-08-28
All 2 recorded results & sources

65.7% · raw 65.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-28

Vendor chart transcription

https://hy4.site/benchmarks/hy4-benchmarks

65.7% · raw 65.7 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
GLM-5.3Z.ai64.6%unknown effort2026-09-08
All 1 recorded result & sources

64.6% · raw 64.6 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
GPT-5.6 TerraOpenAI63.4%unknown effort2026-07-09
All 1 recorded result & sources

63.4% · raw 63.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-07-09

https://openai.com/index/gpt-5-6
Qwen3.8 Omni FlashAlibaba (Qwen)63.3%xhigh effort2026-09-18
All 1 recorded result & sources

63.3% · raw 63.3 %

Headline · xhigh effort · Third-party

Source/record date: 2026-09-18

SWE-bench Pro — launch secondary transcription

https://www.cometapi.com/what-is-qwen3-8-omni-flash-and-how-to-access/
Qwen3.8 FlashAlibaba (Qwen)62.5%xhigh effort2026-08-26
All 1 recorded result & sources

62.5% · raw 62.5 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

SWE-bench Pro — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
Qwen3.8 27BAlibaba (Qwen)61.7%unknown effort2026-08-14
All 1 recorded result & sources

61.7% · raw 61.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

SWE-bench Pro; Claude Code harness (temp=1.0, top_p=0.95, 256K) except Opus uses official

https://huggingface.co/Qwen/Qwen3.8-27B
Nex-N2.5-ProNex AGI61.2%unknown effort2026-09-08
All 1 recorded result & sources

61.2% · raw 61.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-08

NexAU harness

https://huggingface.co/nex-agi/Nex-N2.5-Pro
MiMo-V2.6-Distill-Qwen-9BXiaomi44.6%unknown effort2026-09-21
All 9 recorded results & sources

44.6% · raw 44.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

Released SFT checkpoint, avg@3; MiMo-V2.6 tech report Table 6 (https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf) and HF card Evaluation table. Also plotted (SFT line) in radar figure on https://mimo.mi.com/docs/en-US/news/latest/v2-6. Eval effort not specified (thinking model). HF card label 'SWE Pro'.

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B

45.2% · raw 45.2 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-09-21

Harness: mini-harness1 (training harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf

46.6% · raw 46.6 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-09-21

Harness: mini-harness2 (training harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf

45.1% · raw 45.1 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-09-21

Harness: mini-harness3 (training harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf

45.5% · raw 45.5 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-09-21

Harness: mini-harness4 (training harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf

40.3% · raw 40.3 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-09-21

Harness: codex (held-out harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf

42.2% · raw 42.2 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-09-21

Harness: claude code (held-out harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf

45.6% · raw 45.6 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-09-21

Harness: mini-swe-agent (held-out harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf

44.4% · raw 44.4 %

Alternative · unknown effort · Own vendor

Source/record date: 2026-09-21

Harness: Mean (7 harnesses) (unweighted mean). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf
Claude Fable 5.1Anthropic——No recorded result
DeepSeek V4.1 FlashDeepSeek——No recorded result
Gemini 3.8 FlashGoogle——No recorded result
Gemini 4 ArgonGoogle——No recorded result
GLM-5.3-FlashZ.ai——No recorded result
GPT-6 AstraOpenAI——No recorded result
GPT-6 LunaOpenAI——No recorded result
GPT-6.1 SolOpenAI——No recorded result
Grok 4.7xAI——No recorded result
Kimi K3Moonshot AI——No recorded result
Kolibri-1Aleph Alpha——No recorded result
Ling 3.0 Flash VLinclusionAI——No recorded result
Ling 3.1 FlashinclusionAI——No recorded result
Mercury 2.5Inception——No recorded result
MiMo-V2.6-FlashXiaomi——No recorded result
MiMo-V2.6-ProXiaomi——No recorded result
Muse Spark 1.3Meta——No recorded result
Pareto 26.10 PreviewUnbiased——No recorded result

Benchmark source references

Use the fair-comparison guide before interpreting results from different configurations.