SWE-bench Pro Results Across LLMs
Vendor-reported SWE-bench Pro results across showcased LLMs. Long-horizon software engineering on real GitHub issues beyond the classic SWE-bench Verified set.
What this table contains
Long-horizon software engineering on real GitHub issues beyond the classic SWE-bench Verified set.
12 models with results · 18 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.
Interactive results & effort preference → · Results over time → · Methodology
| Model | Headline score | Source/record date | Evidence |
|---|---|---|---|
| Claude Opus 5.5Anthropic | 89.9%unknown effort | 2026-09-27 | All 1 recorded result & sources89.9% · raw 89.9 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI; card cites the Opus 5.5 system card. https://naive.ai/en/research/ |
| Claude Sonnet 5.5Anthropic | 81.3%max effort | 2026-09-28 | All 1 recorded result & sources81.3% · raw 81.3 % Headline · max effort · Own vendor Source/record date: 2026-09-28 SWE-bench Pro avg over 5 trials; standard config (adaptive thinking max) https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Naive-N0.5-FlashNaiveAI | 73.6%unknown effort | 2026-09-27 | All 1 recorded result & sources73.6% · raw 73.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 PDF chart text and figure pixels read 73.6; the release page alt text says 68.8 and looks stale. Chart value used; harness Claude Code 2.1.207. https://naive.ai/en/research/ |
| Qwen3.8 Max (0902)Alibaba (Qwen) | 67.7%unknown effort | 2026-09-02 | All 3 recorded results & sources67.7% · raw 67.7 % Headline · unknown effort · Third-party Source/record date: 2026-09-02 SWE-Pro / SWE-bench Pro https://www.datacamp.com/blog/qwen3-8-max67.7% · raw 67.7 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp; vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value. https://huggingface.co/nex-agi/Nex-N2.5-Pro67.7% · raw 67.7 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted. https://naive.ai/en/research/ |
| Hy4 previewTencent | 65.7%unknown effort | 2026-08-28 | All 2 recorded results & sources65.7% · raw 65.7 % Headline · unknown effort · Own vendor Source/record date: 2026-08-28 Vendor chart transcription https://hy4.site/benchmarks/hy4-benchmarks65.7% · raw 65.7 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| GLM-5.3Z.ai | 64.6%unknown effort | 2026-09-08 | All 1 recorded result & sources64.6% · raw 64.6 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| GPT-5.6 TerraOpenAI | 63.4%unknown effort | 2026-07-09 | All 1 recorded result & sources63.4% · raw 63.4 % Headline · unknown effort · Own vendor Source/record date: 2026-07-09 https://openai.com/index/gpt-5-6 |
| Qwen3.8 Omni FlashAlibaba (Qwen) | 63.3%xhigh effort | 2026-09-18 | All 1 recorded result & sources63.3% · raw 63.3 % Headline · xhigh effort · Third-party Source/record date: 2026-09-18 SWE-bench Pro — launch secondary transcription https://www.cometapi.com/what-is-qwen3-8-omni-flash-and-how-to-access/ |
| Qwen3.8 FlashAlibaba (Qwen) | 62.5%xhigh effort | 2026-08-26 | All 1 recorded result & sources62.5% · raw 62.5 % Headline · xhigh effort · Own vendor Source/record date: 2026-08-26 SWE-bench Pro — QwenCloud latest-model page https://docs.qwencloud.com/developer-guides/getting-started/latest-model |
| Qwen3.8 27BAlibaba (Qwen) | 61.7%unknown effort | 2026-08-14 | All 1 recorded result & sources61.7% · raw 61.7 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 SWE-bench Pro; Claude Code harness (temp=1.0, top_p=0.95, 256K) except Opus uses official https://huggingface.co/Qwen/Qwen3.8-27B |
| Nex-N2.5-ProNex AGI | 61.2%unknown effort | 2026-09-08 | All 1 recorded result & sources61.2% · raw 61.2 % Headline · unknown effort · Own vendor Source/record date: 2026-09-08 NexAU harness https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | 44.6%unknown effort | 2026-09-21 | All 9 recorded results & sources44.6% · raw 44.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 Released SFT checkpoint, avg@3; MiMo-V2.6 tech report Table 6 (https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf) and HF card Evaluation table. Also plotted (SFT line) in radar figure on https://mimo.mi.com/docs/en-US/news/latest/v2-6. Eval effort not specified (thinking model). HF card label 'SWE Pro'. https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B45.2% · raw 45.2 % Alternative · unknown effort · Own vendor Source/record date: 2026-09-21 Harness: mini-harness1 (training harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf46.6% · raw 46.6 % Alternative · unknown effort · Own vendor Source/record date: 2026-09-21 Harness: mini-harness2 (training harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf45.1% · raw 45.1 % Alternative · unknown effort · Own vendor Source/record date: 2026-09-21 Harness: mini-harness3 (training harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf45.5% · raw 45.5 % Alternative · unknown effort · Own vendor Source/record date: 2026-09-21 Harness: mini-harness4 (training harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf40.3% · raw 40.3 % Alternative · unknown effort · Own vendor Source/record date: 2026-09-21 Harness: codex (held-out harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf42.2% · raw 42.2 % Alternative · unknown effort · Own vendor Source/record date: 2026-09-21 Harness: claude code (held-out harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf45.6% · raw 45.6 % Alternative · unknown effort · Own vendor Source/record date: 2026-09-21 Harness: mini-swe-agent (held-out harness). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf44.4% · raw 44.4 % Alternative · unknown effort · Own vendor Source/record date: 2026-09-21 Harness: Mean (7 harnesses) (unweighted mean). SFT checkpoint, tech report Table 7 (coding performance across agent harnesses; PDF p.36). Not the headline (headline = Table 6 / HF card). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf |
| Claude Fable 5.1Anthropic | — | — | No recorded result |
| DeepSeek V4.1 FlashDeepSeek | — | — | No recorded result |
| Gemini 3.8 FlashGoogle | — | — | No recorded result |
| Gemini 4 ArgonGoogle | — | — | No recorded result |
| GLM-5.3-FlashZ.ai | — | — | No recorded result |
| GPT-6 AstraOpenAI | — | — | No recorded result |
| GPT-6 LunaOpenAI | — | — | No recorded result |
| GPT-6.1 SolOpenAI | — | — | No recorded result |
| Grok 4.7xAI | — | — | No recorded result |
| Kimi K3Moonshot AI | — | — | No recorded result |
| Kolibri-1Aleph Alpha | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | — | — | No recorded result |
| Ling 3.1 FlashinclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | — | — | No recorded result |
| MiMo-V2.6-FlashXiaomi | — | — | No recorded result |
| MiMo-V2.6-ProXiaomi | — | — | No recorded result |
| Muse Spark 1.3Meta | — | — | No recorded result |
| Pareto 26.10 PreviewUnbiased | — | — | No recorded result |
Benchmark source references
- SWE-bench Pro · first_party
- Scale SWE-bench Pro board · leaderboard
- Vendor announce tables · vendor_release
Use the fair-comparison guide before interpreting results from different configurations.