DeepSWE v1.1 Results Across LLMs
Vendor-reported DeepSWE v1.1 results across showcased LLMs. Long-horizon agentic software engineering on original tasks from active open-source repos (Datacurve DeepSWE).
What this table contains
Long-horizon agentic software engineering on original tasks from active open-source repos (Datacurve DeepSWE).
26 models with results · 4 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.
Interactive results & effort preference → · Results over time → · Methodology
| Model | Headline score | Source/record date | Evidence |
|---|---|---|---|
| Gemini 4 ArgonGoogle | 77.9%max effort | 2026-09-30 | All 1 recorded result & sources77.9% · raw 77.9 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed, mini-swe agent harness; state of the art | Highest thinking settings per Google eval methodology. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ |
| Muse Spark 1.3Meta | 75.4%max effort | 2026-09-02 | All 2 recorded results & sources75.4% · raw 75.4 % Headline · max effort · Own vendor Source/record date: 2026-09-02 Meta published table as transcribed by ExplainX; max reasoning; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3 https://research.meta.ai/blog/introducing-muse-spark-1-375.4% · raw 75.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| GPT-6.1 SolOpenAI | 75.2%high effort | 2026-09-29 | All 2 recorded results & sources75.2% · raw 75.2 % Headline · high effort · Own vendor Source/record date: 2026-09-29 Announcement chart: matches Astra at ~1/5 cost; +6.4 over GPT-6 Sol best at lower effort and cost https://openai.com/index/introducing-gpt-6-1-sol/71.9% · raw 71.9 % Alternative · max effort · Own vendor Source/record date: 2026-09-29 Max-effort point on same chart; falls back from high-effort 75.2 https://openai.com/index/introducing-gpt-6-1-sol/ |
| Claude Opus 5.5Anthropic | 74.2%unknown effort | 2026-09-27 | All 2 recorded results & sources74.2% · raw 74.2 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/74.2% · raw 74.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| DeepSeek V4.1 FlashDeepSeek | 74.2%unknown effort | 2026-09-10 | All 3 recorded results & sources74.2% · raw 74.2 % Headline · unknown effort · Third-party Source/record date: 2026-09-10 VentureBeat quoting DeepSeek reports; also deepseekagent.io 74.2 https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-574.2% · raw 74.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/74.2% · raw 74.2 % Alternative · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| GPT-6 AstraOpenAI | 74.1%unknown effort | 2026-09-03 | All 2 recorded results & sources74.1% · raw 74.1 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/74.1% · raw 74.1 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Gemini 3.8 FlashGoogle | 73.7%unknown effort | 2026-09-03 | All 1 recorded result & sources73.7% · raw 73.7 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 DeepMind model card (prefer over OpenAI 73.8 if reconciling) https://deepmind.google/models/model-cards/gemini-3-8-flash |
| MiMo-V2.6-ProXiaomi | 71.9%unknown effort | 2026-09-21 | All 1 recorded result & sources71.9% · raw 71.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL |
| Claude Sonnet 5.5Anthropic | 71%max effort | 2026-09-28 | All 1 recorded result & sources71% · raw 71 % Headline · max effort · Own vendor Source/record date: 2026-09-28 DeepSWE v1.1 avg over 5 trials https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Grok 4.7xAI | 71%high effort | 2026-09-21 | All 1 recorded result & sources71% · raw 71 % Headline · high effort · Own vendor Source/record date: 2026-09-21 Official table marks high-effort with asterisk https://x.ai/news/grok-4-7 |
| Pareto 26.10 PreviewUnbiased | 69.9%unknown effort | 2026-10-01 | All 1 recorded result & sources69.9% · raw 69.9 % Headline · unknown effort · Own vendor Source/record date: 2026-10-01 Unbiased preliminary vendor result; mean task cost $0.24. Unbiased says results may change and denominator and cost methods need confirmation. https://unbiased.ai/blog/pareto-26-10-preview/ |
| GPT-5.6 TerraOpenAI | 69.6%unknown effort | 2026-07-09 | All 1 recorded result & sources69.6% · raw 69.6 % Headline · unknown effort · Own vendor Source/record date: 2026-07-09 https://openai.com/index/gpt-5-6 |
| Qwen3.8 Max (0902)Alibaba (Qwen) | 69.3%unknown effort | 2026-09-02 | All 3 recorded results & sources69.3% · raw 69.3 % Headline · unknown effort · Third-party Source/record date: 2026-09-02 Alibaba published comparison as reported by DataCamp https://www.datacamp.com/blog/qwen3-8-max69.3% · raw 69.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp; vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value. https://huggingface.co/nex-agi/Nex-N2.5-Pro56.6% · raw 56.6 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted. https://naive.ai/en/research/ |
| MiMo-V2.6-FlashXiaomi | 67.9%unknown effort | 2026-09-21 | All 1 recorded result & sources67.9% · raw 67.9 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL |
| Naive-N0.5-FlashNaiveAI | 67.8%unknown effort | 2026-09-27 | All 1 recorded result & sources67.8% · raw 67.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 Vendor chart figure; harness Claude Code 2.1.207, 1M context, temp 1.0, top-p 0.95. https://naive.ai/en/research/ |
| Kimi K3Moonshot AI | 67.5%max effort | 2026-07-16 | All 2 recorded results & sources67.5% · raw 67.5 % Headline · max effort · Own vendor Source/record date: 2026-07-16 DeepSWE v1.1 tasks; Kimi Code harness (mini-SWE-agent board reports 67.3) https://github.com/MoonshotAI/Kimi-K367.5% · raw 67.5 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| Claude Fable 5.1Anthropic | 67.4%unknown effort | 2026-09-03 | All 2 recorded results & sources67.4% · raw 67.4 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-03 As reported by OpenAI https://openai.com/index/gpt-6-astra/67.4% · raw 67.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| GLM-5.3Z.ai | 66.9%unknown effort | 2026-08-14 | All 3 recorded results & sources66.9% · raw 66.9 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 https://docs.z.ai/guides/llm/glm-5.366.9% · raw 66.9 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Z.ai GLM-5.3 docs (docs.z.ai); vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value. https://huggingface.co/nex-agi/Nex-N2.5-Pro66.9% · raw 66.9 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| GPT-6 LunaOpenAI | 66.6%max effort | 2026-09-22 | All 1 recorded result & sources66.6% · raw 66.6 % Headline · max effort · Own vendor Source/record date: 2026-09-22 max effort https://openai.com/index/introducing-gpt-6-sol-and-luna/ |
| Hy4 previewTencent | 64.3%unknown effort | 2026-08-28 | All 2 recorded results & sources64.3% · raw 64.3 % Headline · unknown effort · Own vendor Source/record date: 2026-08-28 Vendor chart transcription (DeepSWE) https://hy4.site/benchmarks/hy4-benchmarks64.3% · raw 64.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| GLM-5.3-FlashZ.ai | 63.4%unknown effort | 2026-08-26 | All 3 recorded results & sources63.4% · raw 63.4 % Headline · unknown effort · Own vendor Source/record date: 2026-08-26 https://z.ai/blog/glm-5.3-flash63.4% · raw 63.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/63.4% · raw 63.4 % Alternative · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| Ling 3.1 FlashinclusionAI | 59.7%unknown effort | 2026-09-30 | All 1 recorded result & sources59.7% · raw 59.7 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| Qwen3.8 FlashAlibaba (Qwen) | 58.7%xhigh effort | 2026-08-26 | All 1 recorded result & sources58.7% · raw 58.7 % Headline · xhigh effort · Own vendor Source/record date: 2026-08-26 DeepSWE 1.1 — QwenCloud latest-model page https://docs.qwencloud.com/developer-guides/getting-started/latest-model |
| Qwen3.8 Omni FlashAlibaba (Qwen) | 57.8%unknown effort | 2026-09-23 | All 1 recorded result & sources57.8% · raw 57.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-23 Qwen3.8-Omni-Flash column on Omni blog DeepSWE 1.1 table (Flash-Next listed 58.7 beside it) https://qwen.ai/blog?id=qwen3.8-omni-flash |
| Nex-N2.5-ProNex AGI | 55.8%unknown effort | 2026-09-08 | All 1 recorded result & sources55.8% · raw 55.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-08 NexAU harness https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| Qwen3.8 27BAlibaba (Qwen) | 42.2%unknown effort | 2026-08-14 | All 1 recorded result & sources42.2% · raw 42.2 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 DeepSWE 1.1; Claude Code harness https://huggingface.co/Qwen/Qwen3.8-27B |
| Kolibri-1Aleph Alpha | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | — | — | No recorded result |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | — | — | No recorded result |
Benchmark source references
- Datacurve DeepSWE (deepswe.datacurve.ai) · first_party
- OpenAI GPT-6 Astra system card · vendor_system_card
- OpenAI GPT-6 Sol & Luna announce · vendor_release
Use the fair-comparison guide before interpreting results from different configurations.