Examenos

DeepSWE v1.1 Results Across LLMs

Vendor-reported DeepSWE v1.1 results across showcased LLMs. Long-horizon agentic software engineering on original tasks from active open-source repos (Datacurve DeepSWE).

Official / vendor Higher is better · unit: %

What this table contains

Long-horizon agentic software engineering on original tasks from active open-source repos (Datacurve DeepSWE).

26 models with results · 4 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.

Interactive results & effort preference → · Results over time → · Methodology

DeepSWE v1.1: vendor-reported results only
ModelHeadline scoreSource/record dateEvidence
Gemini 4 ArgonGoogle77.9%max effort2026-09-30
All 1 recorded result & sources

77.9% · raw 77.9 %

Headline · max effort · Own vendor

Source/record date: 2026-09-30

Self-computed, mini-swe agent harness; state of the art | Highest thinking settings per Google eval methodology.

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
Muse Spark 1.3Meta75.4%max effort2026-09-02
All 2 recorded results & sources

75.4% · raw 75.4 %

Headline · max effort · Own vendor

Source/record date: 2026-09-02

Meta published table as transcribed by ExplainX; max reasoning; primary Meta blog (chart-heavy): https://research.meta.ai/blog/introducing-muse-spark-1-3

https://research.meta.ai/blog/introducing-muse-spark-1-3

75.4% · raw 75.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
GPT-6.1 SolOpenAI75.2%high effort2026-09-29
All 2 recorded results & sources

75.2% · raw 75.2 %

Headline · high effort · Own vendor

Source/record date: 2026-09-29

Announcement chart: matches Astra at ~1/5 cost; +6.4 over GPT-6 Sol best at lower effort and cost

https://openai.com/index/introducing-gpt-6-1-sol/

71.9% · raw 71.9 %

Alternative · max effort · Own vendor

Source/record date: 2026-09-29

Max-effort point on same chart; falls back from high-effort 75.2

https://openai.com/index/introducing-gpt-6-1-sol/
Claude Opus 5.5Anthropic74.2%unknown effort2026-09-27
All 2 recorded results & sources

74.2% · raw 74.2 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

74.2% · raw 74.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
DeepSeek V4.1 FlashDeepSeek74.2%unknown effort2026-09-10
All 3 recorded results & sources

74.2% · raw 74.2 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-10

VentureBeat quoting DeepSeek reports; also deepseekagent.io 74.2

https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5

74.2% · raw 74.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

74.2% · raw 74.2 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
GPT-6 AstraOpenAI74.1%unknown effort2026-09-03
All 2 recorded results & sources

74.1% · raw 74.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/

74.1% · raw 74.1 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Gemini 3.8 FlashGoogle73.7%unknown effort2026-09-03
All 1 recorded result & sources

73.7% · raw 73.7 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

DeepMind model card (prefer over OpenAI 73.8 if reconciling)

https://deepmind.google/models/model-cards/gemini-3-8-flash
MiMo-V2.6-ProXiaomi71.9%unknown effort2026-09-21
All 1 recorded result & sources

71.9% · raw 71.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
Claude Sonnet 5.5Anthropic71%max effort2026-09-28
All 1 recorded result & sources

71% · raw 71 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

DeepSWE v1.1 avg over 5 trials

https://www.anthropic.com/claude-sonnet-5-5-system-card
Grok 4.7xAI71%high effort2026-09-21
All 1 recorded result & sources

71% · raw 71 %

Headline · high effort · Own vendor

Source/record date: 2026-09-21

Official table marks high-effort with asterisk

https://x.ai/news/grok-4-7
Pareto 26.10 PreviewUnbiased69.9%unknown effort2026-10-01
All 1 recorded result & sources

69.9% · raw 69.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-10-01

Unbiased preliminary vendor result; mean task cost $0.24. Unbiased says results may change and denominator and cost methods need confirmation.

https://unbiased.ai/blog/pareto-26-10-preview/
GPT-5.6 TerraOpenAI69.6%unknown effort2026-07-09
All 1 recorded result & sources

69.6% · raw 69.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-07-09

https://openai.com/index/gpt-5-6
Qwen3.8 Max (0902)Alibaba (Qwen)69.3%unknown effort2026-09-02
All 3 recorded results & sources

69.3% · raw 69.3 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-02

Alibaba published comparison as reported by DataCamp

https://www.datacamp.com/blog/qwen3-8-max

69.3% · raw 69.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Alibaba Qwen3.8-Max-0902 comparison table via DataCamp; vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value.

https://huggingface.co/nex-agi/Nex-N2.5-Pro

56.6% · raw 56.6 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI. Chart labels Qwen-3.8-Max; attached to the 0902 snapshot with version ambiguity noted.

https://naive.ai/en/research/
MiMo-V2.6-FlashXiaomi67.9%unknown effort2026-09-21
All 1 recorded result & sources

67.9% · raw 67.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
Naive-N0.5-FlashNaiveAI67.8%unknown effort2026-09-27
All 1 recorded result & sources

67.8% · raw 67.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-27

Vendor chart figure; harness Claude Code 2.1.207, 1M context, temp 1.0, top-p 0.95.

https://naive.ai/en/research/
Kimi K3Moonshot AI67.5%max effort2026-07-16
All 2 recorded results & sources

67.5% · raw 67.5 %

Headline · max effort · Own vendor

Source/record date: 2026-07-16

DeepSWE v1.1 tasks; Kimi Code harness (mini-SWE-agent board reports 67.3)

https://github.com/MoonshotAI/Kimi-K3

67.5% · raw 67.5 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
Claude Fable 5.1Anthropic67.4%unknown effort2026-09-03
All 2 recorded results & sources

67.4% · raw 67.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/

67.4% · raw 67.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
GLM-5.3Z.ai66.9%unknown effort2026-08-14
All 3 recorded results & sources

66.9% · raw 66.9 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

https://docs.z.ai/guides/llm/glm-5.3

66.9% · raw 66.9 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table | Demoted 2026-09-25: duplicate of headline from Z.ai GLM-5.3 docs (docs.z.ai); vendor official preferred over Nex-N2.5-Pro peer comparison table (rule a); same value.

https://huggingface.co/nex-agi/Nex-N2.5-Pro

66.9% · raw 66.9 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
GPT-6 LunaOpenAI66.6%max effort2026-09-22
All 1 recorded result & sources

66.6% · raw 66.6 %

Headline · max effort · Own vendor

Source/record date: 2026-09-22

max effort

https://openai.com/index/introducing-gpt-6-sol-and-luna/
Hy4 previewTencent64.3%unknown effort2026-08-28
All 2 recorded results & sources

64.3% · raw 64.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-28

Vendor chart transcription (DeepSWE)

https://hy4.site/benchmarks/hy4-benchmarks

64.3% · raw 64.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
GLM-5.3-FlashZ.ai63.4%unknown effort2026-08-26
All 3 recorded results & sources

63.4% · raw 63.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash

63.4% · raw 63.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

63.4% · raw 63.4 %

Alternative · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. | Same value as current headline; kept as corroboration, non-headline.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Ling 3.1 FlashinclusionAI59.7%unknown effort2026-09-30
All 1 recorded result & sources

59.7% · raw 59.7 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
Qwen3.8 FlashAlibaba (Qwen)58.7%xhigh effort2026-08-26
All 1 recorded result & sources

58.7% · raw 58.7 %

Headline · xhigh effort · Own vendor

Source/record date: 2026-08-26

DeepSWE 1.1 — QwenCloud latest-model page

https://docs.qwencloud.com/developer-guides/getting-started/latest-model
Qwen3.8 Omni FlashAlibaba (Qwen)57.8%unknown effort2026-09-23
All 1 recorded result & sources

57.8% · raw 57.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-23

Qwen3.8-Omni-Flash column on Omni blog DeepSWE 1.1 table (Flash-Next listed 58.7 beside it)

https://qwen.ai/blog?id=qwen3.8-omni-flash
Nex-N2.5-ProNex AGI55.8%unknown effort2026-09-08
All 1 recorded result & sources

55.8% · raw 55.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-08

NexAU harness

https://huggingface.co/nex-agi/Nex-N2.5-Pro
Qwen3.8 27BAlibaba (Qwen)42.2%unknown effort2026-08-14
All 1 recorded result & sources

42.2% · raw 42.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

DeepSWE 1.1; Claude Code harness

https://huggingface.co/Qwen/Qwen3.8-27B
Kolibri-1Aleph Alpha——No recorded result
Ling 3.0 Flash VLinclusionAI——No recorded result
Mercury 2.5Inception——No recorded result
MiMo-V2.6-Distill-Qwen-9BXiaomi——No recorded result

Benchmark source references

Use the fair-comparison guide before interpreting results from different configurations.