Examenos

Agents' Last Exam Results Across LLMs

Vendor-reported Agents' Last Exam results across showcased LLMs. Hard multi-step agent exam stressing planning, tools, and failure recovery at the frontier.

Official / vendor Higher is better · unit: %

What this table contains

Hard multi-step agent exam stressing planning, tools, and failure recovery at the frontier.

14 models with results · 16 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.

Interactive results & effort preference → · Results over time → · Methodology

Agents' Last Exam: vendor-reported results only
ModelHeadline scoreSource/record dateEvidence
GPT-6 AstraOpenAI59.3%unknown effort2026-09-03
All 3 recorded results & sources

59.3% · raw 59.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/

33.3% · raw 33.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

34.2% · raw 34.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 59.3 (first-party); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
GPT-5.6 TerraOpenAI50.4%unknown effort2026-07-09
All 1 recorded result & sources

50.4% · raw 50.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-07-09

https://openai.com/index/gpt-5-6
Gemini 4 ArgonGoogle39.5%max effort2026-09-30
All 1 recorded result & sources

39.5% · raw 39.5 %

Headline · max effort · Own vendor

Source/record date: 2026-09-30

Self-computed ALE-Claw harness 5h, binary pass rate, safety filters on | Highest thinking settings per Google eval methodology.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Claude Opus 5.5Anthropic34.3%unknown effort2026-09-27
All 2 recorded results & sources

34.3% · raw 34.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/

38.2% · raw 38.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 34.3 (spillover); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Muse Spark 1.3Meta33.3%unknown effort2026-09-27
All 1 recorded result & sources

33.3% · raw 33.3 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
Naive-N0.5-FlashNaiveAI32.4%unknown effort2026-09-27
All 1 recorded result & sources

32.4% · raw 32.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-27

ALE-CLI per card sources; harness Claude Code 2.1.207.

https://naive.ai/en/research/
DeepSeek V4.1 FlashDeepSeek31.8%unknown effort2026-09-10
All 1 recorded result & sources

31.8% · raw 31.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-10

Agent's Last Exam Pass@1

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
MiMo-V2.6-ProXiaomi31.6%unknown effort2026-09-21
All 1 recorded result & sources

31.6% · raw 31.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
GLM-5.3Z.ai28.5%unknown effort2026-08-14
All 2 recorded results & sources

28.5% · raw 28.5 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

https://docs.z.ai/guides/llm/glm-5.3

28.5% · raw 28.5 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
Kimi K3Moonshot AI28.3%max effort2026-07-16
All 2 recorded results & sources

28.3% · raw 28.3 %

Headline · max effort · Own vendor

Source/record date: 2026-07-16

Agents' Last Exam primary pass-rate; Kimi Code harness; from official leaderboard as of 2026-07-23

https://github.com/MoonshotAI/Kimi-K3

28.3% · raw 28.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
MiMo-V2.6-FlashXiaomi27.6%unknown effort2026-09-21
All 1 recorded result & sources

27.6% · raw 27.6 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
GLM-5.3-FlashZ.ai26.3%unknown effort2026-08-26
All 2 recorded results & sources

26.3% · raw 26.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash

26.3% · raw 26.3 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
Hy4 previewTencent22.8%unknown effort2026-08-28
All 2 recorded results & sources

22.8% · raw 22.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-28

ALE-CLI; vendor chart transcription

https://hy4.site/benchmarks/hy4-benchmarks

22.8% · raw 22.8 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-27

As reported by NaiveAI.

https://naive.ai/en/research/
Qwen3.8 27BAlibaba (Qwen)20.4%unknown effort2026-08-14
All 1 recorded result & sources

20.4% · raw 20.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-14

Agents' Last Exam Pass@1 (Score also reported 42.9 on card)

https://huggingface.co/Qwen/Qwen3.8-27B
Claude Fable 5.1Anthropic——No recorded result
Claude Sonnet 5.5Anthropic——No recorded result
Gemini 3.8 FlashGoogle——No recorded result
GPT-6 LunaOpenAI——No recorded result
GPT-6.1 SolOpenAI——No recorded result
Grok 4.7xAI——No recorded result
Kolibri-1Aleph Alpha——No recorded result
Ling 3.0 Flash VLinclusionAI——No recorded result
Ling 3.1 FlashinclusionAI——No recorded result
Mercury 2.5Inception——No recorded result
MiMo-V2.6-Distill-Qwen-9BXiaomi——No recorded result
Nex-N2.5-ProNex AGI——No recorded result
Pareto 26.10 PreviewUnbiased——No recorded result
Qwen3.8 FlashAlibaba (Qwen)——No recorded result
Qwen3.8 Max (0902)Alibaba (Qwen)——No recorded result
Qwen3.8 Omni FlashAlibaba (Qwen)——No recorded result

Benchmark source references

Use the fair-comparison guide before interpreting results from different configurations.