Agents' Last Exam Results Across LLMs
Vendor-reported Agents' Last Exam results across showcased LLMs. Hard multi-step agent exam stressing planning, tools, and failure recovery at the frontier.
What this table contains
Hard multi-step agent exam stressing planning, tools, and failure recovery at the frontier.
14 models with results · 16 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.
Interactive results & effort preference → · Results over time → · Methodology
| Model | Headline score | Source/record date | Evidence |
|---|---|---|---|
| GPT-6 AstraOpenAI | 59.3%unknown effort | 2026-09-03 | All 3 recorded results & sources59.3% · raw 59.3 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/33.3% · raw 33.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/34.2% · raw 34.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 59.3 (first-party); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| GPT-5.6 TerraOpenAI | 50.4%unknown effort | 2026-07-09 | All 1 recorded result & sources50.4% · raw 50.4 % Headline · unknown effort · Own vendor Source/record date: 2026-07-09 https://openai.com/index/gpt-5-6 |
| Gemini 4 ArgonGoogle | 39.5%max effort | 2026-09-30 | All 1 recorded result & sources39.5% · raw 39.5 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Self-computed ALE-Claw harness 5h, binary pass rate, safety filters on | Highest thinking settings per Google eval methodology. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Claude Opus 5.5Anthropic | 34.3%unknown effort | 2026-09-27 | All 2 recorded results & sources34.3% · raw 34.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/38.2% · raw 38.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 34.3 (spillover); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Muse Spark 1.3Meta | 33.3%unknown effort | 2026-09-27 | All 1 recorded result & sources33.3% · raw 33.3 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| Naive-N0.5-FlashNaiveAI | 32.4%unknown effort | 2026-09-27 | All 1 recorded result & sources32.4% · raw 32.4 % Headline · unknown effort · Own vendor Source/record date: 2026-09-27 ALE-CLI per card sources; harness Claude Code 2.1.207. https://naive.ai/en/research/ |
| DeepSeek V4.1 FlashDeepSeek | 31.8%unknown effort | 2026-09-10 | All 1 recorded result & sources31.8% · raw 31.8 % Headline · unknown effort · Own vendor Source/record date: 2026-09-10 Agent's Last Exam Pass@1 https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash |
| MiMo-V2.6-ProXiaomi | 31.6%unknown effort | 2026-09-21 | All 1 recorded result & sources31.6% · raw 31.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL |
| GLM-5.3Z.ai | 28.5%unknown effort | 2026-08-14 | All 2 recorded results & sources28.5% · raw 28.5 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 https://docs.z.ai/guides/llm/glm-5.328.5% · raw 28.5 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| Kimi K3Moonshot AI | 28.3%max effort | 2026-07-16 | All 2 recorded results & sources28.3% · raw 28.3 % Headline · max effort · Own vendor Source/record date: 2026-07-16 Agents' Last Exam primary pass-rate; Kimi Code harness; from official leaderboard as of 2026-07-23 https://github.com/MoonshotAI/Kimi-K328.3% · raw 28.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| MiMo-V2.6-FlashXiaomi | 27.6%unknown effort | 2026-09-21 | All 1 recorded result & sources27.6% · raw 27.6 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL |
| GLM-5.3-FlashZ.ai | 26.3%unknown effort | 2026-08-26 | All 2 recorded results & sources26.3% · raw 26.3 % Headline · unknown effort · Own vendor Source/record date: 2026-08-26 https://z.ai/blog/glm-5.3-flash26.3% · raw 26.3 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| Hy4 previewTencent | 22.8%unknown effort | 2026-08-28 | All 2 recorded results & sources22.8% · raw 22.8 % Headline · unknown effort · Own vendor Source/record date: 2026-08-28 ALE-CLI; vendor chart transcription https://hy4.site/benchmarks/hy4-benchmarks22.8% · raw 22.8 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-27 As reported by NaiveAI. https://naive.ai/en/research/ |
| Qwen3.8 27BAlibaba (Qwen) | 20.4%unknown effort | 2026-08-14 | All 1 recorded result & sources20.4% · raw 20.4 % Headline · unknown effort · Own vendor Source/record date: 2026-08-14 Agents' Last Exam Pass@1 (Score also reported 42.9 on card) https://huggingface.co/Qwen/Qwen3.8-27B |
| Claude Fable 5.1Anthropic | — | — | No recorded result |
| Claude Sonnet 5.5Anthropic | — | — | No recorded result |
| Gemini 3.8 FlashGoogle | — | — | No recorded result |
| GPT-6 LunaOpenAI | — | — | No recorded result |
| GPT-6.1 SolOpenAI | — | — | No recorded result |
| Grok 4.7xAI | — | — | No recorded result |
| Kolibri-1Aleph Alpha | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | — | — | No recorded result |
| Ling 3.1 FlashinclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | — | — | No recorded result |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | — | — | No recorded result |
| Nex-N2.5-ProNex AGI | — | — | No recorded result |
| Pareto 26.10 PreviewUnbiased | — | — | No recorded result |
| Qwen3.8 FlashAlibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Max (0902)Alibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Omni FlashAlibaba (Qwen) | — | — | No recorded result |
Benchmark source references
- OpenAI GPT-6 Astra system card · vendor_system_card
- OpenAI GPT-6 Sol & Luna announce · vendor_release
Use the fair-comparison guide before interpreting results from different configurations.