AutomationBench v1.0.6 Results Across LLMs
Vendor-reported AutomationBench v1.0.6 results across showcased LLMs. General agent automation across multi-step workflows (forms, apps, scripts) rather than pure coding.
What this table contains
General agent automation across multi-step workflows (forms, apps, scripts) rather than pure coding.
18 models with results · 12 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.
Interactive results & effort preference → · Results over time → · Methodology
| Model | Headline score | Source/record date | Evidence |
|---|---|---|---|
| DeepSeek V4.1 FlashDeepSeek | 54.8%unknown effort | 2026-09-10 | All 1 recorded result & sources54.8% · raw 54.8 % Headline · unknown effort · Third-party Source/record date: 2026-09-10 VentureBeat quoting DeepSeek https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5 |
| MiMo-V2.6-ProXiaomi | 53.1%unknown effort | 2026-09-21 | All 1 recorded result & sources53.1% · raw 53.1 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL |
| Ling 3.1 FlashinclusionAI | 52.5%unknown effort | 2026-09-30 | All 1 recorded result & sources52.5% · raw 52.5 % Headline · unknown effort · Third-party Source/record date: 2026-09-30 Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed. https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym |
| MiMo-V2.6-FlashXiaomi | 52.3%unknown effort | 2026-09-21 | All 1 recorded result & sources52.3% · raw 52.3 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL |
| Gemini 4 ArgonGoogle | 51.3%max effort | 2026-09-30 | All 1 recorded result & sources51.3% · raw 51.3 % Headline · max effort · Own vendor Source/record date: 2026-09-30 Private set via official Zapier public leaderboard; ranks #1 | Highest thinking settings per Google eval methodology. https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/ |
| Muse Spark 1.3Meta | 49.6%max effort | 2026-09-02 | All 1 recorded result & sources49.6% · raw 49.6 % Headline · max effort · Own vendor Source/record date: 2026-09-02 from chart/figure Meta Muse Spark 1.3 scorecard; AutomationBench E2E; max effort https://research.meta.ai/blog/introducing-muse-spark-1-3 |
| GLM-5.3Z.ai | 48.8%unknown effort | 2026-09-10 | All 2 recorded results & sources48.8% · raw 48.8 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-10 As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column) https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash48.2% · raw 48.2 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| GLM-5.3-FlashZ.ai | 48.8%unknown effort | 2026-08-26 | All 1 recorded result & sources48.8% · raw 48.8 % Headline · unknown effort · Own vendor Source/record date: 2026-08-26 https://z.ai/blog/glm-5.3-flash |
| Claude Sonnet 5.5Anthropic | 44.7%max effort | 2026-09-28 | All 1 recorded result & sources44.7% · raw 44.7 % Headline · max effort · Own vendor Source/record date: 2026-09-28 Same Zapier 1.0.6 run, mirrored per project convention https://www.anthropic.com/claude-sonnet-5-5-system-card |
| Nex-N2.5-ProNex AGI | 44.2%unknown effort | 2026-09-08 | All 1 recorded result & sources44.2% · raw 44.2 % Headline · unknown effort · Own vendor Source/record date: 2026-09-08 AutomationBench Public v1.0.6 https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| GPT-6 AstraOpenAI | 41.4%unknown effort | 2026-09-03 | All 2 recorded results & sources41.4% · raw 41.4 % Headline · unknown effort · Own vendor Source/record date: 2026-09-03 https://openai.com/index/gpt-6-astra/41.4% · raw 41.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Claude Opus 5.5Anthropic | 40%max effort | 2026-09-22 | All 2 recorded results & sources40% · raw 40 % Headline · max effort · Own vendor Source/record date: 2026-09-22 Anthropic announce table; Zapier AutomationBench without fallback; adaptive thinking max effort https://www.anthropic.com/claude-opus-5-542.5% · raw 42.5 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 40.0 (first-party); kept with provenance, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Qwen3.8 Max (0902)Alibaba (Qwen) | 39.8%unknown effort | 2026-09-08 | All 1 recorded result & sources39.8% · raw 39.8 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-08 As reported in Nex-N2.5-Pro model card comparison table https://huggingface.co/nex-agi/Nex-N2.5-Pro |
| GPT-6.1 SolOpenAI | 36.1%max effort | 2026-09-29 | All 2 recorded results & sources36.1% · raw 36.1 % Headline · max effort · Own vendor Source/record date: 2026-09-29 Max effort; +2.9 over GPT-6 Sol; -5.3 vs Astra 41.4; -6.4 vs Opus 5.5 with fallbacks 42.5 https://openai.com/index/introducing-gpt-6-1-sol/31.7% · raw 31.7 % Alternative · medium effort · Own vendor Source/record date: 2026-09-29 Medium effort: +2.2 over Opus 5.5 (29.5), +4.8 over GPT-6 Sol same setting https://openai.com/index/introducing-gpt-6-1-sol/ |
| Hy4 previewTencent | 32.1%unknown effort | 2026-08-28 | All 1 recorded result & sources32.1% · raw 32.1 % Headline · unknown effort · Own vendor Source/record date: 2026-08-28 Vendor chart transcription https://hy4.site/benchmarks/hy4-benchmarks |
| Claude Fable 5.1Anthropic | 31.4%unknown effort | 2026-09-03 | All 2 recorded results & sources31.4% · raw 31.4 % Headline · unknown effort · Peer vendor Source/record date: 2026-09-03 As reported by OpenAI https://openai.com/index/gpt-6-astra/31.4% · raw 31.4 % Alternative · unknown effort · Peer vendor Source/record date: 2026-09-30 As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline. https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif |
| Kimi K3Moonshot AI | 30.8%max effort | 2026-07-16 | All 1 recorded result & sources30.8% · raw 30.8 % Headline · max effort · Own vendor Source/record date: 2026-07-16 AutomationBench 600-task public subset (mapped to automationbench-v1.0.6) https://github.com/MoonshotAI/Kimi-K3 |
| MiMo-V2.6-Distill-Qwen-9BXiaomi | 30.3%unknown effort | 2026-09-21 | All 1 recorded result & sources30.3% · raw 30.3 % Headline · unknown effort · Own vendor Source/record date: 2026-09-21 Released SFT checkpoint, avg@1; MiMo-V2.6 tech report Table 6 (https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf) and HF card Evaluation table. Also plotted (SFT line) in radar figure on https://mimo.mi.com/docs/en-US/news/latest/v2-6. Eval effort not specified (thinking model). https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B |
| Gemini 3.8 FlashGoogle | — | — | No recorded result |
| GPT-5.6 TerraOpenAI | — | — | No recorded result |
| GPT-6 LunaOpenAI | — | — | No recorded result |
| Grok 4.7xAI | — | — | No recorded result |
| Kolibri-1Aleph Alpha | — | — | No recorded result |
| Ling 3.0 Flash VLinclusionAI | — | — | No recorded result |
| Mercury 2.5Inception | — | — | No recorded result |
| Naive-N0.5-FlashNaiveAI | — | — | No recorded result |
| Pareto 26.10 PreviewUnbiased | — | — | No recorded result |
| Qwen3.8 27BAlibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 FlashAlibaba (Qwen) | — | — | No recorded result |
| Qwen3.8 Omni FlashAlibaba (Qwen) | — | — | No recorded result |
Benchmark source references
- OpenAI / vendor AutomationBench tables · vendor_release
- OpenAI GPT-6 Astra system card · vendor_system_card
Use the fair-comparison guide before interpreting results from different configurations.