Examenos

AutomationBench v1.0.6 Results Across LLMs

Vendor-reported AutomationBench v1.0.6 results across showcased LLMs. General agent automation across multi-step workflows (forms, apps, scripts) rather than pure coding.

Official / vendor Higher is better · unit: %

What this table contains

General agent automation across multi-step workflows (forms, apps, scripts) rather than pure coding.

18 models with results · 12 missing · 30 showcased models. Missing scores are not zero or failed tests. Headline selection is the default; source dates and setups can differ.

Interactive results & effort preference → · Results over time → · Methodology

AutomationBench v1.0.6: vendor-reported results only
ModelHeadline scoreSource/record dateEvidence
DeepSeek V4.1 FlashDeepSeek54.8%unknown effort2026-09-10
All 1 recorded result & sources

54.8% · raw 54.8 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-10

VentureBeat quoting DeepSeek

https://venturebeat.com/technology/deepseek-v4-1-flash-debuts-with-0-003-1m-off-peak-cached-input-rate-and-benchmarks-eclipsing-gpt-5-6-sol-claude-opus-5
MiMo-V2.6-ProXiaomi53.1%unknown effort2026-09-21
All 1 recorded result & sources

53.1% · raw 53.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
Ling 3.1 FlashinclusionAI52.5%unknown effort2026-09-30
All 1 recorded result & sources

52.5% · raw 52.5 %

Headline · unknown effort · Third-party

Source/record date: 2026-09-30

Via ThreatFrontier transcription of Ant launch table; original pixels not yet re-read. Effort undisclosed.

https://threatfrontier.com/articles/ling-3-1-flash-ant-group-best-flash-model-yet-cybergym
MiMo-V2.6-FlashXiaomi52.3%unknown effort2026-09-21
All 1 recorded result & sources

52.3% · raw 52.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
Gemini 4 ArgonGoogle51.3%max effort2026-09-30
All 1 recorded result & sources

51.3% · raw 51.3 %

Headline · max effort · Own vendor

Source/record date: 2026-09-30

Private set via official Zapier public leaderboard; ranks #1 | Highest thinking settings per Google eval methodology.

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/
Muse Spark 1.3Meta49.6%max effort2026-09-02
All 1 recorded result & sources

49.6% · raw 49.6 %

Headline · max effort · Own vendor

Source/record date: 2026-09-02

from chart/figure Meta Muse Spark 1.3 scorecard; AutomationBench E2E; max effort

https://research.meta.ai/blog/introducing-muse-spark-1-3
GLM-5.3Z.ai48.8%unknown effort2026-09-10
All 2 recorded results & sources

48.8% · raw 48.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-10

As reported in DeepSeek-V4.1-Flash HF comparison table (GLM-5.3 column)

https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

48.2% · raw 48.2 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
GLM-5.3-FlashZ.ai48.8%unknown effort2026-08-26
All 1 recorded result & sources

48.8% · raw 48.8 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-26

https://z.ai/blog/glm-5.3-flash
Claude Sonnet 5.5Anthropic44.7%max effort2026-09-28
All 1 recorded result & sources

44.7% · raw 44.7 %

Headline · max effort · Own vendor

Source/record date: 2026-09-28

Same Zapier 1.0.6 run, mirrored per project convention

https://www.anthropic.com/claude-sonnet-5-5-system-card
Nex-N2.5-ProNex AGI44.2%unknown effort2026-09-08
All 1 recorded result & sources

44.2% · raw 44.2 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-08

AutomationBench Public v1.0.6

https://huggingface.co/nex-agi/Nex-N2.5-Pro
GPT-6 AstraOpenAI41.4%unknown effort2026-09-03
All 2 recorded results & sources

41.4% · raw 41.4 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-03

https://openai.com/index/gpt-6-astra/

41.4% · raw 41.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Claude Opus 5.5Anthropic40%max effort2026-09-22
All 2 recorded results & sources

40% · raw 40 %

Headline · max effort · Own vendor

Source/record date: 2026-09-22

Anthropic announce table; Zapier AutomationBench without fallback; adaptive thinking max effort

https://www.anthropic.com/claude-opus-5-5

42.5% · raw 42.5 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Differs from current headline 40.0 (first-party); kept with provenance, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Qwen3.8 Max (0902)Alibaba (Qwen)39.8%unknown effort2026-09-08
All 1 recorded result & sources

39.8% · raw 39.8 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-08

As reported in Nex-N2.5-Pro model card comparison table

https://huggingface.co/nex-agi/Nex-N2.5-Pro
GPT-6.1 SolOpenAI36.1%max effort2026-09-29
All 2 recorded results & sources

36.1% · raw 36.1 %

Headline · max effort · Own vendor

Source/record date: 2026-09-29

Max effort; +2.9 over GPT-6 Sol; -5.3 vs Astra 41.4; -6.4 vs Opus 5.5 with fallbacks 42.5

https://openai.com/index/introducing-gpt-6-1-sol/

31.7% · raw 31.7 %

Alternative · medium effort · Own vendor

Source/record date: 2026-09-29

Medium effort: +2.2 over Opus 5.5 (29.5), +4.8 over GPT-6 Sol same setting

https://openai.com/index/introducing-gpt-6-1-sol/
Hy4 previewTencent32.1%unknown effort2026-08-28
All 1 recorded result & sources

32.1% · raw 32.1 %

Headline · unknown effort · Own vendor

Source/record date: 2026-08-28

Vendor chart transcription

https://hy4.site/benchmarks/hy4-benchmarks
Claude Fable 5.1Anthropic31.4%unknown effort2026-09-03
All 2 recorded results & sources

31.4% · raw 31.4 %

Headline · unknown effort · Peer vendor

Source/record date: 2026-09-03

As reported by OpenAI

https://openai.com/index/gpt-6-astra/

31.4% · raw 31.4 %

Alternative · unknown effort · Peer vendor

Source/record date: 2026-09-30

As reported by Google Gemini 4 Argon chart; peer harness notes in Google eval methodology | Same value as current headline; kept as corroboration, non-headline.

https://storage.googleapis.com/gweb-uniblog-publish-prod/original_images/gemini-4-argon_table_blog.gif
Kimi K3Moonshot AI30.8%max effort2026-07-16
All 1 recorded result & sources

30.8% · raw 30.8 %

Headline · max effort · Own vendor

Source/record date: 2026-07-16

AutomationBench 600-task public subset (mapped to automationbench-v1.0.6)

https://github.com/MoonshotAI/Kimi-K3
MiMo-V2.6-Distill-Qwen-9BXiaomi30.3%unknown effort2026-09-21
All 1 recorded result & sources

30.3% · raw 30.3 %

Headline · unknown effort · Own vendor

Source/record date: 2026-09-21

Released SFT checkpoint, avg@1; MiMo-V2.6 tech report Table 6 (https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf) and HF card Evaluation table. Also plotted (SFT line) in radar figure on https://mimo.mi.com/docs/en-US/news/latest/v2-6. Eval effort not specified (thinking model).

https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Gemini 3.8 FlashGoogle——No recorded result
GPT-5.6 TerraOpenAI——No recorded result
GPT-6 LunaOpenAI——No recorded result
Grok 4.7xAI——No recorded result
Kolibri-1Aleph Alpha——No recorded result
Ling 3.0 Flash VLinclusionAI——No recorded result
Mercury 2.5Inception——No recorded result
Naive-N0.5-FlashNaiveAI——No recorded result
Pareto 26.10 PreviewUnbiased——No recorded result
Qwen3.8 27BAlibaba (Qwen)——No recorded result
Qwen3.8 FlashAlibaba (Qwen)——No recorded result
Qwen3.8 Omni FlashAlibaba (Qwen)——No recorded result

Benchmark source references

Use the fair-comparison guide before interpreting results from different configurations.